<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Classifier ensemble for biomedical document retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Manabu Torii</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hongfang Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Biostatistics</institution>
          ,
          <addr-line>Bioinformatics, and Biomathematics</addr-line>
          ,
          <institution>Georgetown University Medical Center</institution>
          ,
          <addr-line>4000 Reservoir Rd, NW, Washington, DC 20057</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>§Corresponding author</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>important aspect of research activities in the biomedical domain. Machine Learning (ML)
techniques have been explored to retrieve relevant articles from a large literature archive (i.e.,
classifying articles into relevant and irrelevant classes), and to accelerating the literature
review process. Meanwhile, an ensemble classifier, a system that assigns classes based on the
outputs of multiple classifiers, tends to be more robust and has better performance than each
individual classifier. Ensemble classifiers are often composed of classifiers trained on
different training sets (e.g., sampled data sets) or of those using different ML algorithms. In
this paper, we propose a simple ensemble approach where an ensemble is composed of
classifiers using different feature sets for an ML algorithm. We evaluated the approach using
Support Vector Machine (SVM) on two publicly available collections of MEDLINE citations,
the Post-translational modification (PTM) data sets and the Immune Epitope Database (IEDB)
data sets, that resulted from biomedical database curation projects.</p>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>The evaluation showed that ensemble classifiers outperformed their constituent classifiers as
measured by both area under ROC curve (AUC) and precision/recall break-even-point (BEP),
provided with enough training data. We observed that the performance of SVM ensembles
were competitive or better than the best results previously reported for the data sets used.</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>The proposed ensemble approach was found to be effective in improving performance of
SVM</p>
      <p>classifiers. The approach is also simple and easy-to-deploy in document
classification/retrieval tasks. However, improvement of classifiers through the current
approach is still modest. We plan to explore different ways to derive and combine constituent
classifiers, and continue our investigation over other data sets.</p>
      <sec id="sec-3-1">
        <title>Background</title>
        <p>
          Due to rich information embedded in published biomedical articles, literature review has
become an increasingly important aspect of research activities in the biomedical domain, e.g.,
[
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ]. One of the initial steps in literature review is document retrieval, i.e., gathering
documents relevant to the target topic from a large literature archive such as MEDLINE. To
mitigate human effort in retrieving relevant articles, there has been a growing interest in
automatic document retrieval. It has become an active research area in biomedical text
mining. For example, Genomics Track (2003–2007) of Text Retrieval Conference (TREC)
has dedicated to the evaluation of biomedical document retrieval systems [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Also, one of the
tasks in BioCreAtIvE II1 (http://biocreative.sourceforge.net/biocreative_2.html) asked
participants to order biomedical articles based on their relevance to protein interaction
annotation.
        </p>
        <p>
          Document retrieval is to prioritize (i.e., order) documents according to their relevance
to the target topic. Considering the target documents as positive instances and others as
negative instances, the priority order of documents can be obtained using a classifier that
yields a confidence score in assigning positive/negative classes to documents. Machine
Learning (ML) approaches such as Naïve Bayes and Support Vector Machine (SVM) have
enjoyed great success in document classification and retrieval [
          <xref ref-type="bibr" rid="ref22 ref4">4</xref>
          ]. In the biomedical domain,
ML classifiers have been considered for document retrieval in database curation projects
[59], and therefore the efficiency of database curation can be enhanced by improving document
classifiers. For a specific application, improvement of ML classifiers can be attempted
through incorporation of task-specific features/heuristics and/or elaborated domain-specific
features [
          <xref ref-type="bibr" rid="ref3 ref7 ref8">3, 7, 8</xref>
          ]. Alternatively, or in conjunction with such effort, classifier performance can
be improved by combining multiple classifiers, i.e., ensemble of classifiers [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>In this study, we consider a simple and easy-to-deploy ensemble approach for
biomedical document classification tasks where an ensemble is composed of classifiers built
1 Protein Interaction Article Sub-task (IAS) of the Protein-Protein Interaction task (PPI).
with different sets of features. The goal of this study is two fold: i) examine the effectiveness
of our ensemble approach for classification of MEDLINE citations; and ii) report
classification performance on publicly available data sets in the domain.</p>
        <p>
          In the following, we first provide background information for classifier ensemble.
Next, we describe two publicly available data sets used in this study, the Post-translational
modification (PTM) data sets [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and the Immune Epitope Database (IEDB) data sets [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ],
resulted from actual biomedical database curation projects.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Classifier ensemble</title>
      <p>
        It has been observed that “accurate and diverse” classifiers make an ensemble classifier that
outperforms a single classifier [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. A classifier is “accurate” if it performs better than a
random classifier and classifiers are “diverse” if they do not make the same classification
mistakes. Popular ensemble approaches include bagging and boosting. In the bagging
approach [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], constituent classifiers of an ensemble are trained on data sets sampled from the
training data. In the boosting approach [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], constituent classifiers are trained sequentially, in
which misclassified instances are assigned more weights during the training of the next
classifier. The performance of bagging and boosting is dependent on the ML algorithm used.
For example, in the newswire domain, Dong and Han [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] examined the utility of different
ensemble methods including bagging and boosting for document classification. In their work,
although a boosted Naïve Bayes classifier outperformed a single Naïve Bayes classifier, a
boosted SVM classifier performed worse than a single SVM classifier. In fact, since SVM
does not depend on weights/frequencies of instances, boosting may not be an appropriate
choice for SVM (see, e.g., [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]). In Dong and Han, there was little or no improvement
reported also for bagging with Naïve Bayes and with SVM.
      </p>
      <p>In bagging or boosting, constituent classifiers of an ensemble are built by varying
training data sets (i.e., using sampled documents or documents with different weights). In this
study, we propose an ensemble approach that builds constituent classifiers by varying the size
of feature words used (i.e., by varying feature vectors, but using the same document set and
the same ML algorithm).</p>
    </sec>
    <sec id="sec-5">
      <title>Publicly available data sets for biomedical document classification/retrieval</title>
      <p>
        PTM data sets2 – The PTM data sets developed at Protein Information Resource (PIR)
consist of five collections of MEDLINE citations for five different PTM types (acetylation,
glycosylation, hidroxylation, methylation, and phosphorylation), which were labelled by
domain experts as either positive or negative at the level of abstract or full-length article.
Small parts of the abstract-level PTM data sets have been used in a document retrieval study
by Han et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which specifically investigated document retrieval for small data sets. We
used these small data sets for performance comparison purposes. Of five data sets used in Han
et al., we used two data sets, the acetylation and phosphorylation data sets, where there are at
least 50 positive documents3 (Table 1). In the work reported in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Naïve Bayes classifiers
exploiting substring features outperformed SVM classifiers on these data sets.
IEDB data sets [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]4 – The IEDB data sets were developed during the annotation of epitopes
from four different sources with a view to populating the Immune Epitope Database. Thus,
the data sets consist of four sets of MEDLINE citations (abstracts and titles), which were
retrieved from PubMed using “complex queries.” Citations were, then, manually classified as
positive and negative documents. In this study, following the study of Wang et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], we
combined the four data sets and derived a corpus of 20,907 MEDLINE citations. In [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the
authors reported that Naïve Bayes classifiers outperformed SVM
classifiers in the
2 The PTM data sets are available at the PIR iProLink web site
(http://pir.georgetown.edu/cgibin/ipkLitFt.pl?stat=12). In this study, we used smaller parts of these data sets, which were introduced
in the study by Han et al. (http://www.ist.temple.edu/PIRsupplement/).
3 The acetylation data set contains 55 references to MEDLINE citations (PMIDs) for positive
documents. However, when the citations were downloaded, six of them were no longer accessible with
the listed PMIDs (see the caption of Table 1).
4 http://www.biomedcentral.com/1471-2105/8/269/additional/
experiments, and the best classification performance was obtained using domain- and
taskspecific features. Summaries of the data sets are found in Table 1.
      </p>
      <sec id="sec-5-1">
        <title>Methods</title>
        <p>
          We used SVM Light [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] to derive SVM classifiers in this study. Specifically, we used radial
basis function (RBF) with gamma value of 1.0 as the kernel, with the default setting for the
regularization parameter C in SVM Light. We used this setting for its yielding the good
performance in preliminary experiments using the small portion of the data sets. For
classification features, only normalized words in documents are used without any task- or
domain-specific features/heuristics.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Word normalization</title>
      <p>
        For SVM classifiers, we used words in documents as features, except for stop words listed in
the NCBI stopword list5 and rare words that appear in less than three documents in a training
data set. Words are defined as consecutive alphabet letters, numbers, hyphens (-), or slash (/),
e.g., IL-1 is regarded as one word without being tokenized into smaller sequences. All words
were processed with our implementation of S-stemmer [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], which converts plural nouns and
third person singular verbs into their base forms, e.g., receptors Æ receptor, studies Æ study,
IFNs Æ IFN. Also, within each word, alphabet letters were lowercased, a number sequence
was converted to a token “DIGIT”, and Greek alphabets to a token “GREEK”, e.g., KappaB
Æ
      </p>
      <p>GREEKb and Ras1 Æ rasDIGIT.</p>
    </sec>
    <sec id="sec-7">
      <title>Feature vector generation</title>
      <p>
        To select classification features among normalized words identified in a training data set, we
used information gain (IG), also known as expected mutual information, commonly used in
text classification, e.g., [
        <xref ref-type="bibr" rid="ref18 ref19 ref7 ref8">7, 8, 18, 19</xref>
        ]. IG is calculated:
IG(D, w) = H (D) − ∑
d∈Dw+ D
      </p>
      <p>Dw+</p>
      <p>H (Dw+ )− ∑
d∈Dw- D</p>
      <p>Dw−</p>
      <p>H (Dw− )
5 http://www.ncbi.nlm.nih.gov/books/bv.fcgi?indexed=google&amp;rid=helppubmed.table.pubmedhelp.T43
where H(•) is an entropy for having different classes given a document set, and Dw+ and
Dware partitions of a document set D, each of which consists of documents containing (w+) or
not containing (w-) word w, respectively. In building classifiers, we consider the top R
percent of words according to the IG measure. In this study, ten different R values were
considered: R = 1, 4, 9, …, 100, i.e., R=n2 for n = 1, 2, ..., 10. Among the words with higher
IG values, differences of IG values for any two words tend to be greater, while among those
with lower IG values, such differences tend to be smaller6. This motivated us to use the above
choices of R in deriving a set of “diverse” classifiers. Namely, we varied the percentage less
for the smaller values of R, e.g., 1 Æ 4 Æ 9 Æ …, but varied the percentage more for the
larger values of R, e.g., … Æ 64 Æ 81 Æ 100.</p>
      <p>
        A document was represented with a set of selected feature words found therein (a
bag-of-words approach) in a vector format. Each value in a vector associated with a feature
word is a frequency of words (i.e., terms) in the document (TF) weighted by the inverse of the
document frequency (IDF), i.e., TFi, j × log(IDFj ) for word j in document i. As in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], feature
vectors are normalized so that the Euclidean norm of a vector is 1.0.
      </p>
    </sec>
    <sec id="sec-8">
      <title>Classifier ensemble</title>
      <p>Given a set of input documents, a classifier will assign a numeric value to each document,
which is regarded as a confidence score for a positive (or negative) class. With a single
classifier, documents are ranked according to assigned confidence scores. With multiple
classifiers, documents can be raked according to the summation of confidence scores assigned
to each document.
6 To be clear, let w1, w2, w3, .... wn be an ordered list of all words in a corpus such that IG(w1) ≥ IG(w2) ≥
… ≥ IG(wn), where IG(•) is a function mapping a word to an IG value. Then, roughly speaking, we
observed IG(wi)-IG(wi+1) &gt; IG(wi+1)-IG(wi+2) for i=1…n-1. In other words, comparatively speaking, an
SVM classifier exploiting 1% of top IG-value words, say SVMR=1, can differ much from another
classifier SVMR=4, while SVMR=97 and SVMR=100 may be almost the same in terms of their performance.</p>
      <p>In order to derive multiple classifiers from one training data set, we built each
classifier by varying the threshold value R in selecting features. As detailed in the previous
sub-section, a set of ten classifiers are derived using ten different R values (i.e., R=1, 4, 9, 16,
…, 100). Among these single classifiers, a group of two (R=1 and 4), three (R=1, 4, and 9),
four (R=1, 4, 9, and 16), …, ten (R=1, 4, 9, … 100) classifiers were selected so that each
group of classifiers makes an ensemble classifier, i.e., nine ensemble classifiers.</p>
      <p>n0n1</p>
    </sec>
    <sec id="sec-9">
      <title>Classifier evaluation</title>
      <p>
        Two measures were used to evaluate the performance of classifiers: Area under ROC curve
(AUC) and Precision/recall break-even-point (BEP). Given an ordering of documents by a
classifier, AUC is interpreted as the probability that the rank of a positive document d1 is
greater (i.e., more likely to be positive) than that of a negative document d0, where d1 and d0
are documents randomly selected from positive and negative document sets, respectively. The
higher the AUC value, the better the classifier is. After documents (of size n) are ranked from
1 (least likely to be positive) to n (most likely to be positive), AUC can be calculated as
AUC = S − n1(n1 +1) 2 , where S is the sum of the ranks assigned to positive documents, and n0
and n1 are the numbers of negative and positive documents, respectively (see the details in
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]). BEP is a precision (or a recall)7 obtained for a classification threshold where the
precision and the recall become equal (see, e.g., [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]).
      </p>
      <p>While AUC can provide a summary of the overall ordering of positive and negative
documents by a classifier, when comparing classifiers from literature reviewers’ point of
view, a higher AUC value does not necessarily imply the better utility of the classifier. For
example, suppose reviewers can go over at most 100 articles among the list of articles
7 For a fixed threshold on confidence scores assigned by a classifier, documents are classified into two
classes. Then, precision is the number of true positives divided by the total number of true positives
and false positives. Recall is the number of true positives divided by the total number of positive
instances.
retrieved at a time. Then, for reviewers, changes in document ordering matter only when they
take places within the first 100 documents. In this respect, BEP may be the more appropriate
performance measure.</p>
      <p>Each classifier was evaluated in m repeated n-fold cross-validation, and average AUC
and BEP over m×n runs were calculated, i.e., a document collection was split into n
equallysized partitions, and classifiers trained on n-1 partitions were evaluated over the remaining
one partition, which was repeated n times using a different partition as a test set each time.
Then, such n-fold cross-validation test was repeated m times over the data collection. For
each data set, the same partitioning was used for all the evaluated classifiers. To make the
results comparable to the previously reported results, different pairs of m and n were used for
the data sets. (m=20 and n=5 for the PTM data sets, and m=1 and n=10 for the IEDB data
sets).</p>
      <sec id="sec-9-1">
        <title>Results and discussion</title>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Results on the PTM data sets</title>
      <p>
        On the two PTM data sets (acetylation and phosphorylation), we evaluated ten single SVM
classifiers and nine ensemble classifiers as detailed in the Method section. For each single and
ensemble classifier, we repeated 5-fold cross-validation tests 20 times as in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and calculated
average AUC and BEP measures of, thus, 100 runs. In each run, there were about 2,500 and
1,800 unique words in the training set portion of the acetylation and phosphorylation data
sets, respectively. R% of the unique words were used in training a single classifier. The
results of the experiments are reported in Table 2.
      </p>
      <p>
        We found inclusion of excessive word features was harmful for these small data sets,
although SVM can usually exploit a large number of word features including those that are
less informative (e.g., in terms of IG) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. For both of the data sets, the best performance of
single SVM classifiers was obtained at R = 4 in term of AUC measure, and R = 9 in terms of
BEP (Table 2). The best performance of ensemble classifiers was obtained when five single
classifiers (R=1, 4, 9, 16 and 25) were combined for the acetylation data set, and when four
classifiers (R=1, 4, 9 and 16) were combined for the phosphorylation data set. For both of the
data sets, performance of ensemble classifiers was consistently better than single classifiers in
terms of both AUC and BEP.
      </p>
      <p>
        We found the performance of single SVM classifiers and that of the most comparable
SVM classifier in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] (WB-SVM-IG) were still very different. While further investigation is
needed, this difference may be attributed to different classifier settings (e.g., a linear kernel
function in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] vs. an RBF kernel function in our experiments for SVM), document
representation (e.g., we used normalized TF-IDF vectors in our experiments, while it is not
clear in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]), and/or threshold settings in feature selection (i.e., a fixed threshold, IG &gt; 0.02, in
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]).
      </p>
      <p>We also examined the applicability of the ensemble approach on the glycosylation,
hydroxylation and methylation data sets used in Han et al, where there are very small
numbers of positive instances (Table 1). On these data sets, performance of ensemble
classifiers was no better than that of single classifiers, or sometimes even worse. On the
hydroxylation and methylation data sets where there are especially small numbers of positive
and negative documents, BEP of single classifiers were low (e.g., an average BEP of ten
single classifiers was 0.24 and 0.47 for the hydroxylation and the methylation data set,
respectively). We assumed that such classifiers were not reliable enough to contribute to an
ensemble classifier.</p>
    </sec>
    <sec id="sec-11">
      <title>Results on the IEDB data sets</title>
      <p>On the combined IEDB data sets, ten single SVM classifiers and nine ensemble classifiers
were evaluated just like on the PTM data sets. Compared to the PTM data sets used by Han et
al., there are a much larger number of documents in this data set (20,907 MEDLINE citations
as opposed to 916 or 457 citations), and we obtained stable results in a ten-fold
crossvalidation test. In each fold, there were about 22,000 unique words in the training set portion
of the data set, R% of which were used as features in training a single classifier. The results
are shown in Table 3. Figure 1 shows how AUC and BEP change as R changes for single and
ensemble classifiers. Note that, for ensemble classifiers, R is to indicate the largest percentage
of feature words used among constituent classifiers, e.g., an ensemble classifier consists of
single classifiers using 1, 4, 9, …, up to R% of word features.</p>
      <p>As in Figure 1, for single classifiers, the BEP measure peaks at R=16 and it degrades
when R &lt; 16 or R &gt; 16. On the other hand, performance of ensemble classifiers keeps
improving as R gets larger. The ensemble classifier with R=100 outperformed all the single
SVM classifiers in terms of both AUC and BEP (Table 2).</p>
      <p>
        To examine the applicability of this ensemble approach to Naïve Bayes methods, we
repeated the same experiment using the MALLET library [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] to build multinomial Naïve
Bayes classifiers. The results are reported in Table 3 and Figure 2. Table 3 shows that
performance (i.e., AUC) of Naïve Bayes classifiers in this study agrees with that in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
(despite that [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] used binary feature vectors and we used TF feature vectors, see, e.g., [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]).
As in Table 3 and Figure 2, although the proposed ensemble approach improved classification
performance of Naïve Bayes classifiers in terms of AUC, it did not improve in terms of BEP.
      </p>
      <p>
        While these results need to be confirmed on other data sets, the success of the
proposed ensemble approach may be attributed to the property of SVM classifiers that they
can exploit a large number of features (even less informative features in terms of IG) with
hardly over-fitting to data sets [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Namely, given a larger number of words as features,
SVM classifiers will yield a globally well-ordered document list without over-fitting. On the
other hand, given a small number of the top IG value words, SVM classifiers will identify
apparently positive and apparently negative documents confidently. Thus, the ensemble
classifiers will take advantage of the both ranking schemes. This did not hold for Naïve Bayes
classifiers, whose performance (BEP) degraded when a large number of features were used.
      </p>
      <sec id="sec-11-1">
        <title>Conclusions</title>
        <p>In this study, we examined a simple and easy-to-deploy classifier ensemble approach for
biomedical document classification/retrieval tasks. In the proposed approach, constituent
classifiers were built by varying the sizes of the feature set for an ML algorithm. Note that
even when a single classifier is employed in a database curation project, a number of
classifiers with different sizes of feature sets would be built anyway before the best
performing system is selected. The proposed approach suggests combining such intermediate
classifiers. In our experiments, SVM ensembles outperformed all the constituent classifiers in
terms of both AUC and BEP. Using this approach, we updated the classification performance
previously reported on the benchmarking data sets, and set new baseline performance for the
data sets. However, the ensemble approach was not effective when there was no sufficient
data to train reliable constituent classifiers or when it was applied to Naïve Bayes classifiers.</p>
        <p>In the current ensemble method, the way we derive constituent classifiers is based on
our observation of the list of feature words. We plan to explore systematic ways in selecting
different sets of features, and different approach to combining resulted classifiers. We also
plan to investigate the effectiveness of the method using different data sets and different ML
algorithms.
final manuscript.</p>
      </sec>
      <sec id="sec-11-2">
        <title>Authors' contributions</title>
        <p>MT carried out the experiments and drafted the manuscript. HL participated in the design of
the study and helped draft and revise the manuscript. MT and HL both read and approved the</p>
      </sec>
      <sec id="sec-11-3">
        <title>Acknowledgements</title>
        <p>We thank Zhangzhi Hu at PIR as well as anonymous reviewers of LBM for helpful comments
on the manuscript. We also thank those researchers who made the corpora and the machine
learning software publicly available. This project was supported by IIS-0639062 from the
National Science Foundation.
Identifiers (PMIDs) assigned to MEDLINE citations. Some of the citations, however, are no
longer accessible with the listed PMIDs. In this table, numbers in parentheses are the
document counts reported in the original papers introducing the data sets. We used two of the
five data sets from Han et al., the acetylation and phosphorylation data sets (indicated by * in
the table), where there are (originally) more than 50 positive documents.</p>
        <sec id="sec-11-3-1">
          <title>Data sets</title>
          <p>Data sets in Han et la.
Acetylation*
Glycosylation
Hydroxylation
Methylation
Phosphorylation*</p>
        </sec>
        <sec id="sec-11-3-2">
          <title>Positives</title>
          <p>Each figure shows how AUC or BEP changes as the setting of R changes for single SVM
classifiers (▲) and SVM ensemble classifiers (■).
0.87
C
U</p>
          <p>A
0.865
0.855
0.82
0.815
0.81
1
The each figure shows how AUC or BEP changes as the setting of R changes for single
Naïve Bayes classifiers (▲) and Naïve Bayes ensemble classifiers (■).</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Farriol-Mathis</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garavelli</surname>
            <given-names>JS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boeckmann</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duvaud</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gasteiger</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gateau</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veuthey</surname>
            <given-names>AL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bairoch</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>Annotation of post-translational modifications in the SwissProt knowledge base</article-title>
          .
          <source>Proteomics</source>
          <year>2004</year>
          ,
          <volume>4</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1537</fpage>
          -
          <lpage>1550</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Peri</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navarro</surname>
            <given-names>JD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kristiansen</surname>
            <given-names>TZ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amanchy</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Surendranath</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muthusamy</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gandhi</surname>
            <given-names>TK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrika</surname>
            <given-names>KN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deshpande</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suresh</surname>
            <given-names>S</given-names>
          </string-name>
          et al:
          <article-title>Human protein reference database as a discovery resource for proteomics</article-title>
          .
          <source>Nucleic Acids Res</source>
          <year>2004</year>
          ,
          <volume>32</volume>
          (Database issue):
          <fpage>D497</fpage>
          -
          <lpage>501</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hersh</surname>
            <given-names>W</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhupatiraju</surname>
            <given-names>RT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corley</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>Enhancing access to the Bibliome: the TREC Genomics Track</article-title>
          .
          <source>Medinfo</source>
          <year>2004</year>
          ,
          <volume>11</volume>
          (Pt 2):
          <fpage>773</fpage>
          -
          <lpage>777</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Sebastiani</surname>
            <given-names>F</given-names>
          </string-name>
          :
          <article-title>Machine learning in automated text categorization</article-title>
          .
          <source>Acm Computing Surveys</source>
          <year>2002</year>
          ,
          <volume>34</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Donaldson</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Bruijn</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wolting</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lay</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tuekam</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baskin</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bader</surname>
            <given-names>GD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michalickova</surname>
            <given-names>K</given-names>
          </string-name>
          et al:
          <article-title>PreBIND and Textomy--mining the biomedical literature for protein-protein interactions using a support vector machine</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2003</year>
          , 4:
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dobrokhotov</surname>
            <given-names>PB</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goutte</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veuthey</surname>
            <given-names>AL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaussier</surname>
            <given-names>E</given-names>
          </string-name>
          :
          <string-name>
            <surname>Combining</surname>
            <given-names>NLP</given-names>
          </string-name>
          <article-title>and probabilistic categorisation for document and term selection for Swiss-Prot medical annotation</article-title>
          .
          <source>Bioinformatics</source>
          <year>2003</year>
          , 19
          <issue>Suppl 1</issue>
          :
          <fpage>i91</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Wang</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morgan</surname>
            <given-names>AA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>Q</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sette</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            <given-names>B</given-names>
          </string-name>
          :
          <article-title>Automating document classification for the Immune Epitope Database</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2007</year>
          ,
          <volume>8</volume>
          (
          <issue>1</issue>
          ):
          <fpage>269</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Han</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Obradovic</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hu</surname>
            <given-names>ZZ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            <given-names>CH</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vucetic</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>Substring selection for biomedical document classification</article-title>
          .
          <source>Bioinformatics</source>
          <year>2006</year>
          ,
          <volume>22</volume>
          (
          <issue>17</issue>
          ):
          <fpage>2136</fpage>
          -
          <lpage>2142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Shah</surname>
            <given-names>PK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bork</surname>
            <given-names>P</given-names>
          </string-name>
          :
          <article-title>LSAT: learning about alternative transcripts in MEDLINE</article-title>
          .
          <source>Bioinformatics</source>
          <year>2006</year>
          ,
          <volume>22</volume>
          (
          <issue>7</issue>
          ):
          <fpage>857</fpage>
          -
          <lpage>865</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Dietterich</surname>
            <given-names>TG</given-names>
          </string-name>
          :
          <article-title>Ensemble methods in machine learning</article-title>
          .
          <source>Multiple Classifier Systems</source>
          <year>2000</year>
          ,
          <year>1857</year>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Chung</surname>
            <given-names>YS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            <given-names>DF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            <given-names>CY</given-names>
          </string-name>
          :
          <article-title>On the Diversity-Performance Relationship for Majority Voting in Classifier Ensembles</article-title>
          .
          <source>In: 7th International Workshop on Multiple Classifier Systems (MCS)</source>
          <year>2007</year>
          ; Springer Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Breiman</surname>
            <given-names>L</given-names>
          </string-name>
          :
          <article-title>Bagging predictors</article-title>
          .
          <source>Machine Learning</source>
          <year>1996</year>
          ,
          <volume>24</volume>
          (
          <issue>2</issue>
          ):
          <fpage>123</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Freund</surname>
            <given-names>Y</given-names>
          </string-name>
          :
          <article-title>Boosting a Weak Learning Algorithm by Majority</article-title>
          .
          <source>Information and Computation</source>
          <year>1995</year>
          ,
          <volume>121</volume>
          (
          <issue>2</issue>
          ):
          <fpage>256</fpage>
          -
          <lpage>285</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Dong</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Han</surname>
            <given-names>K</given-names>
          </string-name>
          :
          <article-title>A Comparison of Several Ensemble Methods for Text Categorization</article-title>
          .
          <source>In: Services Computing</source>
          ,
          <source>2004 IEEE International Conference on (SCC)</source>
          <year>2004</year>
          : pp.
          <fpage>419</fpage>
          -
          <lpage>422</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Kudo</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsumoto</surname>
            <given-names>Y</given-names>
          </string-name>
          :
          <article-title>Chunking with Support Vector Machines</article-title>
          . In:
          <article-title>The Second Meeting of North American Chapter of Association for Computational Linguistics (NAACL)</article-title>
          <year>2001</year>
          : pp.
          <fpage>192</fpage>
          -
          <lpage>199</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Joachims</surname>
            <given-names>T</given-names>
          </string-name>
          :
          <article-title>Making large-Scale SVM Learning Practical</article-title>
          : MIT-Press;
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Harman</surname>
            <given-names>D</given-names>
          </string-name>
          :
          <article-title>How effective is suffixing</article-title>
          ?
          <source>Journal of the American Society for Information Science</source>
          <year>1991</year>
          ,
          <volume>42</volume>
          (
          <issue>1</issue>
          ):
          <fpage>7</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>McCallum</surname>
            <given-names>AK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nigam</surname>
            <given-names>K</given-names>
          </string-name>
          :
          <article-title>A Comparison of Event Models for Naive Bayes Text Classification</article-title>
          . In: AAAI/ICML-98 Workshop on Learning for Text Categorization: AAAI Press;
          <year>1998</year>
          : pp.
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Joachims</surname>
            <given-names>T</given-names>
          </string-name>
          :
          <article-title>Text Categorization with Support Vector Machines: Learning with Many Relevant Features</article-title>
          . University of Dortmund;
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Hand</surname>
            <given-names>DJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Till</surname>
            <given-names>RJ</given-names>
          </string-name>
          :
          <article-title>A Simple Generalisation of the Area Under the ROC Curve for Multiple Class Classification Problems</article-title>
          .
          <source>Machine Learning</source>
          <year>2001</year>
          ,
          <volume>45</volume>
          (
          <issue>2</issue>
          ):
          <fpage>171</fpage>
          -
          <lpage>186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>McCallum</surname>
            <given-names>AK</given-names>
          </string-name>
          :
          <article-title>MALLET: A Machine Learning for Language Toolkit</article-title>
          . http://mallet.cs.umass.edu;
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>4 9 16 25 36 49 64</source>
          81 100
          <string-name>
            <given-names>R</given-names>
            <surname>(</surname>
          </string-name>
          <article-title>"1 up to R" for ensembles)</article-title>
          <source>Single Ensemble 4 9 16 25 36 49 64</source>
          81 100
          <string-name>
            <given-names>R</given-names>
            <surname>(</surname>
          </string-name>
          <article-title>"1 up to R" for ensembles) Single Ensemble</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>