<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AUTH-Atypon at BioASQ 3: Large-Scale Semantic Indexing in Biomedicine</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yannis Papanikolaou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grigorios Tsoumakas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manos Laliotis</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikos Markantonatos</string-name>
          <email>nikos@atypon.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ioannis Vlahavas</string-name>
          <email>vlahavasg@csd.auth.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aristotle University of Thessaloniki</institution>
          ,
          <addr-line>Thessaloniki 54124</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Atypon Hellas</institution>
          ,
          <addr-line>Dimitrakopoulou 7, Agia Paraskevi 15341, Athens</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Atypon</institution>
          ,
          <addr-line>5201 Great America Parkway Suite 510, Santa Clara, CA 95054</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present the methods and the approaches employed in terms of our participation to the BioASQ Challenge 2015 and more speci cally in task 3a, concerning the automatic semantic annotation of scienti c abstracts. Based on the successful approaches of the previous years we considered a variety of ensembles, incorporated journalspeci c semantic information and developed an approach to handle the concept drift within the BioASQ corpus. The o cial results demonstrate a consistent advantage of our approaches against the BioASQ and the National Library of Medicine (NLM) baselines. Speci cally, the systems proposed by our team ranked among the top tier ones along the competition, obtaining the second place in 10 out of 15 weeks.</p>
      </abstract>
      <kwd-group>
        <kwd>semantic indexing</kwd>
        <kwd>BioASQ</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The BioASQ project [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] aims to provide a challenge framework for researchers
dealing with classi cation (semantic indexing) and natural language processing
(question answering) tasks in the eld of bio-medicine. The challenge, similar
to the previous two years, is divided in two tasks: automated semantic indexing
(3a) and question answering (3b). In Task 3a participants are given a set of
new, unannotated articles and are required to automatically predict the relevant
MeSH terms for each one of them in a given time. For each article only the
abstract along with some meta-information is provided (journal, year and title).
This task is particularly di cult, as the MeSH taxonomy comprises of a large
number of labels ( 27000), with the label set following a power-law similar
distribution. Furthermore the terms are subject to a signi cant concept drift
along time.
      </p>
      <p>
        A number of di erent approaches have been pursued along the previous
challenges, in order to automatically annotate new articles. The NLM Medical Text
Indexer (MTI/MTIFL) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], is a system that incorporates multiple rule-based and
machine learning methods in order to e ectively provide MeSH label
recommendations for new articles. Other approaches include Learning-to-Rank methods
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], hierarchical classi cation [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] or multi-label ensemble approaches [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        In this work we build on the previous year's methods [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], employing
ensemble techniques for the semantic indexing task. The rest of the paper is organized
as follows. In Section 2, we present the methods used throughout the
semantic indexing part of the challenge. Section 3 shows the relevant results. Final
considerations and conclusions are drawn in Section 4.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>
        In this section we present the methods that we used for the semantic indexing
task. We rst provide a brief description of those approaches that were used
also in our previous challenge participation [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and then provide the various
extensions of our work with more detail.
      </p>
      <p>In this year's participation, we used as a training set the last 1 million articles
and reserved the last 20 thousand as a validation set. For pre-processing of the
articles, a similar pipeline was used as in the previous years; the abstract and
the title were concatenated, one-grams and bi-grams were used as features and
stop-words as well as features with less than ve occurrences in the corpus were
removed. Following the above steps we obtained 257,197 one-grams and 478,533
bi-grams. The tf-idf representation was used for the features. Also, zoning of the
features belonging to the title and those equal to a MeSH label was performed;
speci cally, we increased the tf-idf value of features that belonged to the title
by log2 and those being equal to a label by log1:25.</p>
      <p>
        The above features were used in order to train several multi-label learning
models. We used the Meta-Labeler [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a set of Binary Relevance (BR) models
with Linear SVMs (both tuned and with default parameters) and a Labeled LDA
variant, Prior LDA [8]. Speci cally for the SVM models, we used di erent values
for the C parameter and handled class imbalance by penalizing more heavily
false negative errors than false positive ones by adjusting properly the weight
parameter [9].
2.1
      </p>
      <sec id="sec-2-1">
        <title>Rule-Based Journal Model</title>
        <p>Along with the previously mentioned models, we developed a rule-based model,
exploiting the journal-speci c distributions of labels. The BioASQ corpus
contains scienti c papers from more than 5000 journals, that cover diverse scienti c
domains and topics and therefore we expect the MeSH terms distributions to
greatly vary among them. Furthermore, articles belonging to a speci c journal
may contain one or more MeSH labels particular to that journal, e.g. we expect
an article belonging to the journal "Pediatrics", to contain the MeSH terms
"Infant" or "Infant, Newborn" with a rather high probability.</p>
        <p>Given the above observations, we rst studied the label distributions among
di erent journals and we observed that speci c labels appear with very high
probabilities ( 0:75) in every journal. Subsequently, we implemented a
rulebased journal model in which labels are divided in two categories, frequent (for
instance with more than 100,000 appearances out of the entire corpus of 4:2
million documents) and non-frequent. Then, each instance, according to the
journal it belongs, is assigned automatically a frequent label if it has a probability
of more than 0.95 in that journal and similarly a non-frequent label if the relevant
probability is greater than 0.75. Naturally, multiple frequent and non-frequent
labels can be assigned to the same instance. The above values were heuristically
chosen, based on small-scale experiments.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Ensembles</title>
        <p>
          The systems used throughout the challenge, were mainly based on ensemble
methods, similar to the previous year participation. We used the MULE
framework [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and further experimented on voting systems. In the following, we
describe the details.
        </p>
        <p>
          MULE MULE [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] is a statistical signi cance multi-label ensemble that performs
classi er selection. The key idea is to combine a set of multi-label classi ers
aiming to optimize a selected measure (for the purpose of this challenge, we are
mainly interested in the micro-F measure) and validate this combination through
a statistical signi cance test; McNemar's test. This way, each label of the
multilabel problem is predicted with a speci c component model, the one that (a)
contributes to the greatest improvement to the evaluation metric of interest and
(b) is validated from the statistical test to indeed produce the aforementioned
improvement. If the null hypothesis of the statistical test is not rejected for
a given label (i.e. if the improvement for a speci c component model is not
statistically signi cant), we predict that label with the globally optimal model.
Voting ensembles We further considered three voting ensembles, which decide
whether to assign a label to an article or not based on the votes of the component
models. The rst voting ensemble relied on the majority vote, while the others
on two and three votes respectively.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Full-Text retrieval</title>
        <p>In the PubMed interface4, for a number of journals, open-access to the full
text of the articles is available, through the PubMed Central (PMC) web page.
The motivation is that the full text of an article will provide more semantic
information and more features in order to learn MeSH terms, especially those
occurring more rarely. After retrieval of a total of 160,691 entire articles (out of
4 http://www.ncbi.nlm.nih.gov/pubmed
the entire corpus which consisted of 4:2 million abstracts) and having trained a
Meta-Labeler model on the new data set, we used the model for prediction of
new articles for which the full-text was also available.</p>
        <p>In order to study the e ect of including the full-text to learn a model, we
considered the following strategies:
{ FF: stands for use of the full text for both training and testing documents.
{ FA: stands for using the full text only for documents in the training set ( for
documents in the test set we use only the abstract).
{ AA: stands for using only the abstract for both training and testing
{ AF: for using full text only for the test set documents</p>
        <p>Table 1 shows the relevant results. We can easily see that including the full
text of an article yields an improvement in Micro-F but not necessarily in the
Macro-F measure. Also, the model does not seem to bene t from the respective
combinations (AF, FA) In short, we would propose using the full-text on a
similar semantic indexing task, only if the full text is available for both the
training and the testing documents. Finally, we note that as the training data
set for the full-text model was a lot smaller that the default abstracts data set
( 1m), the performance for these particular instances was signi cantly worse so
this approach was not further considered during the BioASQ challenge.
The BioASQ corpus extends over a period of almost 70 years (1946-2015) and
thus we expect signi cant changes in the meaning and the context of concepts
(i.e. MeSH terms). For instance, a disease in 1970 and in 2000 can be connected to
totally di erent causes. Furthermore, the MeSH ontology is subject to changes
and additions of new terms every year. The above factors a ect the MeSH
word distributions and consequently a machine learning model performance. In
order to handle this phenomenon, we trained classi ers with variable training
sizes and extending across various time periods (2012-2014, 2010-2014,
20072014) and combined them through the MULE framework (Sect. 2.2) along with
the rest of the models. In this manner, we managed to use the useful semantic
information across large portions of the corpus, at the same time smoothing out
the e ect of the concept drift in the model's performance. A more detailed study
of the temporal aspects of the data along with their e ect on performance can
be found in [10].
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>In this section we present and discuss some key aspects with respect to the o
cial challenge results (http://participants-area.bioasq.org/results/3a/,
concerning our systems.</p>
      <p>We made submissions in 14 out of a total of 15 weeks. During the rst and
the third batch, we steadily obtained the second place for both Micro-F and
LCA-F metrics (9 out of 10 weeks) while in the second batch we ranked in the
third place. Table 2 shows the results in terms of the Micro-F and the LCA-F
measures, for the best performing model of each of the AUTH-Atypon, NLM
and NCBI teams, during the rst batch (the respective results for the two other
batches are available at the BioASQ challenge website). Results are shown for
the already annotated articles, as of May, 11. In total, we outperformed both
NLM's (MTI/MTIFL) and NCBI's (MeSH Now BF/HR) systems in terms of
Micro-F, while in terms of LCA-F, we outperformed the NCBI systems in 4 out
of 5 weeks and NLM systems throughout the batch.</p>
      <p>In order to indicate the improvement of our systems with respect to last
year, in Table 3, we additionally compare the mean results of this year's
challenge (BioASQ 3) to the respective ones from last year (BioASQ 2). Results
are shown for each year's data sets in terms of the mean performance in terms
of Micro-F and LCA-F, for the top performing model across all weeks. We can
observe a signi cant improvement between our two participations, mainly
related to our use of a wider variety of component models, as well as to di erent
parameterizations of the MULE ensembles.
In this paper we presented the participation of the AUTH-Atypon team in the
BioASQ challenge 2015. Building on the successful approaches in the past two
challenges, we further extended our line of work to improve the performance of
our systems, employing ensemble techniques for a number of component models.
The o cial challenge results demonstrate a clear advantage of our methods over
the BioASQ baseline as well as the NLM and NCBI systems.
8. Rubin, T.N., Chambers, A., Smyth, P., Steyvers, M.: Statistical Topic Models for</p>
      <p>Multi-label Document Classi cation. Mach. Learn. 88(1-2) (July 2012) 157{208
9. Lewis, D.D., Yang, Y., Rose, T.G., Li, F.: RCV1: A New Benchmark Collection
for Text Categorization Research. J. Mach. Learn. Res. 5 (2004) 361{397
10. Papanikolaou, Y., Tsoumakas, G., Laliotis, M., Markantonatos, N., Vlahavas, I.:
Large-scale semantic indexing of biomedical papers via a statistical signi cance
multi-label ensemble. Journal of Biomedical Semantics (2015) Manuscript accepted
for publication.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Balikas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Partalas</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krithara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paliouras</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Results of the BioASQ Track of the Question Answering Lab at CLEF 2014</article-title>
          . In: Working Notes for CLEF 2014 Conference,
          <article-title>She eld</article-title>
          ,
          <source>UK, September 15-18</source>
          ,
          <year>2014</year>
          . (july
          <year>2014</year>
          )
          <volume>1181</volume>
          {
          <fpage>1193</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Mork</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Recent enhancements to the NLM medical text indexer</article-title>
          .
          <source>In: Working Notes for CLEF 2014 Conference</source>
          ,
          <article-title>She eld</article-title>
          , UK. (
          <year>2014</year>
          )
          <volume>1328</volume>
          {
          <fpage>1336</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>NCBI at the 2014 BioASQ Challenge Task: Large-scale Biomedical Semantic Indexing and Question Answering</article-title>
          . In: Working Notes for CLEF 2014 Conference,
          <article-title>She eld</article-title>
          ,
          <source>UK, September 15-18</source>
          ,
          <year>2014</year>
          . (
          <year>2014</year>
          )
          <volume>1319</volume>
          {
          <fpage>1327</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The Fudan-UIUC Participation in the BioASQ Challenge Task 2a: The Antinomyra system</article-title>
          .
          <source>In: Working Notes for CLEF 2014 Conference, She eld, UK, September 15-18</source>
          ,
          <year>2014</year>
          . (
          <year>2014</year>
          )
          <volume>1311</volume>
          {
          <fpage>1318</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ribadas-Pena</surname>
            ,
            <given-names>F.J.</given-names>
          </string-name>
          , de Campos Iban~ez,
          <string-name>
            <given-names>L.M.</given-names>
            ,
            <surname>Bilbao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.M.D.</given-names>
            ,
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.E.</surname>
          </string-name>
          :
          <article-title>CoLe and UTAI Participation at the 2014 BioASQ Semantic Indexing Challenge</article-title>
          . In: Working Notes for CLEF 2014 Conference,
          <article-title>She eld</article-title>
          ,
          <source>UK, September 15-18</source>
          ,
          <year>2014</year>
          . (
          <year>2014</year>
          )
          <volume>1361</volume>
          {
          <fpage>1374</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Papanikolaou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dimitriadis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsoumakas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laliotis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markantonatos</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vlahavas</surname>
            ,
            <given-names>I.P.</given-names>
          </string-name>
          :
          <article-title>Ensemble Approaches for Large-Scale Multi-Label Classi cation and Question Answering in Biomedicine</article-title>
          . In: Working Notes for CLEF 2014 Conference,
          <article-title>She eld</article-title>
          ,
          <source>UK, September 15-18</source>
          ,
          <year>2014</year>
          . (
          <year>2014</year>
          )
          <volume>1348</volume>
          {
          <fpage>1360</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rajan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Narayanan</surname>
            ,
            <given-names>V.K.</given-names>
          </string-name>
          :
          <article-title>Large scale multi-label classi cation via metalabeler</article-title>
          .
          <source>In: WWW '09: Proceedings of the 18th international conference on World wide web</source>
          , New York, NY, USA, ACM (
          <year>2009</year>
          )
          <volume>211</volume>
          {
          <fpage>220</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>