<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Fudan participation in the 2015 BioASQ Challenge: Large-scale Biomedical Semantic Indexing and Question Answering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shengwen Peng</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ronghui You</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhikai Xie</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beichen Wang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yanchun Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shanfeng Zhu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Applied Informatics, College of Engineering and Science, Victoria University</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science, Fudan University</institution>
          ,
          <addr-line>Shanghai 200433</addr-line>
          ,
          <country country="CN">P. R. China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Shanghai Key Lab of Data Science, Fudan University</institution>
          ,
          <addr-line>Shanghai 200433</addr-line>
          ,
          <country country="CN">P. R. China</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Shanghai Key Lab of Intelligent Information Processing, Fudan University</institution>
          ,
          <addr-line>Shanghai 200433</addr-line>
          ,
          <country country="CN">P. R. China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article describes the participation of Fudan team in the 2015 BioASQ challenge. The challenge consists of two tasks, largescale biomedical semantic indexing (task 3a) and biomedical question answering (task 3b). In task 3a, our method, MeSHLabeler, achieved the rst place in all 15 weeks of three batches. Based on 3215 annotated citations (June 6, 2015) out of all 4435 citations in the o cial test set (batch 1, week 2), our submission best achieved 0.6194 in at Micro F-measure. This is 0.0576(10.25%) higher than 0.5618, obtained by the o cial NLM solution Medical Text Indexer (MTI). Task 3b includes two phases. Given the questions raised by a team of biomedical experts from around Europe, the main task of phase A is to nd relevant documents, snippets, concepts and RDF triples, while the main task of phase B is to provide exact and ideal answers. In the phase A of task 3b, our submission, fdu, achieved the rst place in both document and snippet retrieval in batch 5 (June 6, 2015).</p>
      </abstract>
      <kwd-group>
        <kwd>MeSH Indexing</kwd>
        <kwd>Learning to Rank</kwd>
        <kwd>Multi-Label Classi cation</kwd>
        <kwd>Biomedical Question Answering</kwd>
        <kwd>Information Retrieval</kwd>
        <kwd>Information Extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        BioASQ 2015 is the third year of BioASQ challenge, which is an established
international competition for large-scale biomedical semantic indexing and question
answering since 2013[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It consists of two tasks, A) automatic indexing new
MEDLINE citations using current Medical Subject Headings (MeSH), and B)
? Corresponding author
answering biomedical questions raised by biomedical expert from around
Europe. Task 3a includes three rounds, with each taking ve weeks. In each week,
the organizers provide thousands of new MEDLINE citations to the competition
participants, who are required to submit MeSH annotations in 21 hours. Our
system, MeSHLabeler, achieved the rst place in all 15 weeks of three rounds.
Task 3b includes 5 rounds. In each round, around 100 biomedical questions are
provided to the competition participants. There are two phases in each round.
In phase A, the competition participants are required to submit relevant
documents, snippets, concepts and RDF triples in 24 hours. The organizers then
release gold (correct) relevant articles and snippets. In phase B, the competition
participants are required to submit exact and ideal answers of these questions.
In the phase A of task 3b, our best submission, fdu, achieved the rst place in
both document and snippet retrieval in round 5, with MAP of 0.2035 and 0.1226,
respectively.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Problem</title>
      <sec id="sec-2-1">
        <title>Task3a: Large-scale MeSH Indexing</title>
        <p>
          MeSH [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is the NLM controlled vocabulary thesaurus used for indexing almost
all of the citations in MEDLINE5. It is widely used for facilitating biomedical
information retrieval and knowledge discovery[
          <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
          ]. MeSH is organized
according to the hierarchical structure, and it is slightly updated every year. In
2015, there are more than 27000 MeSH headings. With the dramatic growth of
biomedical documents, the number of citations in MEDLINE has reached to 23
million 6. To reduce the time and nancial cost, NLM has developed a software
package, MTI (Medical Text Indexer), for assisting MeSH annotation[
          <xref ref-type="bibr" rid="ref7 ref8 ref9">7, 8, 9</xref>
          ].
Automatic MeSH indexing is a very challenging problem. The di culty of this
problem comes from the following three factors: (1) the large number of distinct
MeSH and their uneven distribution in the MEDLINE; (2) the large variation
in the number of MeSH of each citation; and (3) insu cient information, such
as full text.
2.2
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>MeSHLabeler</title>
      <p>
        Each MeSH can be viewed as a label. The MeSH indexing problem is actually
a large-scale multi-label classi cation problem. We have developed a novel
algorithm, MeSHLabeler, for solving this problem, which has also achieved the
rst place in the round 3 of BioASQ 2014 Task2a [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. The basic idea of
MeSHLabeler is to integrate multiple evidence by learning to rank to achieve
accurate MeSH annotation. It consists of two components, MeSHRanker and
MeSHNumber. Given a test citation x, MeSHRanker returns an ordered list L
of candidate MeSH headings, and MeSHNumber predicts the actual number of
      </p>
      <sec id="sec-3-1">
        <title>5 http://www.ncbi.nlm.nih.gov/mesh 6 http://www.nlm.nih.gov/pubs/factsheets/medline.html</title>
        <p>MeSH headings annotated, K. Then top K MeSH headings of L is returned as
the predicted MeSH annotation for x.</p>
        <p>
          Multiple evidence has been integrated into MeSHLabeler by learning to rank.
These evidence can be mainly divided into ve di erent types, local evidence,
global evidence, pattern matching, MTI and MeSH dependency. Local evidence
considers only a small number of citations that are most similar to the test
citation. Global evidence refers to the MeSH classi ers trained from the whole
MEDLINE collection. We train a distinct classi er by logistic regression for
each MeSH heading. A novel score normalization method is developed to make
the prediction scores of di erent classi ers comparable. Pattern matching tries
to scan the title and abstract of test citation to see if it includes any MeSH
headings and its synonyms. MTI is a mixture of local evidence, pattern matching
and indexing rules. Incorporating MTI into MeSHLabeler can take advantage of
domain knowledge in MeSH indexing. In addition, taking MeSH dependency
into consideration is a distinct feature of MeSHLabeler, which has not been
explored in previous studies. Please refer to [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] for the detailed description of
MeSHLabeler.
2.3
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data Processing and Implementation</title>
      <p>
        We downloaded the entire database of MEDLINE in Feb 2015, including 23,343,329
citations. After removing the citations without abstract, title or MeSH
annotations, there are 13,156,128 citations. We extracted the latest 20,000 citations
as the test set and validation set of logistic regression classi ers. In addition,
LTR(Learning to Rank) dataset were extracted from the latest annotated citations
during the competition. The system was mainly written in C++. Referenced
external libraries includes LibLinear 7 for Logistic Regression [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], RankLib 8 for
LambdaMART, JsonCpp 9 for the input/output of json les and OpenMP 10 for
parallel processing. Our server has 4 * Intel XEON E5-4650 2.7GHzs CPU, and
128G memory. After extracting features, both training LTR model and making
prediction are very quick. It takes 5 days to train 27,000 MeSH classi ers by
Logistic Regression, but less than 2 hours to annotate 10,000 citations.
2.4
      </p>
    </sec>
    <sec id="sec-5">
      <title>Performance</title>
      <p>
        The main evaluation metrics for BioASQ are label-based micro F-measure (MiF)
and the Lowest Common Ancestor F-measure (LCA-F)[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. During the whole
competition, MeSHLabeler kept the rst place on both metrics. As shown in
Table 1, we compare the performance of MeSHLabeler with two baselines of
BioASQ 2015, NCBI system(MeSH Now) and the MTI system, in batch 3 and
week 5 of batch 2. In the week 5 of batch 2, MTI is not incorporated into the
      </p>
      <sec id="sec-5-1">
        <title>7 http://www.csie.ntu.edu.tw/ cjlin/liblinear/ 8 http://sourceforge.net/p/lemur/wiki/RankLib/ 9 http://jsoncpp.sourceforge.net/ 10 http://openmp.org/wp/</title>
        <p>MeSHLabeler. MeSHLabeler achieved an MiF of 0.6247, which is 0.0426(7.3%)
higher than 0.5821, obtained by MTIDEF. On the other hand, by incorporating
MTI into MeSHLabeler in batch 3, the performance of MeSHLabeler is
significantly improved, which is on average 8.9% higher than MTI in terms of MiF.
From this we can clearly see the advantage of integrating diverse evidence, such
as MTI, in MeSHLabeler for the accurate MeSH indexing. As shown in Table 2,
we further compare the performance among MeSHLabeler, AUTH and
MeSHUK systems, which are the top 3 systems in the last batch.
There are 4 types of questions in task 3b: yes/no, factoid, list and summary.
In the phase A of task 3b, the participants are required to submit the lists of
relevant documents, concepts, snippets and RDF triples. In each list, at most 10
relevant items can be returned. In the phase B of task 3b, the participants are
required to return the exact and ideal answers.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Task 3b Phase A: Find relevant documents and snippets</title>
      <p>Here we brie y describe several important factors considered in our retrieval
system.</p>
      <p>
        Retrieval Model We chose a statistical language model [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], query likelihood
model, as the underlying model to retrieve relevant documents. The open source
software package Indri 11 was used to build our system, which achieves good
performance in many applications.
      </p>
      <p>Term Weight Optimization As most of nouns in the query are concepts
and phrases, we think these keywords are more important than other keywords.
To emphasize these keywords, we put higher weights on these keywords in the
model.</p>
      <p>Occurrence of Query Keywords We check the occurrence of query keywords
in the retrieved documents. A document that includes all keywords in the query
should be scored higher than other documents with very few query keywords.
Stemming We nd that stemming greatly a ects the performance of the
retrieval system. In most cases, no stemming is a better choice. However, for some
questions, stemming could improve the retrieval performance.</p>
      <p>Phrase If two terms are adjacent in the query, they may be a part of phrase.
In this case, the documents that includes the phrase should be emphasized.
Pseudo Relevance Feedback In our information retrieval, we use the
pseudo relevance feedback to select the top k documents. Some keywords in these
documents are used for query expansion.
3.2</p>
    </sec>
    <sec id="sec-7">
      <title>Task 3b Phase B: Provide exact and ideal answer</title>
      <p>
        Since the golden relevant documents, concepts, snippets and RDF triples are
available for the participants in phase B, we use these to extract exact answers
and ideal answers. After checking previous work of other teams [
        <xref ref-type="bibr" rid="ref14 ref15 ref16">14, 15, 16</xref>
        ], we
design our question answering system. As shown in Fig. 1, our system
architecture of question answering consists of three main components, question analysis,
candidates generating, and candidates ranking.
11 http://www.lemurproject.org/indri.php
Question analysis Question analysis is mainly responsible for extracting
answer types of questions. It is important to understand what the question is
actually asking about. For Factoid and List-type questions, they usually carry key
information of answer type. We classify questions into several types of desired
answers: 1) disease; 2) drug; 3) gene/protein; 4) mutation; 5) number; 6) choice.
Based on the above strategies and historical data of BioASQ 2013 and BioASQ
2014 [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], we develop a set of rules to recognize answer types of given questions.
For example, A Factoid question is as follows: Which gene is associated with
Muenke syndrome? The corresponding rule is \Which gene (.)*".
Candidates generating We could use di erent methods to generate candidates
of di erent questions. For diseases, drugs, gene, protein, mutation, and other
biomedical questions, we can employ PubTator [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] tool to identify and extract
biomedical concepts of corresponding answer type as candidates. For number
and other questions, we use Stanford POS tool 12 to tag relevant snippets, noun,
noun phrases and numbers as candidates.
      </p>
      <p>Candidates ranking For each candidate, we count its word frequency in
relevant documents and snippets, which is the basis of ranking. We then return
the maximum number of allowed answers (e.g., no more than 100 answers for
List-type question).
3.3</p>
    </sec>
    <sec id="sec-8">
      <title>Performance</title>
      <p>In Phase A of task 3b, our best submission achieved the rst place in nding
relevant documents and snippets in batch 5, with a MAP score of 0.2035 and
0.1226, respectively. The performance of fdu and the best MAP of SNUMedinfo
in document level is shown in Table 3. Moreover, Table 4 illustrates the best
performance of top 3 teams in terms of exact answer on task 3b Phase B (July
12 http://nlp.stanford.edu/software/tagger.shtml
6, 2015). According to o cial measures of di erent types of questions, we obtain
the best performance for the Yes/No-type questions in the Batch 5, and for the
List-type questions in Batch 2. Overall our system achieved good performance
in all three types of questions in all ve batches.
Although MeSHLabeler achieved great performance in automatic MeSH
indexing, there are several issues to be explored in the future. Firstly, we found some
citations lack of similar citations from Entrez elink. The data inconsistency can
lead to bad prediction accuracy for these citations. Considering the importance
of KNN score, we may nd similar citations in other ways for these citations.
Secondly, we did not take advantage of the MeSH hierarchical structure, which
could be used to optimize the LCA-F score. Finally, although we attempted to
integrate other information into MeSHLabeler, the performance varies only
slightly. We would like to know the upper limit of MeSHLabeler and if we can
nd an e ective method to integrate other types of information.</p>
      <p>For the Phase A of task 3a, we nd that stemming sometimes improves
the performance. The interesting future work is to automatically judge whether
stemming is a good choice for a speci c question. For the Phase B of task 3b, we
make use of PubTator to identify biomedical concepts. However, some biomedical
concepts cannot be recognized. An accurate biomedical concept identi cation
tool with high coverage would be important for the success of biomedical question
answer system.
5</p>
      <sec id="sec-8-1">
        <title>Acknowledgement</title>
        <p>This work has been partially supported by National Natural Science Foundation
of China (61332013, 61272480 and 61170097), and Scienti c Research Starting
Foundation for Returned Overseas Chinese Scholars, Ministry of Education,
China. We would like to thank Ke Liu, Junning Gao and Renyu Liu for their helpful
suggestions and insightful discussion.</p>
        <p>1
2
3
4
5
fa1(0.8125)
fdu(0.8125)
fa1(0.9655)
fdu(0.9655)
fdu(0.9600)
fdu(0.7143)
fa1(0.6786)
Batch</p>
        <p>Yes/No(Accuracy)</p>
        <p>Factoid(MRR)</p>
        <p>List(F-Measure)
fa1(0.8458)
fdu(0.8458)
main system(0.8458)
main system(0.8125) main system(0.1198)
main system(0.1938) main system(0.1362)
fdu(0.1423)
fa1(0.0769)
fdu (0.0859)
fa1(0.0313)</p>
        <p>fdu(0.0756)
HPI-S2(0.0650)
fdu(0.1160)
main system(0.1081)</p>
        <p>HPI-S2(0.0262)
main system(0.9655) oaqa-3b-3(0.1615)
main system(0.1587)
oaqa-3b-4(0.9600)
oaqa-3b-4(0.5155)
oaqa-3b-4(0.3168)
main system(0.9600)
main system(0.3201)</p>
        <p>fdu(0.2192)
main system(0.1346)
fdu(0.0846)</p>
        <p>fdu(0.1319)
oaqa-3b-3-e(0.0969)
fdu(0.2299)</p>
        <p>main system(0.1349)
oaqa-3b-4(0.2727)</p>
        <p>oaqa-3b-5(0.1875)
fdu(0.2500)</p>
        <p>YodaQA base(0.1631)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Tsatsaronis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balikas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malakasiotis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Partalas</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zschunke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alvers</surname>
            ,
            <given-names>M. R.</given-names>
          </string-name>
          , et al. (
          <year>2015</year>
          ).
          <article-title>An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition</article-title>
          .
          <source>BMC bioinformatics</source>
          ,
          <volume>16</volume>
          (
          <issue>1</issue>
          ),
          <fpage>138</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Nelson</surname>
            ,
            <given-names>S. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schopen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savage</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulman</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Arluk</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          :
          <article-title>The MeSH translation maintenance system: structure, interface design, and implementation</article-title>
          .
          <source>Medinfo</source>
          ,
          <volume>11</volume>
          (
          <issue>Pt 1</issue>
          ),
          <volume>67</volume>
          {
          <fpage>69</fpage>
          . (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Gu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mamitsuka</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>E cient Semi-supervised MEDLINE document clustering with MeSH semantic and global content constraints</article-title>
          .
          <source>IEEE Transactions on Cybernetics</source>
          ,
          <volume>43</volume>
          (
          <issue>4</issue>
          ),
          <volume>1265</volume>
          {
          <fpage>1276</fpage>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mamitsuka</surname>
          </string-name>
          , H.:
          <article-title>Enhancing MEDLINE document clustering by incorporating MeSH semantic similarity</article-title>
          .
          <source>Bioinformatics</source>
          <volume>25</volume>
          (15):
          <year>1944</year>
          {
          <year>1951</year>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takigawa</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mamitsuka</surname>
          </string-name>
          , H.:
          <article-title>Field independent probabilistic model for clustering multi- eld documents</article-title>
          .
          <source>Information Processing &amp; Management</source>
          .
          <volume>45</volume>
          (
          <issue>5</issue>
          ):
          <volume>555</volume>
          {
          <fpage>570</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Enhanced clustering of biomedical documents using ensemble non-negative matrix factorization</article-title>
          .
          <source>Information Science</source>
          <volume>181</volume>
          (
          <issue>11</issue>
          ):
          <volume>2293</volume>
          {
          <fpage>2302</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Mork</surname>
            ,
            <given-names>J. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimeno-Yepes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A. R.:</given-names>
          </string-name>
          <article-title>The NLM Medical Text Indexer System for Indexing Biomedical Literature</article-title>
          . In BioASQ@ CLEF. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mork</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gay</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Humphrey</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>W. Then</given-names>
          </string-name>
          <article-title>NLM indexing initiative's medical text indexer</article-title>
          .
          <source>Stud Health Technol Inform. 107(Pt 1)</source>
          .
          <fpage>268</fpage>
          -
          <lpage>272</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Mork</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Recent</surname>
          </string-name>
          <article-title>Enhancements to the NLM Medical Text Indexer</article-title>
          .
          <source>CLEF (Working Notes)</source>
          <year>2014</year>
          :
          <fpage>1328</fpage>
          -
          <lpage>1336</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The Fudan-UIUC Participation in the BioASQ Challenge Task 2a: The Antinomyra system</article-title>
          .
          <source>CLEF (Working Notes)</source>
          <year>2014</year>
          :
          <fpage>1311</fpage>
          -
          <lpage>1318</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mamitsuka</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>: MeSHLabeler: Improving the Accuracy of Large-scale MeSH indexing by Integrating Diverse Evidence</article-title>
          .
          <source>Bioinformatics</source>
          <volume>31</volume>
          (12):
          <fpage>339</fpage>
          -
          <lpage>347</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>R. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>K. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsieh</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C. J.:</given-names>
          </string-name>
          <article-title>LIBLINEAR: A library for large linear classi cation</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>9</volume>
          :
          <fpage>1871</fpage>
          <string-name>
            <surname>{</surname>
          </string-name>
          <fpage>1874</fpage>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Statistical Language Models for Information Retrieval: A Critical Review at Foundations</article-title>
          and
          <source>Trends in Information Retrieval</source>
          <volume>2</volume>
          (
          <issue>3</issue>
          ):
          <fpage>137</fpage>
          -
          <lpage>213</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Weissenborn</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsatsaronis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schroeder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Answering factoid questions in the biomedical domain</article-title>
          . In 1st BioASQ Workshop:
          <article-title>A challenge on largescale biomedical semantic indexing and question answering</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>NCBI at the 2014 BioASQ challenge task: large-scale biomedical semantic indexing and question answering</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Classi cation and Retrieval of Biomedical Literatures: SNUMedinfo at CLEF QA track BioASQ2014</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Balikas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Partalas</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A. C. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krithara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaussier</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Paliouras</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Results of the BioASQ Track of the Question Answering Lab at CLEF 2014</article-title>
          . CLEF.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>C.-H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kao</surname>
          </string-name>
          , H.-Y.,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>PubTator: a web-based text mining tool for assisting biocuration</article-title>
          .
          <source>Nucleic acids research</source>
          41,
          <fpage>W518</fpage>
          -
          <lpage>W522</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Kosmopoulos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Partalas</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaussier</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paliouras</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Androutsopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Evaluation measures for hierarchical classi cation: a uni ed view and novel approaches</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          ,
          <volume>29</volume>
          (
          <issue>3</issue>
          ),
          <fpage>820</fpage>
          -
          <lpage>865</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>