<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Biomedical question-focused multi-document summarization: ILSP and AUEB at BioASQ3</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Prodromos Malakasiotis</string-name>
          <email>malakasiotis@ilsp.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emmanouil Archontakis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ion Androutsopoulos</string-name>
          <email>ion@aueb.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitrios Galanis</string-name>
          <email>galanisd@ilsp.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harris Papageorgiou</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Informatics, Athens University of Economics and Business</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Language and Speech Processing, Research Center `Athena'</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Question answering systems aim to nd answers to natural language questions by searching in document collections (e.g., repositories of scienti c articles or the entire Web) and/or structured data (e.g., databases, ontologies). Strictly speaking, the answer to a question might sometimes be simply `yes' or `no', a named entity, or a set of named entities. In practice, however, a more elaborate answer is often also needed, ideally a summary of the most important information from relevant documents and structured data. In this paper, we focus on generating summaries from documents that are known to be relevant to particular questions. We describe the joint participation of AUEB and ILSP in the corresponding subtask of the bioasq3 competition, where participants produce multi-document summaries of given biomedical articles that are relevant to English questions prepared by biomedical experts.</p>
      </abstract>
      <kwd-group>
        <kwd>biomedical question answering</kwd>
        <kwd>text summarization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Biomedical experts are extremely short of time. They also need to keep up with
scienti c developments happening at a pace that is probably faster than in any
other science. The online biomedical bibliographic database PubMed currently
comprises approximately 21 million references and was growing at a rate
often exceeding 20,000 articles per week in 2011.3 Figure 1 shows the number
of biomedical articles indexed by PubMed per year since 1964. Rich sources
of structured biomedical information, like the Gene Ontology, umls, or
Diseasesome are also available.4 Obtaining su cient and concise answers from this
wealth of information is a challenging task for traditional search engines, which
instead of answers return lists of (possibly) relevant documents that the experts
3 Consult http://www.ncbi.nlm.nih.gov/pubmed/.
4 See http://www.geneontology.org/, http://www.nlm.nih.gov/research/umls/,
http://diseasome.eu/.
themselves have to study. Consequently, there is growing interest for
biomedical question answering (QA) systems [3, 4], which aim to produce more concise
answers. To foster research in biomedical QA, the bioasq project constructs
benchmark datasets, evaluation services, and organizes international biomedical
QA competitions since 2012 [20].5
900
800
700
)
s
nd600
a
s
huo500
t
(
 
lse 400
c
i
t
r
 A300
d
e
lish200
b
u
P 100</p>
      <p>Given a question expressed in natural language, QA systems aim to provide
answers by searching in document collections (e.g., repositories of scienti c
articles or the entire Web) and/or structured data (e.g., databases, ontologies).
Strictly speaking, the answer to a question might sometimes be simply a `yes' or
`no' (e.g., in biomedical questions like \Do CpG islands co-localize with
transcription start sites?"), a named entity (e.g., in \What is the methyl donor of DNA
(cytosine-5)-methyltransferases?"), or a set of named entities (e.g., in \Which
species may be used for the biotechnological production of itaconic acid?").
Following the terminology of bioasq, we call short answers of this kind `exact'
answers. In practice, however, a more elaborate answer is often needed, ideally
a paragraph summarizing the most important information from relevant
documents and structured data; bioasq calls answers of this kind `ideal' answers. In
this paper, we focus on generating `ideal' answers (summaries) from documents
that are known to be relevant to particular questions. We describe our
participation in the corresponding subtask of the bioasq3 competition (Task 3b, Phase
B, generation of `ideal' answers), where the participants produce summaries of
5 See also http://www.bioasq.org/.
biomedical articles that are relevant to English questions prepared by
biomedical experts. In this particular subtask, the input is a question along with the
PubMed articles that a biomedical expert identi ed as relevant to the
question; in e ect, a perfect search engine is assumed (see Fig. 2). More precisely, in
bioasq3 only the abstracts of the articles were available; hence, we summarize
sets of abstracts (one set per question). We also note that the abstracts contain
annotations showing the snippets (one or more consecutive sentences each) that
the biomedical experts considered most relevant to the corresponding questions.
We do not use the snippet annotations of the experts, since our system includes
its own mechanisms to assess the importance of each sentence. Hence, our system
may be at a disadvantage compared to systems that use the snippet annotations
of the experts. Nevertheless, experimental results we present indicate that it still
performs better than its competitors.</p>
      <p>Question: Do CpG islands co-localize with transcription start sites?</p>
    </sec>
    <sec id="sec-2">
      <title>Query: e.g., “CpG islands” AND “transcription start sites”</title>
    </sec>
    <sec id="sec-3">
      <title>Search Engine</title>
      <p>Documents,
RDF triples …</p>
      <p>QA,
summarization,</p>
      <p>NLG
“Exact” answer: Yes.
“Ideal” answer (summary): Yes. It is generally
known that the presence of a CpG island around
the TSS is related to the expression pattern of the
gene. CGIs (CpG islands) often extend into
downstream transcript regions. This provides an
explanation for the observation that the exon at
the 5' end of the transcript, flanked with the
transcription start site, shows a remarkably higher</p>
      <p>CpG density than the downstream exons.</p>
      <p>We also note that when relevant structured information is also available (e.g.,
rdf triples), concept to text natural language generation (nlg) [1] can also be
used to produce `ideal' answers or texts to be given as additional input
documents to the summarizer. We did not consider nlg, however, since in bioasq3
the questions were not accompanied by manually selected (by the biomedical
experts) relevant structured information, unlike bioasq1 and bioasq2, and we
do not yet have mechanisms to select structured information automatically.</p>
      <p>Section 2 below describes the di erent versions of the multi-document
summarizer that we used. Section 3 reports our experimental results. Section 4
concludes and provides directions for future work.</p>
      <p>Our question-focused multi-document summarizer
We now discuss how the `ideal' answers (summaries) of our system are produced.
Recall that for each question, a set of documents (article abstracts) known to be
relevant to the question is given. Our system is an extractive summarizer, i.e., it
includes in each summary sentences of the input documents, without rephrasing
them. The summarizer attempts to select the most relevant (to the question)
sentences, also trying to avoid including in the summary redundant sentences,
i.e., pairs of sentences that convey the same information. bioasq restricts the
maximum size of each `ideal' answer to 200 words; including redundant sentences
wastes space and is also penalized when experts manually assess the responses
of the systems [20]. The summarizer does not attempt to repair (e.g., replace
pronouns by their referents), order, or aggregate the selected sentences [6]; we
leave these important issues for future work.
2.1</p>
      <p>Baseline 1 and Baseline 2
As a starting point, we used the extractive summarizer of Galanis et al. [7, 8].
Two versions of the summarizer, known as Baseline 1 and Baseline 2, have been
used as baselines for `ideal' answers in all three years of the bioasq competition.6
Both versions employ a Support Vector Regression (svr) model [5] to assign a
relevance score rel (si) to each sentence si of the relevant documents of a question
q.7 An svr learns a function f : Rn ! R in order to predict a real value yi 2 R
given a feature vector ~xi 2 Rn that represents an instance. In our case, ~xi
is a feature vector representing a sentence si of the relevant documents of a
question q, and yi is the relevance score of si. Consult Galanis et al. [7, 8] for a
discussion of the features that were used in the svr of Baseline 1 and Baseline
2. During training, for each q we compute the rouge-2 and rouge-su4 scores
[13] between each si and the gold (provided by an expert) `ideal' answer of
q, and we take yi to be the average of the rouge-2 and rouge-su4 scores.
The motivation for using these scores is that they are the two most commonly
used measures for automatic evaluation of machine-generated summaries against
gold ones. Roughly speaking, both measures compute the word bigram recall
of the summary (or sentence) being evaluated against, possibly multiple, gold
summaries. However, rouge-su4 also considers skip bigrams (pairs of words
with other ignored intervening words) with a maximum distance of 4 words
between the words of each skip bigram. Both measures have been found to
correlate well with human judgements in extractive summarization [13] and,
hence, training a component (e.g., an svr) to predict the rouge score of each
sentence can be particularly useful. Intuitively, a sentence with a high rouge
score has a high overlap with the gold summaries; and since the gold summaries
6 Baseline 1 and Baseline 2 are the ilp2 and greedy-red methods, respectively, of</p>
      <p>Galanis et al. [8]. Baseline 2 had also participated in TAC 2008 [9].
7 We use the svr implementation of libsvm (see http://www.csie.ntu.edu.tw/
~cjlin/libsvm/) with an rbf kernel and libsvm's parameter tuning facilities.
contain the sentences that human authors considered most important, a sentence
with a high rouge score is most likely also important.</p>
      <p>Baseline 1 uses Integer Linear Programming (ilp) to jointly maximize the
relevance and diversity (non-redundancy) of the selected sentences si,
respecting at the same time the maximum allowed summary length. The ilp model
maximizes the following objective function:8
subject to:
max
b;x
n
X
i=1
i</p>
      <p>
        li
lmax
xi + (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) XjBj bi
      </p>
      <p>n
i=1
X bj
gj2Bi</p>
      <p>X xi
si2Sj
n
X lixi
i=1</p>
      <p>
        lmax
jBij xi; for i = 1; : : : ; n
bj ; for j = 1; : : : ; jBj
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
where i is the relevance score rel (si) of sentence si normalized in [0; 1]; li is the
word length of si; lmax is the maximum allowed summary length in words; n is the
number of input sentences (sentences in the given relevant documents); B is the
set of all the word bigrams in the input sentences; xi and bi show which sentences
si and word bigrams, respectively, are present in the summary; Bi is the set of
word bigrams that occur in sentence si; gj ranges over the word bigrams in Bi;
and Sj is the set of sentences that contain bigram gj . Constraint (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) ensures that
the maximum allowed summary length is not exceeded. Constraint (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) ensures
that if an input sentence is included in the summary, then all of its word bigrams
are also included. Constraint (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) ensures that if a word bigram is included in the
summary, than at least a sentence that contains it is also included. The rst sum
of Eq. 1 maximizes the total relevance of the selected sentences. The second sum
maximizes the number of distinct bigrams in the summary, in e ect minimizing
the redundancy of the included sentences. Finally, 2 [0; 1] controls how much
the model tries to maximize the total relevance of the selected sentences at the
expense of non-redundancy and vice versa. Consult Galanis et al. [7, 8] for a
more detailed explanation of the ilp model.
      </p>
      <p>Baseline 2 rst uses the trained svr to rank the sentences si of the relevant
documents of q by decreasing relevance rel (si). It then greedily examines each si
from highest to lowest rel (si). If the cosine similarity between si and any of the
sentences that have already been added to the summary exceeds a threshold t,
then si is discarded; the cosine similarity is computed by representing each
sentence as a bag of words (using Boolean features), and t is tuned on development
8 We use the implementation of the Branch and Cut algorithm of the gnu Linear</p>
      <p>Programming Kit (glpk); consult http://sourceforge.net/projects/winglpk/.
data. Otherwise, if si ts in the remaining available summary space, it is added
to the summary; if it does not t, the summary construction process stops.</p>
      <p>Baselines 1 and 2 were trained on news articles, as discussed by Galanis et
al. [7, 8], and were used in bioasq without retraining and without modifying
the features of their svr. However, there are many di erences between news
and biomedical articles, and many of the features that were used in the svr of
Baselines 1 and 2 are irrelevant to biomedical articles. For example, Baselines
1 and 2 use a feature that counts the names of organizations, persons, etc. in
sentence si, as identi ed by a named entity recognizer that does not support
biomedical entity types (e.g., names of genes, diseases). They also use a feature
that considers the order of si in the document it was extracted from, based
on the intuition that news articles usually list the most important information
rst, a convention that does not always hold in biomedical abstracts. Hence, we
also experimented with modi ed versions of Baselines 1 and 2, discussed below,
which were trained on bioasq datasets and used di erent feature sets.
2.2</p>
      <p>The ILP-SUM-0 and ILP-SUM-1 summarizers
The rst new version of our summarizer, called ilp-sum-0, is the same as
Baseline 1 (the baseline that uses ilp, with the same features in its svr), but was
trained on bioasq data, as discussed in Section 3 below.</p>
      <p>Another version, ilp-sum-1, is the same as ilp-sum-0, it was also trained on
bioasq data, but uses a di erent feature set in its svr, still close to the features
of Baselines 1 and 2 [7, 8], but modi ed for biomedical questions and articles.
The features of ilp-sum-1 are the following. All the features of all the versions
of the summarizer, including Baselines 1 and 2, are normalized in [0; 1].
(1.1) Word overlap: The number of common words between the question q
and each sentence si of the relevant documents of q, after removing stop
words and duplicate words from q and si.
(1.2) Stemmed word overlap: The same as Feature (1.1), but the words of q
and si are stemmed, after removing stop words.
(1.3) Levenshtein distance: The Levenshtein distance [11] between q and si,
taking insertions, deletions, and replacements to operate on entire words.
(1.4) Stemmed Levenshtein distance: The same as Feature (1.3), but the
words of q and si are stemmed, before computing the Levensthein distance.
(1.5) Content word frequency: The average frequency CF (si) of the content
words of sentence si in the relevant documents of q, as de ned by Schilder
and Ravikumar [18]:</p>
      <p>CF (si) =</p>
      <p>Pjc(=s1i) pc(wj )</p>
      <p>c(si)
where c(si) is the number of content words in sentence si, pc(w) = Mm , m
is the number of occurrences of content word wj in the relevant documents
of q, and M is the total number of content word occurences in the relevant
documents of q.
(1.6) Stemmed content word frequency: The same as Feature (1.5), but
the content words of the relevant documents of q (and their sentences si)
are stemmed before computing CF (si).
(1.7) Document frequency: The average document frequency of the content
words of sentence si in the relevant documents of q, as de ned by Schilder
and Ravikumar [18]:</p>
      <p>DF (si) =</p>
      <p>Pjc(=s1i) pd(wj)</p>
      <p>c(si)
where pd(w) = Dd , d is the number of relevant documents of q that contain
the content word wj, and D is the number of relevant documents of q.
(1.8) Stemmed document frequency: The same as Feature (1.7), but the
content words of the relevant documents of q (and their sentences si) are
stemmed before computing DF (si).
2.3</p>
      <p>The ILP-SUM-2 and GR-SUM-2 summarizers
In recent years, continuous space vector representations of words, also known as
word embeddings, have been found to capture several morphosyntactic and
semantic properties of words [12, 14{17]. bioasq employed the popular word2vec
tool [14{16] to construct embeddings for a vocabulary of 1,701,632 words
occurring in biomedical texts, using a corpus of 10,876,004 English abstracts of
biomedical articles from PubMed.9 The ilp-sum-2 and gr-sum-2 versions of
our summarizer use the following features in their svr, which are based on the
bioasq word embeddings, in addition to Features (1.1){(1.8) of ilp-sum-1.
ilpsum-2 also uses the ilp model (like Baseline 1, ilp-sum-0, ilp-sum-1), whereas
gr-sum-2 uses the greedy approach of Baseline 2 instead (see Section 2.1).
(2.1) Euclidean similarity of centroids: This is computed as:
ES (q; si) =</p>
      <p>1
1 + ED(~q; ~si)
where ~q, ~si are the centroid vectors of q and si, respectively, de ned below,
and ED(~q; ~si) is the Euclidean distance between ~q and ~si. The centroid ~t
of a text t (question or sentence) is computed as:</p>
      <p>jV j
t P w~j TF(wj; t)
~t = 1 Xjj w~i = j=1
jtj i=1 jPVj TF(wj; t)</p>
      <p>
        j=1
where jtj is the number of words (tokens) in t, and w~i is the embedding
(vector) of the i-th word (token) of t, jV j is the number of (distinct) words
9 See https://code.google.com/p/word2vec/ and http://bioasq.lip6.fr/tools/
BioASQword2vec/ for further details.
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
in the vocabulary, and TF(wj; t) is the term frequency (number of
occurrences) of the j-th vocabulary word in the text t.10
(2.2) Euclidean similarity of IDF-weighted centroids: The same as
Feature (2.1), except that the centroid of a text t (question or sentence) now
also takes into account the inverse document frequencies of the words in t:
jV j
      </p>
      <p>P w~j TF(wj; t) IDF(wj)
~t = j=1
jV j
P TF(wj; t) IDF(wj)
j=1
(7)
where IDF(wj) = log jDj(Dwjj)j , jDj = 10; 876; 004 is the total number of
abstracts in the corpus the word embeddings were obtained from, and
jD(wj)j is the number of those abstracts that contain the word wj.
(2.3) Pairwise Euclidean similarities: To compute this set of features (8
features in total), we create two bags, one with the tokens (word occurrences)
of the question q and one with the tokens of the sentence si. We then
compute the similarity ES (w; w0) (as in Eq. 5) for every pair of tokens w; w0
of q and si, respectively, and we construct the following features:
- the average of the similarities ES (w; w0), for all the token pairs w; w0
of q and si, respectively,
- the median of the similarities ES (w; w0),
- the maximum similarity ES (w; w0),
- the average of the two largest similarities ES (w; w0),
- the average of the three largest similarities ES (w; w0),
- the minimum similarity ES (w; w0),
- the average of the two smallest similarities ES (w; w0),
- the average of the three smallest similarities ES (w; w0).
(2.4) IDF-weighted pairwise Euclidean similarities: The same set of
features (8 features) as Features (2.3), but the Euclidean similarity ES (w; w0)
of each pair of tokens w; w0 is multiplied with IDF(w) IDF(w0) to reward pairs
maxidf2
with high idf scores. The idf scores are computed as in Feature (2.2), and
maxidf is the maximum idf score of the words we have embeddings for.
3</p>
      <sec id="sec-3-1">
        <title>Experimental results</title>
        <p>We used the datasets of bioasq1 and bioasq2 to train and tune the four new
versions of our summarizer (ilp-sum-0, ilp-sum-1, ilp-sum-2, gr-sum-2). We
then used the dataset of bioasq3 to test the two best new versions of our
summarizer (ilp-sum-2, gr-sum-2) on unseen data, and to compare them against
Baseline 1, Baseline 2, and the other systems that participated in bioasq3.
10 Tokens for which we have no embeddings are ignored when computing the features
of this section.
3.1</p>
        <p>Experiments on BioASQ1 and BioASQ2 data
The bioasq1 and bioasq2 datasets consist of 3 and 5 batches, respectively,
called Batches 1{3 and Batches 4{8 in this section. Each batch contains
approximately 100 questions, along with relevant documents, and `ideal' answers
provided by the biomedical experts.</p>
        <p>In a rst experiment, we aimed to tune the parameter of ilp-sum-0,
ilp-sum-1, and ilp-sum-2, which use the ilp model of Section 2.1, and
compare the three systems. Figure 3 shows the average rouge scores of the three
systems on Batches 4{6, for di erent values of , using Batches 1{3 to train
them (train their svrs); Batches 7{8 were reserved for another experiment,
discussed below. In more detail, we rst computed the rouge-2 and
rougesu4 scores on Batch 4, training the systems on Batches 1{3 and 5{6. We then
computed the average of the rouge-2 and rouge-su4 scores of Batch 4, i.e.,
rouge(Batch4) = 12 (rouge-2(Batch4)+rouge-su4(Batch4)), for each value.
We repeated the same process for Batches 5 and 6, obtaining rouge(Batch5)
and rouge(Batch6), for each value. Finally, we computed (and show in Fig. 3)
the average 31 (rouge(Batch4) + rouge(Batch5) + rouge(Batch6)), for each
valuer. Figure 3 shows that ilp-sum-2 performs better than ilp-sum-1, which in
turn outperforms ilp-sum-0. The di erences in the rouge scores are larger for
greater values of , because greater values place more emphasis on the rel (si)
scores returned by the svr, which are a ected by the di erent feature sets of
the three systems. For &gt; 0:8, the rouge scores decline, because the systems
place too much emphasis on avoiding redundant sentences. The best of the three
systems, ilp-sum-2, achieves its best performance for = 0:8.</p>
        <p>In a second experiment, we compared ilp-sum-2, which is the best of our new
versions that use the ilp model, against gr-sum-2, which uses the same features,
but the greedy approach instead of the ilp model. We set = 0:8 in ilp-sum-2,
based on Fig. 3. In gr-sum-2, we set the cosine similarity threshold (Section 2.1)
to t = 0:4, based on Galanis et al. [7, 8]. Figure 4 shows the average rouge-2 and
rouge-su4 score of each system on Batches 7 and 8, using an increasingly larger
training dataset, consisting of Batches 1{3, 1{4, 1{5, or 1{6. A rst observation
is that ilp-sum-2 outperforms gr-sum-2. Moreover, it seems that both systems
would bene t from more training data.
3.2</p>
        <p>Experiments on BioASQ3 data
In bioasq3, we participated with ilp-sum-2 (with = 0:8) and gr-sum-2 (with
t = 0:4), both trained on all 8 batches of bioasq1 and bioasq2. Baseline 1 and
Baseline 2, which are also versions of our own summarizer, were used again as
the o cial baselines for `ideal' answers, as in bioasq1 and bioasq2, i.e., without
modifying their features or retraining them for biomedical data. The test dataset
of bioasq3 contained ve new batches, hereafter called bioasq3 Batches 1{5;
these are di erent from Batches 1{8 of bioasq1 and bioasq2.</p>
        <p>For each bioasq3 batch, Table 1 shows the rouge-2, rouge-su4, and
average of rouge-2 and rouge-su4 scores of the four versions of our summarizer
0.38
0.37
0.36
0.35
0.34
0.33
0.32
0.43
0.41
0.39
0.37
0.35
0.33
0.31
0.29
0.27
0.25</p>
        <p>Average ROUGE scores </p>
        <p>Average ROUGE scores 
(ilp-sum-2, gr-sum-2, Baseline 1, Baseline 2), ordered by decreasing average
rouge-2 and rouge-su4. The results of the three other best (in terms of average
rouge-2 and rouge-su4) participants per batch are also shown, as
part-sys1, part-sys-2, part-sys-3; part-sys-1 is not necessarily the same system in
all batches, and similarly for part-sys-2 and part-sys-3.11 The four versions
of our summarizer are the best four systems in all ve batches of Table 1.</p>
        <p>As in the experiments of Section 3.1, Table 1 shows that ilp-sum-2
consistently outperforms gr-sum-2. Similarly, Baseline 2 (which uses the greedy
approach) performs better than Baseline 1 (which uses the ilp model) only in
the third batch. It is also surprising that ilp-sum-2 and gr-sum-2 do not always
perform better than Baselines 1 and 2, even though the former systems were
tailored for biomedical data by modifying their features and retraining them on the
datasets of bioasq1 and bioasq2. This may be due to the fact that Baseline 1
and Baseline 2 were trained on larger datasets than ilp-sum-2 and gr-sum-2
[7, 8]. Hence, training our summarizer on more data, even from another domain
(news) may be more important than training it on data from the application
domain (biomedical data, in the case of bioasq) and modifying its features.</p>
        <p>It would be interesting to check if the conclusions of Table 1 continue to hold
when the systems are ranked by the manual (provided by biomedical experts)
evaluation scores of their `ideal' summaries, as opposed to using rouge scores.
At the time this paper was written, the manual evaluation scores of the `ideal'
answers of bioasq3 had not been announced.
4</p>
      </sec>
      <sec id="sec-3-2">
        <title>Conclusions and future work</title>
        <p>We presented four new versions (ilp-sum-0, ilp-sum-1, ilp-sum-2, gr-sum-2)
of an extractive question-focused multi-document summarizer that we used to
construct `ideal' answers (summaries) in bioasq3. The summarizer employs an
svr to assign relevance scores to the sentences of the given relevant abstracts,
and an ilp model or an alternative greedy strategy to select the most
relevant sentences avoiding redundant ones. The two o cial bioasq baselines for
`ideal' answers, Baseline 1 and Baseline 2, are also versions of the same
summarizer; they use the ilp model or the greedy approach, respectively, but they
were trained on news articles and their features are not always appropriate for
biomedical data. By contrast the four new versions were trained on data from
bioasq1 and bioasq2. ilp-sum-0, ilp-sum-1, and ilp-sum-2 all use the ilp
model, but ilp-sum-0 uses the original features of Baselines 1 and 2, ilp-sum-1
uses a slightly modi ed feature set, and ilp-sum-2 uses a more extensive feature
set that includes features based on biomedical word embeddings. gr-sum-2 uses
the same features as ilp-sum-2, but with the greedy mechanism.</p>
        <p>A preliminary set of experiments on bioasq1 and bioasq2 data indicated
that ilp-sum-2 performs better than ilp-sum-0 and ilp-sum-1, showing the
importance of modifying the feature set. ilp-sum-2 was also found to perform
11 The results of all the systems can be found at http://participants-area.bioasq.
org/results/3b/phaseB/.
better than gr-sum-2, which uses the same feature set, showing the bene t of
using the ilp model instead of the greedy approach. Our experiments also
indicated that ilp-sum-2 and gr-sum-2 would probably bene t from more training
data. In bioasq3, we participated with ilp-sum-2 and gr-sum-2, tuned and
trained on bioasq1 and bioasq2 data. Along with Baselines 1 and 2, which are
also versions of our own summarizer, ilp-sum-2 and gr-sum-2 were the best
four systems in terms of rouge scores in all ve batches of bioasq3. Again,
ilp-sum-2 consistently outperformed gr-sum-2, but surprisingly ilp-sum-2 and
gr-sum-2 did not always perform better than Baselines 1 and 2. This may be
due to the fact that Baselines 1 and 2 were trained on more data, suggesting that
the size of the training set may be more important than improving the feature
set or using data from the biomedical domain.</p>
        <p>Future work could consider repairing, ordering, or aggregating the sentences
of the `ideal' answers, as already noted. The centroid vectors of ilp-sum-2 and
gr-sum-2 could also be replaced by paragraph vectors [10] or vectors obtained
by using recursive neural networks [19]. Another possible improvement could
be to use metamap [2], a tool that maps biomedical texts to concepts derived
from umls.12 We could then compute new features that measure the similarity
between a question and a sentence in terms of biomedical concepts.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Acknowledgements</title>
        <p>The work of the rst author was funded by the Athens University of Economics
and Business Research Support Program 2014-2015, \Action 2: Support to
Postdoctoral Researchers".</p>
        <p>7. Galanis, D.: Automatic generation of natural language summaries. Ph.D. thesis,</p>
        <p>Department of Informatics, Athens University of Economics and Business (2012)
8. Galanis, D., Lampouras, G., Androutsopoulos, I.: Extractive multi-document
summarization with integer linear programming and support vector regression. In:
Proceedings of COLING 2012. pp. 911{926. Mumbai, India (2012)
9. Galanis, D., Malakasiotis, P.: AUEB at tac 2008. In: Proceedings of the Text
Analysis Conference. pp. 42{47. Gaithersburg, MD (2008)
10. Le, Q., Mikolov, T.: Distributed representations of sentences and documents. In:
Proceedings of the 31st International Conference on Machine Learning. pp. 1188{
1196. Beijing, China (2014)
11. Levenshtein, V.: Binary codes capable of correcting deletions, insertions, and
reversals. Soviet Physice-Doklady 10, 707{710 (1966)
12. Levy, O., Goldberg, Y., Dagan, I.: Improving distributional similarity with lessons
learned from word embeddings. Transactions of the Association for Computational
Linguistics 3, 211{225 (2015)
13. Lin, C.Y.: ROUGE: A package for automatic evaluation of summaries. In:
Proceedings of the ACL workshop `Text Summarization Branches Out'. pp. 74{81.</p>
        <p>Barcelona, Spain (2004)
14. Mikolov, T., Chen, K., Corrado, G., Dean, J.: E cient estimation of word
representations in vector space. In: Proceedings of Workshop at International Conference
on Learning Representations. Scottsdale, AZ, USA (2013)
15. Mikolov, T., Yih, W., Zweig, G.: Distributed representations of words and phrases
and their compositionality. In: Proceedings of the Conference on Neural
Information Processing Systems. Lake Tahoe, NV (2013)
16. Mikolov, T., Yih, W., Zweig, G.: Linguistic regularities in continuous space word
representations. In: Proceedings of the Conference of the North American Chapter
of the Association for Computational Linguistics - Human Language Technologies.</p>
        <p>Atlanta, GA (2013)
17. Pennington, J., Socher, R.and Manning, C.D.: GloVe: Global vectors for word
representation. In: Proceedings of the Conference on Empirical Methods on Natural
Language Processing. Doha, Qatar (2014)
18. Schilder, F., Kondadadi, R.: Fastsum: Fast and accurate query-based
multidocument summarization. In: Proceedings of 46th Annual Meeting of the
Association for Computational Linguistics - Human Language Technologies, Short Papers.
pp. 205{208. Columbus, Ohio (2008)
19. Socher, R., Huval, B., Manning, C.D., Ng, A.Y.: Semantic compositionality
through recursive matrix-vector spaces. In: Proceedings of the 2012 Joint
Conference on Empirical Methods in Natural Language Processing and Computational
Natural Language Learning. pp. 1201{1211. Jeju Island, Korea (2012)
20. Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers,
M., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., Almirantis,
Y., Pavlopoulos, J., Baskiotis, N., Gallinari, P., Artieres, T., Ngonga, A., Heino, N.,
Gaussier, E., Barrio-Alvers, L., Schroeder, M., Androutsopoulos, I., Paliouras, G.:
An overview of the BioASQ large-scale biomedical semantic indexing and question
answering competition. BMC Bioinformatics 16(138) (2015)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Androutsopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lampouras</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galanis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Generating natural language descriptions from OWL ontologies: the NaturalOWL system</article-title>
          .
          <source>Journal of Arti cial Intelligence Research</source>
          <volume>48</volume>
          ,
          <volume>671</volume>
          {
          <fpage>715</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.:</given-names>
          </string-name>
          <article-title>E ective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program</article-title>
          .
          <source>In: Proceedings of the American Medical Informatics Association Symposium</source>
          . pp.
          <volume>18</volume>
          {
          <fpage>20</fpage>
          .
          <string-name>
            <surname>Washington</surname>
            <given-names>DC</given-names>
          </string-name>
          , USA (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Athenikos</surname>
          </string-name>
          , S., Han, H.:
          <article-title>Biomedical question answering: A survey</article-title>
          .
          <source>Computer Methods and Programs in Biomedicine 99(1)</source>
          ,
          <volume>1</volume>
          {
          <fpage>24</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bauer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berleant</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Usability survey of biomedical question answering systems</article-title>
          .
          <source>Human Genomics</source>
          <volume>6</volume>
          (
          <issue>1</issue>
          )(17) (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Drucker</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burges</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaufman</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>Smola</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vapnik</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , et al.:
          <article-title>Support vector regression machines</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          <volume>9</volume>
          ,
          <fpage>155</fpage>
          {
          <fpage>161</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Filippova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strube</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Sentence fusion via dependency graph compression</article-title>
          .
          <source>In: Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>177</volume>
          {
          <fpage>185</fpage>
          .
          <string-name>
            <surname>Honolulu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hawaii</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>