<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Description of the POMELO System for the Task 2 of QALD-4</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thierry Hamon</string-name>
          <email>thierry.hamon@limsi.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natalia Grabar</string-name>
          <email>natalia.grabar@univ-lille3.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fleur Mougin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frantz Thiessard</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LIMSI-CNRS</institution>
          ,
          <addr-line>Orsay</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>STL UMR8163 CNRS, Universite Lille 3</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universite Bordeaux, ISPED, Centre INSERM U897</institution>
          ,
          <addr-line>ERIAS</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universite Paris 13</institution>
          ,
          <addr-line>Sorbonne Paris Cite</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <fpage>1212</fpage>
      <lpage>1223</lpage>
      <abstract>
        <p>In this paper, we present the POMELO system developed for participating in the task 2 of the QALD-4 challenge. For translating natural language questions in SPARQL queries we exploit Natural Language Processing methods, semantic resources and RDF triples description. We designed a four-step method which pre-processes the question, performs an abstraction of the question, then builds a representation of the SPARQL query and nally generates the query. The system was ranked second out of three participating systems. It achieves good performance with 0.85 F-measure on the set of 25 test questions.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Language Processing</kwd>
        <kwd>SPARQL</kwd>
        <kwd>biomedical domain</kwd>
        <kwd>semantic resources</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Biomedical knowledge is disseminated in knowledge bases which become
increasingly available on the Web. These knowledge bases usually focus on a given type
of information: chemical, pharmacological and target information on drugs in
Drugbank [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], clinical studies in ClinicalTrials.gov5, drugs and their side e ects
in Sider [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], etc. The connection of such life-science knowledge bases is crucial for
obtaining more global and comprehensive view on the links that may exist
between di erent biomedical components, factors and actors. Moreover, this allows
inducing and producing new knowledge from the already available data.
Particularly, the creation of ne-grained links between the existing knowledge bases
related to drugs is a great challenge that is being addressed by the project Linked
Open Drug Data (LODD) for instance6. In this project, the knowledge recorded
in the knowledge bases and dataset interlinks is represented as RDF triples, on
the basis of which the linked data can then be queried through a SPARQL
endpoint. However, typical users of this knowledge, such as physicians, life-science
researchers or even patients, cannot manage the syntactic and semantic
requirements of the SPARQL language neither can they manage the structure of various
knowledge bases. This situation impedes the e cient use of knowledge bases and
the retrieval of useful information. Therefore, it is important to design friendly
interfaces that mediate the technical and semantic complexity of the task and
provide simple approaches for querying the knowledge bases.
      </p>
      <p>
        For instance, it has been shown that for querying the knowledge bases and
the Semantic Web data, the use of full and standard sentences is preferred to
the use of keywords, menus or graphs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. While this study is conducted on
general knowledge data, we assume this observation is also relevant for the users
of biomedical knowledge bases. Up to now the design of friendly user interface
is mainly addressed for general knowledge bases [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We can also mention
another work which aim is to translate medical questions issued from a journal
into SPARQL queries [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>In relation with such research problems, the Question Answering over Linked
Data (QALD-4) campaign proposes a task dedicated to the retrieval of precise
biomedical information in linked knowledge bases according to questions in
natural language. We present in this paper the methodology we propose to translate
natural language questions in SPARQL queries and the system we developed for
our participation to the challenge.</p>
      <p>We start with the de nitions of the main terms used in the proposed
presentation (Sect. 2). Then, we describe the semantic resources available and
developed for enriching the questions (Sect. 3). The methodology and the system
are described in Sect. 4. The evaluation of the system on the QALD-4 queries is
presented in Sect. 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Terminology</title>
      <p>The main terms are used with the following meaning:
Question The questions are the natural language expressions uttered by human
users in order to formulate their information need.</p>
      <p>Query The queries respect the SPARQL syntax and semantics. They are
created automatically on the basis of (natural language) questions.
Semantic type Semantic types are indicative of the word meaning. The
semantic types are de ned from URI pre xes such as disease for rickets, drug
for Cetuximab, etc.</p>
      <p>Frame Frames are de ned as linguistic representations of RDF schema. Usually
the frames contain one predicate and at least two elements with associated
semantic types.</p>
      <p>Linguistic annotation The linguistic annotations are obtained with the
Natural Language Processing (NLP) tools and provide lemma, part-of-speech
categories, and terms.</p>
      <p>Semantic annotation Semantic annotations are obtained by tagging questions
with semantic resources.
The main challenge of the proposed work is to design the optimal
methodology for an easy and reproducible rewriting of (natural language) questions in
SPARQL queries.
3</p>
    </sec>
    <sec id="sec-3">
      <title>De nition of the Semantic Resources</title>
      <p>Some resources used are provided by the challenge organizers (Sect. 3.1), others
are collected and built speci cally for the challenge (Sect. 3.2) to support the
method. We use these resources for rewriting the questions in queries.
3.1</p>
      <sec id="sec-3-1">
        <title>Resources Provided by the QALD Challenge</title>
        <p>
          Three datasets are provided by the QALD challenge:
{ Drugbank7 is dedicated to drugs [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. It merges chemical, pharmacological
and pharmaceutical information from other available knowledge bases. We
exploited the documentation8 of this resource to de ne rewriting rules and
regular expressions in our named entity recognizer.
{ Diseasome9 is dedicated to diseases and genes linked among them by known
disorder/gene associations [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. It provides a single framework with all known
phenotypes and disease gene associations, indicating the common genetic
origin of many diseases. We exploited the RDF triples and the documentation
of the resource to de ne the rewriting rules.
{ Sider10 is dedicated to adverse drug e ects [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. It contains information on
marketed medicines and their recorded adverse drug reactions. The
information is extracted from public documents and package inserts. The available
information includes side e ect frequency, drug and side e ect classi cations
as well as links to other information, for example drug-target relations.
The content of each resource is provided in a speci c format: RDF triples subject
predicate object, so that they encode the useful and usable frame elements.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Resources Collected and Built for the QALD Challenge</title>
        <p>On the basis of the RDF triples, we build frames from the RDF schema where
the RDF predicate is the frame predicate, and subject and object of the RDF
triples are the core frame elements. This also includes the OWL sameAs triples.
Several types of frame entities are isolated:
{ As indicated, subject, object and predicate become semantic entities. They
may occur in questions: in this way, the frames are the main resource for
rewriting questions in queries.
7 http://www.drugbank.ca
8 http://www.drugbank.ca/documentation
9 http://diseasome.eu
10 http://sidee ects.embl.de
{ The vocabulary speci c to questions is also built. It covers for instance
aggregation operators, negation and types of questions.
{ RDF literals, issued from named entity recognizer or term extractor,
complete the resources. The RDF literals are detected with speci cally designed
automata that may rely on the source knowledge base documentation.
These entities are associated with the expected semantic type, which allows
creating the queries and rewriting the RDF triples in the SPARQL queries. In that
respect, we can consider IRI, strings, common datatype or regular expressions
when literals are expected.</p>
        <p>Most of the entities are considered and processed through their semantic
type, although some ambiguous entities (e.g. interaction or class) are considered
atomically. For these, the rewriting rules will be applied contextually to generate
the semantic entities corresponding to the frames (see Sect. 4.2). When using
the queries, the semantic types are variables and are used for connecting the
edges of queries.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>System Description</title>
      <p>We design a four-step process based on NLP methods, semantic resources and
RDF triple description (see Fig. 1):
1. We pre-process the questions in order to enrich them with linguistic and
semantic information (Sect. 4.1).
2. We perform a question abstraction (Sect. 4.2).
3. We use the abstracted question to construct the corresponding SPARQL
query representation (Sect. 4.3);
4. We generate the SPARQL query (Sect. 4.4).</p>
      <p>
        The process is implemented as a module within the NLP platform Ogmios [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
4.1
      </p>
      <sec id="sec-4-1">
        <title>Pre-processing</title>
        <p>
          The pre-processing of the questions is the same for the training and test sets. The
annotation of the questions consists in recognition of numerical values (such as
numbers and solubility values), word segmentation, part-of-speech tagging and
lemmatization of the words with TreeTagger [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Then, we apply the TermTagger
Perl module11 for identifying semantic entities, i.e. terms with associated
semantic types. TermTagger exploits the semantic resources (see Sect. 3) to recognize
semantic entities such as disease names, side e ects, etc.
        </p>
        <p>
          However, as we realized during our preliminary experiments on the training
set, the coverage of the terminological entities that appear in questions is not
su cient. We also apply the term extractor YATEA12 [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] to improve the coverage
of our method. The term extractor performs shallow parsing of the POS-tagged
11 http://search.cpan.org/~thhamon/Alvis-TermTagger/
12 http://search.cpan.org/~thhamon/Lingua-YaTeA/
Pre-processing
        </p>
        <p>question</p>
        <sec id="sec-4-1-1">
          <title>Linguistic and semantic annotation</title>
          <p>SPARQL query generation</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>Question abstraction</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>Query construction</title>
        </sec>
        <sec id="sec-4-1-4">
          <title>Query generation</title>
        </sec>
        <sec id="sec-4-1-5">
          <title>SPARQL query</title>
        </sec>
        <sec id="sec-4-1-6">
          <title>Number</title>
        </sec>
        <sec id="sec-4-1-7">
          <title>Solubility</title>
        </sec>
        <sec id="sec-4-1-8">
          <title>Drugbank</title>
        </sec>
        <sec id="sec-4-1-9">
          <title>Diseasome</title>
        </sec>
        <sec id="sec-4-1-10">
          <title>Sider</title>
        </sec>
        <sec id="sec-4-1-11">
          <title>Query vocabulary</title>
        </sec>
        <sec id="sec-4-1-12">
          <title>Argument description</title>
        </sec>
        <sec id="sec-4-1-13">
          <title>Predicate description</title>
          <p>sameAs description
resources in blue)
test set because it is time consuming.</p>
          <p>test set because it is time consuming.
and lemmatized text by chunking it according to syntactic frontiers (pronouns,
and lemmatized text by chunking it according to syntactic frontiers (pronouns,
conjugated verbs, typographic marks, etc.) in order to identify noun phrases.
conjugated verbs, typographic marks, etc.) in order to identify noun phrases.</p>
          <p>13
Then, parsing patterns 13that take into account the morpho-syntactic variation,</p>
          <p>Then, parsing patterns that take into account the morpho-syntactic variation,
are recursively applied and provide parsed terminological entities, usually noun
are recursively applied and provide parsed terminological entities, usually noun
phrases relevant for the targeted domain. Each term is represented as a syntactic
phrases relevant for the targeted domain. Each term is represented as a syntactic
tree, and sub-terms are also considered as terms in the current con guration.</p>
          <p>tree, and sub-terms are also considered as terms in the current configuration.
illustrates the linguistic and semantic annotation of a question.</p>
          <p>illustrates the linguistic and semantic annotation of a question.</p>
          <p>No semantic types are associated to the terms extracted by Y T A. Figure 2
No semantic types are associated to the terms extracted by AYATEEA. Figure 2
On the training step, we observed that the lemmatizer and POS-tagger Genia</p>
          <p>
            On the training step, we observed that the lemmatizer and POS-tagger Genia
Tagger [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] performs a better lemmatization. However, we did not use it on the
Tagger [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] performs a better lemmatization. However, we did not use it on the
Give
          </p>
          <p>me drugs
VB PRO</p>
          <p>N
[drug]
in the</p>
          <p>P</p>
          <p>DT
gaseous</p>
          <p>A
[drugbank/state/Gas]
state</p>
          <p>.</p>
          <p>N
[state]
bobxoexsesrerperpersesnetnt
htehewworodrdanadndsesmemanatnitcicenetnitietse.s.TThehesusubbscsrcirpiptttetxextstsaarereththeePPaartr-to-of-f</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.42.2 QQueusetsitoinonAAbsbtsrtarcatcitoinon</title>
        <p>ThTihsisstsetpepaiami mssatatidiednetnitfyifiynigngthtehererleelveavnatnteleelmemenentstswwitihthininththeeqquueestsitoionnssaanndd
bubiuldilidnigngthtehererperperseesnetnattaitoinonofofthtehseeseeleelmemenetnst.s.AAt tththisisstsetpep,,wweeuuseseththeelilningguuisistticic
anadndsesme manatnicticanannontoattaitoinosnsasassoscoicaitaetdedtotoththeeququesetsitoinonwwoordrdssininththeepprerevvioiouussstsetepp. .</p>
        <p>BeBfeofroerethteheidiednetniticfiactaitoinonofofrerleelveavnatnteleelmemenentst,s,ththeeddisiasammbbigiguuaatitoionnooffththeeaann-fofrotrhtehenenxetxsttsetpesp.sF.oFrorinisntsatnacnec,ew,weekekeepeplalragregrertetremrms swwhhicihchddoonnoottininccluluddeeooththeerr
sesme manatnicticenetnititeise.s.</p>
        <p>TThuhsu,sw,weededenfiendedrerwe wrirtitnigngrurluelsesininorodrderertotommooddifiyfyoorrddeleelteeteththeesesemmaanntticic
mmodoidf
yifyorordedlelteetethteheenetnitiytyacaccocrodridnigngtotoththeecocnontetxext.t.FFoorrininstsatannccee,,ththeesesemmaann-cocnotnetxetxtcocnotnatianisnsa atetremrmwwitihththtehesesmemanatnitcictytpypeefofodod..OOnnththeewwhhoolele,,wweeddeefinneedd
4444cocnotnetxetxutaularlerwewrirtitnigngrurluelsesbabsaesdedononththeevovcoacbabuulalrayryuusesdedininththeeqquueestsitoionnssaanndd
1313
dedpeepnednednecnycyinintetremr minionlolgoigciaclaelnetnititeise.s.
on the documentation of knowledge bases, mainly the one from Drugbank14.
Besides the rewriting rules, additional disambiguation of the annotations is also
performed during the query construction step when the arguments of the
predicate are connected by selecting the correct semantic types.</p>
        <p>For performing the abstraction of questions, we identify information related
to the query structure:
1. De nition of the Result form: the question is scanned for identifying words
expressing the negation, e.g. no, and its scope, the aggregation operation
on the results, e.g. number for count, mean for avg or higher for max, and
speci c result form such as boolean queries (ASK). The information
concerning the presence of the negation and aggregation operators or of speci c
result form is recorded in data structures to be used at the end of the query
construction step or during the query generation step.
2. Identi cation of the Question topic: we consider the rst semantic entity
with a given expected semantic type to be the question topic. The expected
semantic types are those provided by the RDF subjects and objects in the
Drugbank, Diseasome and Sider. This information will be used during the
query construction step.
3. Identi cation of Predicate and Argument: according to our internal
framebased representation of the three resources, the potential predicates, subjects
and objects are identi ed among the semantic entities and described into a
symbol table.15 At this step, the subjects and objects are fully described in
the symbol table. Concerning the predicates, only the semantic types of their
arguments are instantiated in the symbol table within the RDF schema. The
subjects and objects can be URI, RDF typed literals (numerical values or
strings) and extracted terms (these are considered as elements of regular
expressions).
The objective of the query construction step is to connect previously identi ed
elements and to build the SPARQL graph pattern (introduced by the keyword
WHERE). Thus, the symbols of the predicate arguments are instantiated by either
URI associated with the subjects and objects, variables, numerical values or
strings.</p>
        <p>For each question, we perform several connections:
14 http://www.drugbank.ca/documentation
15 Similarly to a compiler, RDF schema of the predicates and semantic types of their
arguments are considered as entries of the symbol table. Information associated with
the symbols (in ected and lemmatized form of the word or term, corresponding</p>
        <sec id="sec-4-2-1">
          <title>SPARQL type of the symbol, semantic types of the arguments, indicators of the use of the symbol as object or subject of a predicate, etc.) are recorded in a data structure.</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>Agregation operator:</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>Question topic:</title>
        </sec>
        <sec id="sec-4-2-4">
          <title>Predicates:</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>Arguments:</title>
          <p>drug (sider/drugs)</p>
        </sec>
        <sec id="sec-4-2-6">
          <title>Frame</title>
          <p>drugbank/drugs state STRING</p>
        </sec>
        <sec id="sec-4-2-7">
          <title>Semantic type Word SPARQL type</title>
          <p>drugbank/state/Gas gaseous STRING/Gas
1. The question topic is connected to the predicate(s). A variable is associated
with the question topic and the predicate arguments that matched the
semantic type of the question topic. Note that at the end of this stage, the
question topic may remain non-associated with any predicate;
2. The predicate arguments are connected to the subjects and objects identi ed
during the question abstraction: they concern elements referring to URI;
3. The predicates are connected between them through their subjects and
objects. The connection between two predicates is then represented by a
variable;
4. The predicates from di erent datasets are connected. We use the sameAs
description to identify URI referring to the same element. New variables are
de ned to connect two predicates;
5. The remaining question topic is connected to arguments of the sameAs
predicate;
6. The arguments are connected to the string type to extracted terms
annotated in the question. We assume these arguments will be related to the string
matching operator REGEX. Thus, terms are considered as string expressions.</p>
          <p>At this point, the predicate arguments which remain unassociated are
replaced by new variables in order to avoid empty literals. Finally, the negation
operator is processed: the predicates are marked as negated and the arguments
within the negation scope are included in a new predicate rdf:type if required.</p>
          <p>At this stage, the question is fully translated into data structures
representing the SPARQL query. Figure 4 illustrates the construction of the query
corresponding to the question 22 of the test set.</p>
        </sec>
        <sec id="sec-4-2-8">
          <title>Agregation operator:</title>
        </sec>
        <sec id="sec-4-2-9">
          <title>Question topic: ?v0</title>
        </sec>
        <sec id="sec-4-2-10">
          <title>Predicates:</title>
        </sec>
        <sec id="sec-4-2-11">
          <title>Arguments:</title>
        </sec>
        <sec id="sec-4-2-12">
          <title>Frame</title>
          <p>?v0 state STRING/Gas</p>
        </sec>
        <sec id="sec-4-2-13">
          <title>Semantic type</title>
          <p>drugbank/state/Gas gaseous STRING/Gas</p>
        </sec>
        <sec id="sec-4-2-14">
          <title>Word SPARQL type Fig. 4. Example of query construction (question#22 of the test set)</title>
          <p>This nal step aims at generating the SPARQL query string based on the data
structures built during the query construction step. The output of this step is
the string corresponding to the SPARQL query. It is composed of two parts:
1. The generation of the result form which takes into account the expected type
of the result form (ASK or SELECT), the presence of aggregation operators and
the variable associated to the question topic;
2. The generation of the graph pattern. Basically, the part of the query
generation consists in generating the strings representing each RDF triple and
lter if predicates are negated. But when aggregation operators are used,
we also need to recursively generate sub-queries computing the subsets of
expressions, before their aggregation.</p>
          <p>The SPARQL queries have been submitted without retrieving the answers.
We let this task to the evaluation tool. Figure 5 presents the generated query
which corresponds to the question 22 of the test set.</p>
          <p>SELECT DISTINCT ?v0
WHERE {
?v0 &lt;http://www4.wiwiss.fu-berlin.de/drugbank/resource/drugbank/state&gt; "Gas".
}</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <sec id="sec-5-1">
        <title>Evaluation Metrics</title>
        <p>The automatically generated SPARQL queries are evaluated with the online
evaluation tool16. The answer of each query q is compared with the gold
standard. The evaluation measures (F-measure, precision and recall) are computed.
The system results are evaluated with macro-measures. The challenge provides
25 questions for the training and 25 questions for the test.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Global Results</title>
        <p>In Tab. 1, we present the overall results on the training and test sets. Our system
was ranked 2nd out of 3 submissions. We can observe that our results are similar
for these two sets of queries (training and test sets).</p>
        <p>We can observe that in both cases, for 19 questions, the system exactly
provides the expected answers. For 4 questions from the training set and 3 questions
from the test set, we obtain partial answers. In the training set, 2 questions
receive no answer while in the test set, we have 3 such questions.
16 http://greententacle.techfak.uni-bielefeld.de/~cunger/qald/index.php?x=evaltool&amp;q=4
Error Analysis To our opinion, the reference SPARQL query for the question
19 of the training set is not correct: the expected result of the question is a list
of drugs, while the reference SPARQL query returns a list of diseases. On the
test set, we can propose two observations on the limitations of our system:
{ In question 1,17 the contextual rewriting rules cannot be correctly applied
because the semantic entity (gene ... associated) is discontinued;
{ In question 18,18 the system correctly detects the semantic entities and the
predicates, including the sameAs predicate. The remaining problem is that
the system assumes that the sameAs predicate is re exive while in the
resources provided, the instances of this predicate do not encode the re exivity
of the relation.
We analyzed the system performance on a standard computer (2.7GHz dual-core
CPU and 4 Gb of memory). Figure 6 presents the running time for each query
ac17 Which genes are associated with Endothelin receptor type B?
18 List the number of distinct side-e ects of drugs which target genes whose general
function involves cell division.
)
s
d
n
o
c
e
s
(
e
m
i
t
n
u
r
3
2
1
0
5
10
15
20</p>
        <p>25
number of questions
cording to the pre-processing sub-steps (named entity recognition, word and
sentence segmentation, POS tagging, semantic entity tagging and term extraction)
and the question translation into SPARQL queries (Question2SPARQLQuery).
Most of the processing time is dedicated to the TermTagger which aims at
recognizing the semantic entities. With the internal Ogmios processing (i.e. mainly
the input/output), each question is processed in 1.94 seconds on the average on
the training set and 1.97 seconds on the average on the test set.
We proposed a four-step process based on Natural Language Processing
methods, semantic resources and RDF triples description. The system achieves good
performance with 0.85 F-measure on the test set of 25 questions. It is ranked
2nd out of 3 submissions. Further work includes optimization of the running time
for the question processing and the extension of the SPARQL syntax by taking
into account the operators on collections of sets. We also plan to investigate the
integration of other biomedical resources such as Dailymed or RxNorm, and the
use of the developped system for the text mining applications.
This work was partly funded through the project POMELO (PathOlogies,
MEdicaments, aLimentatiOn) funded by the MESHS (Maison europeenne des sciences
de l'homme et de la societe) under the framework Projets Emergents.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Medical question answering: Translating medical questions into sparql queries</article-title>
          .
          <source>In: ACM SIGHIT International Health Informatics Symposium (IHI</source>
          <year>2012</year>
          )
          <article-title>(</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Aubin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Improving term extraction with terminological resources</article-title>
          . In: Salakoski,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Ginter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Pyysalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Pahikkala</surname>
          </string-name>
          , T. (eds.)
          <source>Advances in Natural Language Processing (5th International Conference on NLP, FinTAL</source>
          <year>2006</year>
          ). pp.
          <volume>380</volume>
          {
          <fpage>387</fpage>
          . No. 4139
          <string-name>
            <surname>in</surname>
            <given-names>LNAI</given-names>
          </string-name>
          , Springer (
          <year>August 2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Damljanovic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agatonovic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cunningham</surname>
          </string-name>
          , H.:
          <article-title>Natural language interfaces to ontologies: Combining syntactic analysis and ontology-based lookup through the user interaction</article-title>
          .
          <source>In: Proceedings of the 7th International Conference on The Semantic Web: Research and Applications - Volume Part I</source>
          . pp.
          <volume>106</volume>
          {
          <fpage>120</fpage>
          . ESWC'
          <volume>10</volume>
          , Springer-Verlag, Berlin, Heidelberg (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hamon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nazarenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poibeau</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aubin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derivire</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A robust linguistic platform for e cient and domain speci c web content analysis</article-title>
          .
          <source>In: Proceedings of RIAO 2007</source>
          . Pittsburgh, USA (
          <year>2007</year>
          ), http://riao.free.fr/papers/64.pdf, 15 pages
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Janji</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prulj</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          :
          <article-title>The core diseasome</article-title>
          .
          <source>Mol Biosyst</source>
          <volume>8</volume>
          (
          <issue>10</issue>
          ),
          <volume>2614</volume>
          {2625 (Aug
          <year>2012</year>
          ), http://dx.doi.org/10.1039/c2mb25230a
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kaufmann</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>How useful are natural language interfaces to the semantic web for casual end-users?</article-title>
          <source>In: Proceedings of the Forth European Semantic Web Conference (ESWC</source>
          <year>2007</year>
          ). Innsbruck,
          <string-name>
            <surname>Austria</surname>
          </string-name>
          (
          <year>June 2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kuchmann-Beauger</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aufaure</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Natural language interfaces for datawarehouses</article-title>
          . In:
          <article-title>8mes journes francophones sur les Entrepts de Donnes et l'Analyse en ligne (EDA 2012), Bordeaux</article-title>
          . RNTI, vol. B-
          <volume>8</volume>
          , pp.
          <volume>83</volume>
          {
          <fpage>92</fpage>
          . Hermann, Paris (
          <year>Juin 2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kuhn</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campillos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Letunic</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bork</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A side e ect resource to capture phenotypic e ects of drugs</article-title>
          .
          <source>Molecular Systems Biology</source>
          <volume>6</volume>
          (
          <issue>1</issue>
          ) (
          <year>2010</year>
          ), http://msb.embopress.org/content/6/1/343
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Schmid</surname>
          </string-name>
          , H.:
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          . In: Jones,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Somers</surname>
          </string-name>
          , H. (eds.)
          <article-title>New Methods in Language Processing Studies in Computational Linguistics (</article-title>
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tsuruoka</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tateishi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohta</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McNaught</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsujii</surname>
          </string-name>
          , J.:
          <article-title>Developing a robust part-of-speech tagger for biomedical text</article-title>
          .
          <source>In: Proceedings of Advances in Informatics - 10th Panhellenic Conference on Informatics</source>
          . pp.
          <volume>382</volume>
          {
          <fpage>392</fpage>
          . LNCS 3746 (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wishart</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knox</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>A.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shrivastava</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hassanali</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stothard</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Woolsey</surname>
          </string-name>
          , J.:
          <article-title>Drugbank: a comprehensive resource for in silico drug discovery and exploration</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>34</volume>
          ,
          <issue>D668D672</issue>
          (
          <year>2006</year>
          ), database issue
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>