<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Miracl at Clef 2014 : Ehealth Information Retrieval Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nesrine KSENTINI</string-name>
          <email>ksentini.nesrine@ieee.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamed Tmar</string-name>
          <email>mohamed.tmar@isimsf.rnu.tn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fa¨ıez GARGOURI</string-name>
          <email>faiez.gargouri@isimsf.rnu.tn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>MIRACL Laboratory, City ons Sfax, University of Sfax</institution>
          ,
          <addr-line>B.P.3023 Sfax</addr-line>
          <country country="TN">TUNISIA</country>
        </aff>
      </contrib-group>
      <fpage>203</fpage>
      <lpage>209</lpage>
      <abstract>
        <p>This paper presents our rst participation in user-centred health information retrieval task at the CLEFeHealth 2014. This task has as objective the information retrieval to answer patients' questions when reading clinical reports. We have submitted only the mandatory run (baseline system). The obtained results are motivating with map=0:1677 and p@10=0:5460 but can be improved.</p>
      </abstract>
      <kwd-group>
        <kwd>information retrieval</kwd>
        <kwd>vector space model</kwd>
        <kwd>medical documents</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>With the increasing number of documents on the web, searching relevant
documents to users queries becomes a difficult task especially if the queries are short
and do not represent well the user need. The process of searching documents in
a corpus in order to meet a need is called Information Retrieval (IR).
This process (IR) focuses on the representation, storage, organization of
information that should allow the user quick and easy access to information.
To automate the task of RI, Information Retrieval Systems (IRS) are developed
to provide all necessary functions for information retrieval illustrates in the
figure 1.</p>
      <p>The system builds an index of the documents which is an essential data structure
because it allows fast searching over large volumes of data.</p>
      <p>Given that the document database is indexed, the searching process can be
started. The user specifies a query which is parsed and transformed by the same
operations applied to the document database. After that the system retrieves
documents that are relevant to the query from the index and displays them to
the user with a ranked way.</p>
      <p>
        The user can examine the ranked documents. He can identify a subset of the
documents seen as relevant and initiate a relevance feedback step. The system
uses the documents selected by the user in order to refine the query.
The evaluation of such systems appears to be a necessity. This evaluation is
based on the notion of relevance [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To improve the relevance of IR in IRS,
several studies have been proposed several IR models: the Boolean model, the
probabilistic model, the vector model.
      </p>
      <p>{ Boolean Model:</p>
      <p>It is a model which is based on the set theory. Thus, the user query is
represented as a logical equation composed of keywords and Boolean (logic)
operators (OR, AND, NOT) and documents are represented by a list of keywords.
The research process with this model performs operations of union,
intersection and difference, defined by the existence or absence of index terms,
to achieve an exact match between the documents and the equation of the
query.
{ Probabilistic model:</p>
      <p>The probabilistic model addresses the problem of information retrieval in
a probabilistic framework. It allows modeling of the selection process
documents in an IRS based on probability theory. The basic principle of the
probabilistic model is to present the search results in an IRS order (O)
based on the probability of relevance of a document with respect to a query.
The basic idea is first to calculate two conditional probabilities P (R=D) and
P (N R=D) with a given query R is the relevance (all relevant documents)
and N R irrelevance (all irrelevant documents). The terms are not weighted,
but only take the value 0 (term absent) or 1 (this term). The principle of
the model is to calculate an ordering function (O(D)) that would classify
documents that meet a query with
{ Vector Space Model:</p>
      <p>
        The vector space model (VSM) was developed by Salton for the SMART
information retrieval system [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It is an algebraic model that uses a
vector representation for documents as well as queries in a large dimensional
space. The performance of this model depends on the weighting of terms
in these vectors which represents the degree of relationship between a term
and a document. The relevance of a document relative to a query is defined
by distance measurements in a vector space. The most widely used numeric
similarity measure is the cosine of the angle between the vector and the
vectors [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 4, 3</xref>
        ] .
      </p>
      <p>In our participation, we used Vector Space Model because it use the weighted
vector not the binary one and it allows computing a degree of similarity between
documents and queries; also it returns ranking documents according to their
relevance.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Material and Methods</title>
      <p>2.1</p>
      <sec id="sec-2-1">
        <title>Database</title>
        <p>
          Document Collection: The data set for the ehealth information retrieval task
consists of a set of medical-related documents, provided by the Khresmoi project
[
          <xref ref-type="bibr" rid="ref12 ref13 ref8">13, 12, 8</xref>
          ]. This collection contains documents covering a broad set of medical
topics, and does not contain any patient information. The documents in the
collection come from several online sources, including the Health On the Net
organisation certified websites, as well as well-known medical sites and databases
(e.g. Genetics Home Reference, ClinicalTrial.gov, Diagnosia).
        </p>
        <p>
          Topics: The test set consists of 50 medical professional queries containing the
following fields [
          <xref ref-type="bibr" rid="ref13 ref8">13, 8</xref>
          ]:
{ Title: text of the query.
{ Description: longer description of what the query means.
{ Narrative: expected content of the relevant documents.
{ Profile: main information on the patient (age, gender, condition).
{ Discharge summary: ID of the matching discharge summary provided by
Ehealth Information Extraction task.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Methods</title>
        <p>Our work was to index and search the top-1000 relevant medical documents for
each topic using a plain-text search engine.</p>
        <p>
          We chose to use the terrier platform in our case [
          <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
          ]. It is an efficient, effective
and flexible open source search engine, easily deployable on large-scale collections
of documents [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Terrier implements state-of-the-art of indexing and retrieval
functionalities. It is an open source, and a comprehensive and transparent
platform for research and experimentation in text retrieval. Terrier is written in Java,
and is developed at the School of Computing Science, University of Glasgow.
The first step made before you begin indexing was the extraction of
individual HTML documents from raw files. Each of individual document was saved in
its own file having the name of its UID mentioned in the raw file.
After extracting medical documents, we started the indexing step. It happened
offline and contains four steps:
Tokenization: is a process of breaking a text up into words called tokens.
Stop wording: most common words that would appear in the list of tokens are
excluded because they have a little value when searching documents to a user
need.
        </p>
        <p>Stemming: is a process of linguistic normalization for reducing variant forms
of tokens to their stem or root form.</p>
        <p>Storing: information (terms) were stored in file with special structure called
inverted file by specifying the number of occurrences (term frequency (TF) and
their location. This file allows rapid access during query time.</p>
        <p>The third step is searching relevant documents for a particular query; it was
carried online. Firstly, we processed query in the same way documents were
processed during indexing. Then we represented documents and query with weighted
vectors of terms which represents the degree of relationship between a term and
a document in the whole collection. This degree is obtained by multiplying the
two measures (TF) and (IDF) with</p>
        <p>IDFtermi = log</p>
        <p>j D j
j dj : termi 2 dj j
(2)
Where
j D j : total number of documents in the collection.
j dj : termi 2 dj j: number of documents where the termi appears
Then we calculated similarity (scores) between vectors using cosine similarity
measure and ranked the documents according to their scores in descending
order. The top 1000 documents with the highest scores were the best match to the
query and returned as relevant documents.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>In table 1 we show results obtained by taking account of pooling all search
results.</p>
      <p>Our system can retrive 38% (1189 among 3209) of relevant documents and has
a measure of map equal to 0:1677 and of p@10 equal to 0:5460.</p>
      <p>
        These results are motivating but can be improved by including for example the
notion of semantics in the search process [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">11, 10, 9</xref>
        ].
      </p>
      <p>The plots below in figure 2 compares our run against the median and best
performance (p@10) across all systems submitted to Ehealth task for each query.
In particular, for each query, the height of a bar represents the gain/loss of your
system and the best system (for that query) over the median system. The height
of a bar in then given by:</p>
      <p>medianp@10(q)
medianp@10(q)
We can see that our run has 3 queries that achieve the best performance, 17
queries perform better than the median, while 26 queries were worse than the
median, and other 5 queries perform in the median line. It means that in a
specific task such as medical document retrieval, results are motivating but need
improvement.</p>
      <p>In terms of the best performance of each query, some queries can achieve a very
high result like query 14,22,32,33,39,48 and 49 (more than 70%), while some
queries perform near the median like query 2 and 20 (only about 10%).</p>
      <p>We show in figure 3, the average recall-precision curve for all 50 topics. There
is a tradeoff between recall and precision. We notice that when increasing recall
by retrieving more, precision decreases.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and future work</title>
      <p>
        In our first participation in ehealth information retrieval task at the
CLEFeHealth 2014, we obtained motivating results for the search in a large collection
of medical documents to answer patients’ questions when reading clinical
reports. For future work, we will try to improve the obtained results by including
the notion of semantics between terms of documents [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Wei</surname>
          </string-name>
          , Xing and Croft, W Bruce:
          <article-title>Investigating retrieval performance with manuallybuilt topic models</article-title>
          . pp.
          <volume>333</volume>
          {
          <issue>349</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Salton</surname>
          </string-name>
          , Gerard and Wong, Anita and Yang,
          <string-name>
            <surname>Chung-Shu</surname>
          </string-name>
          :
          <article-title>A vector space model for automatic indexing</article-title>
          .
          <source>In :Journal of Communications of the ACM</source>
          . pp.
          <volume>613</volume>
          {
          <issue>620</issue>
          , vol.
          <volume>18</volume>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Aswani</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Ch and Radvansky, M and Annapurna</article-title>
          ,
          <string-name>
            <surname>J:</surname>
          </string-name>
          <article-title>Analysis of a Vector Space Model, Latent Semantic Indexing and Formal Concept Analysis for Information Retrieval</article-title>
          .
          <source>In :Journal of Cybernetics and Information Technologies</source>
          . pp.
          <volume>34</volume>
          {
          <issue>48</issue>
          , vol.
          <volume>12</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Turney</surname>
            ,
            <given-names>Peter D</given-names>
          </string-name>
          and
          <article-title>Pantel, Patrick and others: From frequency to meaning: Vector space models of semantics</article-title>
          .
          <source>In :Journal of arti cial intelligence research</source>
          . pp.
          <volume>141</volume>
          {
          <issue>188</issue>
          , vol.
          <volume>37</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <surname>Rodrygo</surname>
            <given-names>LT</given-names>
          </string-name>
          and McCreadie, Richard and Plachouras, Vassilis:
          <article-title>Large-scale information retrieval experimentation with Terrier</article-title>
          . pp.
          <volume>2601</volume>
          {
          <issue>2602</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Amati</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Plachouras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Macdonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Terrier: A High Performance and Scalable Information Retrieval Platform</article-title>
          . (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ;Lioma,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Plachouras</surname>
          </string-name>
          ,V.: Research Directions in Terrier. In :Journal of Novatica/UPGRADE Special Issue on Web Information Access, Ricardo Baeza-Yates et al. (Eds), Invited Paper. (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Liadh</given-names>
            <surname>Kelly</surname>
          </string-name>
          and
          <article-title>Lorraine Goeuriot and Hanna Suominen and Danielle L. Mowery and Sumithra Velupillai and Wendy W. Chapman and Guido Zuccon and Joao Palotti: Overview of the ShARe/CLEF eHealth Evaluation Lab 2014</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          <year>2014</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Agirre</surname>
          </string-name>
          ,
          <article-title>Eneko and Alfonseca, Enrique and Hall, Keith and Kravalova, Jana and Pasca, Marius and Soroa, Aitor: A study on similarity and relatedness using distributional and WordNet-based approaches</article-title>
          .
          <source>In: Proceedings of Human Language Technologies</source>
          . pp.
          <volume>19</volume>
          {
          <issue>27</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <article-title>Eneko Agirre and Montse Cuadros and German Rigau and Aitor Soroa: Exploring Knowledge Bases for Similarity</article-title>
          .
          <source>In: Proceedings of the Seventh International Conference on Language Resources and Evaluation</source>
          . (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <article-title>Ksentini Nesrine and Tmar Mohamed and Gargouri Faiez: Detection of Semantic Relationships between Terms with a New Statistical Method</article-title>
          .
          <source>In: Proceedings of the 10th International Conference on Web Information Systems and Technologies (WEBIST</source>
          <year>2014</year>
          ). pp.
          <volume>340</volume>
          {
          <issue>343</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Goeuriot</surname>
          </string-name>
          , Lorraine and Hanbury, Allan and Jones,
          <string-name>
            <surname>Gareth</surname>
            <given-names>JF</given-names>
          </string-name>
          and
          <article-title>Kelly, Liadh and Kriewel, Sascha and Martinez Rodriguez, Ivan and Muller, Henning and Tinte, Miguel: Supporting collaborative improvement of resources in the Khresmoi health information system</article-title>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <article-title>Lorraine Goeuriot and Liadh Kelly and Wei Li and Joao Palotti and Pavel Pecina and Guido Zuccon and Allan Hanbury and Gareth Jones and Henning Mueller: ShARe/CLEF eHealth Evaluation Lab 2014, Task 3: User-centred health information retrieval</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          <year>2014</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>