<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring effective information retrieval technique for the medical web documents: SNUMedinfo at CLEFeHealth2014 Task 3</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sungbin Choi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jinwook Choi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Medical Informatics Laboratory, Seoul National University</institution>
          ,
          <addr-line>Seoul</addr-line>
          ,
          <country>Republic of Korea</country>
        </aff>
      </contrib-group>
      <fpage>167</fpage>
      <lpage>175</lpage>
      <abstract>
        <p>This paper describes the participation of the SNUMedinfo team at the CLEFeHealth2014 task 3. We submitted 7 runs to Task3a (monolingual information retrieval): 1 baseline run using query likelihood model in Indri search engine; 3 runs applying UMLS based lexical query expansion utilizing discharge summary as an expansion term filter; 3 runs applying learning to rank technique utilizing various document features. We submitted 4 runs to Task3b (multilingual information retrieval): 1 baseline run using Google Translate for the English translation; 3 runs applying learning to rank technique on the translated query.</p>
      </abstract>
      <kwd-group>
        <kwd>Learning to Rank</kwd>
        <kwd>Document quality</kwd>
        <kwd>Query expansion</kwd>
        <kwd>UMLS</kwd>
        <kwd>Web document</kwd>
        <kwd>Medical information retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In this paper, we describe the methods used for our participation of the
CLEFeHealth2014 Task 3 User-centered health information retrieval. For detailed task
description, please see the overview paper of task3 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <sec id="sec-2-1">
        <title>Baseline run</title>
        <p>
          We submitted 1 baseline run (SNUMEDINFO_EN_Run.1) using unigram
language model with Dirichlet smoothing [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] in the Indri search engine [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Default
parameter setting is used. Documents are indexed with Indri without stopword removal.
Only title field is used as a query. The queries are stopped at the query time using the
standard 418 INQUERY stopword list, case-folded, and stemmed using Porter
stemmer.
        </p>
        <p>
          In task3b, queries are expressed in French, German and Czech. We used Google
Translate [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] for the translation of queries into English. Then, this translated query is
applied to the unigram language model.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Lexical query expansion using discharge summary as an expansion term filter</title>
        <p>We submitted 3 runs applying UMLS based lexical query expansion, utilizing
discharge summary as an expansion term filter. In this method, document content of
discharge summary is assumed as user context information.</p>
        <p>
          Firstly, MetaMap [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] is applied to the query, and UMLS concepts are recognized. We
took all of the concepts in the top-scoring final mapping results from MetaMap.
Preferred terms per each UMLS concepts are extracted as a candidate expansion terms.
Candidate expansion terms which do not occur in the discharge summary are
removed, and remaining expansion terms are used as expansion terms. For example, if
UMLS concept C0559769 : ‘Pelvic cavity structure’ is recognized, but discharge
summary text does not contain term ‘cavity’ and ‘structure’, only ‘pelvic’ is used as
an expansion term.
        </p>
        <p>Original query part is weighted by 0.9 and expansion query part is weight by 0.1.
e.g., Indri query example
#weight (
0.9 #combine ( convalescence open pelvic fracture right superior rami fracture )
0.1 #combine ( fractures open pelvis pelvic right superior ) )
We applied different parameter settings per each run. In Run Number 2 (which
corresponds to SNUMEDINFO_EN_Run.2), we used title field as a query. In Run Number
3, we used title, description and narrative field as a query. In Run Number 4, we used
title field as a query, and Dirichlet prior parameter is changed to 1,000.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Applying learning to rank technique</title>
        <p>We applied learning to rank to incorporate various document features into the ranking
model.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Document features.</title>
        <p>The following features are extracted per each top 1,000 documents.
1. Relevance score of query likelihood model from fulltext index
2. Rank of query likelihood model from fulltext index
3. Relevance score of query likelihood model from title+url index
4. Rank of query likelihood model from title+url index
5. Document quality features
Feature 1 and 2 refers to the relevance score and rank acquired from baseline retrieval
model on fulltext index.</p>
        <p>Feature 3 and 4 refers to the retrieval result acquired from the index built from
keyword content of document. We assumed that terms occurring in the title (e.g., &lt;title&gt;
Community-acquired pneumonia &lt;/title&gt;) and the url field (e.g.,
http://www.merckmanuals.com/home/lung_and_airway_disorders/pneumonia/commu
nity-acquired_pneumonia.html) as keywords of document.</p>
        <p>With regard to the feature 5, we hypothesized a certain notion of document content
quality of medical web documents, which defines ideal characteristics that relevant
documents could have in common across different types of queries. For example,
content reliability, comprehensiveness, whether writing is well structured and so on.</p>
        <p>
          For general web search task, there are several prior studies trying to assess web
page quality [
          <xref ref-type="bibr" rid="ref10 ref6 ref7 ref8 ref9">6-10</xref>
          ]. Many of them utilized link analysis techniques such as
PageRank, textual features, webpage design features and so on. In [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], Bendersky et al.
tried quality-biased ranking. They used several features to evaluate web document
quality such as number of visible terms on the page, depth of the URL path or fraction
of table text on the page. These features are informative to assess readability and
layout of web page. They combined document quality assessment with query-document
relevance ranking
        </p>
        <p>
          In the medical domain, several prior studies tried to evaluate methodological
quality of medical literatures [
          <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
          ], or reliability of medical web document [
          <xref ref-type="bibr" rid="ref14 ref15 ref16 ref17 ref18">14-18</xref>
          ].
Those prior works focused on the quality classification task only.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], Choi et al. tried to incorporate document evidence quality into document
ranking for the high quality literature search. Document quality is defined by the
methodological evidence quality of the literature. Target corpus is research literatures
in MEDLINE, not a web document. Target users are professionals such as medical
doctors and researchers, not the general public. We think that criteria for evaluating
document quality in [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] is different from our task.
        </p>
        <p>In this study, we tried to combine medical web page quality assessment with
topical relevance ranking. We used only textual features for medical webpage quality
assessment. We tried to identify terms possibly relevant to the medical document
quality evaluation. Each term’s document term frequency is used as a feature value.
First author arbitrarily collected about 80 words which considered to be relevant to
the document quality (Table 5). These terms are punctuation removed, case-folded,
tokenized and stemmed, so finally 82 unique terms are used. We included terms such
as ‘etiology’, ‘prognosis’, and ‘treatment’. When we write any text, we often asked to
write in terms of 6W such as what, why, where. We thought that these ‘etiology’,
‘prognosis’ terms could be necessary condition to be well-organized medical
document, analogous to the 6W. We also included terms like ‘md’. When medical doctors
wrote document content, in many cases they write their name and qualification at the
end of text. We thought that these information could give some assurance about the
quality of document content.</p>
        <p>
          We used CLEFeHealth 2013’ Task 3 [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] test collection as a training set. Relevance
assessments are conducted on a 4 point scale (0-4), and a scale 1 correspond to a
document that was topical to the query, but unreliable. There are 50 queries, and 6,217
query-document paired relevance assessment in CLEFeHealth 2013’. Using baseline
method described in Section 2.1, we retrieved 1,000 documents per each query and
prepared dataset. Documents whose relevance assessment is not conducted on
CLEFeHealth 2013’ gold standard is presumed as non-relevant document (scale 0).
        </p>
        <p>
          We used random forest algorithm [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] as a learning to rank algorithm.
Evaluation metric to optimize on the training data was NDCG@10.
        </p>
        <p>We applied different parameter settings per each run. In its default setting,
minimum leaf support parameter is 1, and number of bags parameter is 300. In Run
Number 5, we set minimum leaf support parameter to 10. In Run Number 6, we set the
number of bags parameter to 50. In Run Number 7, we set the minimum leaf support
parameter to 50 (For French, German and Czech language queries in task 3b,
parameter setting is same as English query).
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion</title>
      <sec id="sec-3-1">
        <title>Evaluation results</title>
        <p>Experimental results are described in Table 1, Table 2, Table 3 and Table 4.
In Task3a (Table1), primary evaluation metric was P@10 and secondary evaluation
metric was NDCG@10. Our baseline P@10 was 0.738.</p>
        <p>SNUMEDINFO_EN_Run.2 (UMLS query expansion method) and
SNUMEDINFO_EN_Run.5 (Learning to rank method)’s performance is slightly
improved than SNUMEDINFO_EN_Run.1 (Baseline). The amount of performance gain
was not large enough: on average, 1~2% improved in terms of P@10 compared to the
baseline. If we compare SNUMEDINFO_EN_Run.2 to the baseline performance in
terms of P@10, 5 queries improved, 41 queries unaffected, 4 queries harmed. If we
compare SNUMEDINFO_EN_Run.5 to the baseline performance in terms of P@10,
14 queries improved, 32 queries unaffected, 4 queries harmed. But both of them failed
to attain significant improvement against baseline method when we performed
twotailed paired t-test. Nevertheless, we think that both of our methods; (1) lexical query
expansion using discharge summary as context filter; (2) learning to rank approach
utilizing medical web page quality features; showed good potentials of improvement
against tough baseline. We will try to improve our methods in our future research.</p>
        <p>Regarding Task3b (Table2, 3 and 4), we just used Google Translate to translate
French, German, and Czech query. These translated queries showed comparable
performance compared to the original English query.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In CLEFeHealth 2014’ Task 3, we tried to test various retrieval technique. We tried
lexical query expansion methods and learning to rank approach utilizing various
features for medical web document. Evaluation results shows potentially promising
results. We will try to improve our methods with more experiments in the future study.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This study was supported by a grant of the Korean Health Technology R&amp;D
Project, Ministry of Health &amp; Welfare, Republic of Korea. (No. HI11C1947).
5
ethnicity
etiology
evidence
follow up
followup
frequency
frequent
gender
guide
guideline
history
hospital
illness
introduction
journal
literature
management
md
medical doctor
medication
medicine
morbidity
mortality
ms
occurrence
op
operation
overview
pathology
pathophysiology
phd
physiology
prevent
prevention
problem
prognosis
proof
publication
publish
race
radiology
recommend
recommendation
research article
risk factor
sex
sign
statistic
study
sx
symptom
test
therapy
treat
treatment
tx
update
w/u
work up
workup</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Lorraine</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          , et al.
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2014</year>
          ,
          <article-title>Task 3: Usercentered health information retrieval</article-title>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>J. Lafferty</surname>
          </string-name>
          ,
          <article-title>A study of smoothing methods for language models applied to Ad Hoc information retrieval</article-title>
          ,
          <source>in Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          .
          <year>2001</year>
          , ACM: New Orleans, Louisiana, USA. p.
          <fpage>334</fpage>
          -
          <lpage>342</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Strohman</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , et al.
          <article-title>Indri: A language model-based search engine for complex queries</article-title>
          .
          <source>in Proceedings of the International Conference on Intelligent Analysis</source>
          .
          <year>2005</year>
          .
          <article-title>McLean, VA</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Google</given-names>
            <surname>Translate</surname>
          </string-name>
          .
          <year>2014</year>
          [cited 2014 June 12]; Available from: https://translate.google.com/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>F.-M. Lang</surname>
          </string-name>
          ,
          <article-title>An overview of MetaMap: historical perspective and recent advances</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          ,
          <year>2010</year>
          .
          <volume>17</volume>
          (
          <issue>3</issue>
          ): p.
          <fpage>229</fpage>
          -
          <lpage>236</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , et al.,
          <article-title>The PageRank Citation Ranking: Bringing Order to the Web</article-title>
          .
          <year>1999</year>
          , Stanford InfoLab.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Olteanu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.,
          <string-name>
            <surname>Web</surname>
            <given-names>Credibility</given-names>
          </string-name>
          :
          <article-title>Features Exploration and Credibility Prediction</article-title>
          , in Advances in Information Retrieval,
          <string-name>
            <given-names>P.</given-names>
            <surname>Serdyukov</surname>
          </string-name>
          , et al.,
          <source>Editors</source>
          .
          <year>2013</year>
          , Springer Berlin Heidelberg. p.
          <fpage>557</fpage>
          -
          <lpage>568</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rubin</surname>
            ,
            <given-names>V.L.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>E.D.</given-names>
            <surname>Liddy</surname>
          </string-name>
          .
          <article-title>Assessing Credibility of Weblogs</article-title>
          . in AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs.
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Agichtein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , et al.,
          <article-title>Finding high-quality content in social media</article-title>
          ,
          <source>in Proceedings of the 2008 International Conference on Web Search and Data Mining</source>
          .
          <year>2008</year>
          , ACM: Palo Alto, California, USA. p.
          <fpage>183</fpage>
          -
          <lpage>194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kleinberg</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <article-title>Authoritative sources in a hyperlinked environment</article-title>
          .
          <source>J. ACM</source>
          ,
          <year>1999</year>
          .
          <volume>46</volume>
          (
          <issue>5</issue>
          ): p.
          <fpage>604</fpage>
          -
          <lpage>632</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bendersky</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>W.B.</given-names>
            <surname>Croft</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Diao</surname>
          </string-name>
          ,
          <article-title>Quality-biased ranking of web documents</article-title>
          ,
          <source>in Proceedings of the fourth ACM international conference on Web search and data mining</source>
          .
          <year>2011</year>
          , ACM: Hong Kong, China. p.
          <fpage>95</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Aphinyanaphongs</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , et al.,
          <article-title>Text Categorization Models for High-Quality Article Retrieval in Internal Medicine</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          ,
          <year>2005</year>
          .
          <volume>12</volume>
          (
          <issue>2</issue>
          ): p.
          <fpage>207</fpage>
          -
          <lpage>216</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kilicoglu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , et al.,
          <source>Towards Automatic Recognition of Scientifically Rigorous Clinical Research Evidence. Journal of the American Medical Informatics Association</source>
          ,
          <year>2009</year>
          .
          <volume>16</volume>
          (
          <issue>1</issue>
          ): p.
          <fpage>25</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Abbasi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.,
          <article-title>Detecting Fake Medical Web Sites Using Recursive Trust Labeling</article-title>
          .
          <source>ACM Trans. Inf</source>
          . Syst.,
          <year>2012</year>
          .
          <volume>30</volume>
          (
          <issue>4</issue>
          ): p.
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Sondhi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>V.G.V.</given-names>
            <surname>Vydiswaran</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <article-title>Reliability Prediction of Webpages in the Medical Domain</article-title>
          , in Advances in Information Retrieval,
          <string-name>
            <given-names>R.</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          , et al.,
          <source>Editors</source>
          .
          <year>2012</year>
          , Springer Berlin Heidelberg. p.
          <fpage>219</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Aphinyanaphongs</surname>
            , Y. and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Aliferis</surname>
          </string-name>
          .
          <article-title>Text categorization models for identifying unproven cancer treatments on the web</article-title>
          .
          <source>in Medinfo 2007: Proceedings of the 12th World Congress on Health (Medical) Informatics; Building Sustainable Health Systems</source>
          .
          <year>2007</year>
          . IOS Press.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Richard</surname>
          </string-name>
          ,
          <article-title>Rule-based automatic criteria detection for assessing quality of online health information</article-title>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Price</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>W.R.</given-names>
            <surname>Hersh</surname>
          </string-name>
          .
          <article-title>Filtering Web pages for quality indicators: an empirical approach to finding high quality consumer health information on the World Wide Web</article-title>
          .
          <source>in Proceedings of the AMIA Symposium</source>
          .
          <year>1999</year>
          . American Medical Informatics Association.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.,
          <article-title>Combining relevancy and methodological quality into a single ranking for evidence-based medicine</article-title>
          .
          <source>Information Sciences</source>
          ,
          <year>2012</year>
          .
          <volume>214</volume>
          (
          <issue>0</issue>
          ): p.
          <fpage>76</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , et al.,
          <source>ShARe/CLEF eHealth Evaluation Lab</source>
          <year>2013</year>
          ,
          <article-title>Task 3: Information retrieval to address patients' questions when reading clinical reports</article-title>
          .
          <source>Online Working Notes of CLEF, CLEF</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Random</given-names>
            <surname>Forests</surname>
          </string-name>
          .
          <source>Machine Learning</source>
          ,
          <year>2001</year>
          .
          <volume>45</volume>
          (
          <issue>1</issue>
          ): p.
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Dang</surname>
          </string-name>
          , V. RankLib. [cited 2014 June 12]; Available from: http://sourceforge.net/p/lemur/wiki/RankLib/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>