<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SNUMedinfo at CLEF QA track BioASQ 2015</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sungbin Choi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Biomedical Engineering, Seoul National University</institution>
          ,
          <addr-line>Seoul</addr-line>
          ,
          <country>Republic of Korea</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our participation at the BioASQ Task 3b of CLEF 2015 Question Answering track. We participated at the document retrieval subtask in Phase A and the ideal answer generation subtask in Phase B. As of previous year, in the document retrieval task, we mostly experimented with semantic concept-enriched dependence model and sequential dependence model. In the ideal answer generation task, relevant passages are selected and combined to automatically produce answer text.</p>
      </abstract>
      <kwd-group>
        <kwd>Information retrieval</kwd>
        <kwd>Semantic concept-enriched dependence model</kwd>
        <kwd>Sequential dependence model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        This paper describes the participation of the SNUMedinfo at the CLEF 2015 BioASQ
task 3b. We experimented with almost similar method as of our previous participation
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Task 3b was about biomedical semantic question answering task. For a detailed
introduction of the task, please see the overview paper of CLEF Question Answering track
BioASQ 2015’ [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <sec id="sec-2-1">
        <title>Task 3b Phase A – Document retrieval</title>
        <p>
          In Task 3b Phase A, we participated at the document retrieval subtask only. We used
Indri search engine [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The queries are stopped at the query time using the standard
418 INQUERY stopword list, case-folded, and stemmed using Porter stemmer. We
used unigram language model with Dirichlet prior smoothing [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] as our baseline
retrieval method (referred as QL: query likelihood model).
        </p>
        <p>
          We experimented with semantic concept-enriched dependence model (SCDM) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and
sequential dependence model (SDM) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. For a detailed description of our retrieval
method, please see our previous paper [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <sec id="sec-2-1-1">
          <title>Sequential dependence model (SDM)</title>
          <p>SDM Indri query example for the original query ‘What is the inheritance pattern of
Emery-Dreifuss muscular dystrophy?’ can be described as follows.
#weight (
λT</p>
          <p>#combine( inheritance pattern emery dreifuss muscular dystrophy )</p>
          <p>λ U #combine( #uw8(inheritance pattern) #uw8(pattern emery) #uw8(emery
dreifuss) #uw8(dreifuss muscular) #uw8(muscular dystrophy) ) )
λT, λO, λU are weight parameters for single terms, ordered phrases and unordered
phrases, respectively.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Semantic concept-enriched dependence model (SCDM)</title>
          <p>SCDM Indri query example can be described as follows.
 SCDM type C (single + multi-term, all-in-one)
#weight (
λT #combine( inheritance pattern emery dreifuss muscular dystrophy )</p>
          <p>λO_SC #combine( #od1(inheritance pattern) #od1(emery dreifuss muscular
dystrophy) )</p>
          <p>λU_SC #combine(#uw8(inheritance pattern) #uw16(emery dreifuss muscular
dystrophy) ) )
λT, λO, λU, λO_SC, λU_SC are weight parameters for single terms, ordered phrases and
unordered phrases of sequential query term pairs, ordered phrases and unordered
phrases of semantic concepts, respectively.
 SCDM type D (single+multi-term, pairwise)
#weight (
λT #combine( inheritance pattern emery dreifuss muscular dystrophy )</p>
          <p>λ U #combine( #uw8(inheritance pattern) #uw8(pattern emery) #uw8(emery
dreifuss) #uw8(dreifuss muscular) #uw8(muscular dystrophy) )</p>
          <p>λ O_SC #combine(#od1(inheritance pattern) #od1(emery dreifuss) #od1(dreifuss
muscular) #od1(muscular dystrophy) )</p>
          <p>λ U_SC #combine(#uw8(inheritance pattern) #uw8(emery dreifuss) #uw8(dreifuss
muscular) #uw8(muscular dystrophy) ) )</p>
          <p>We experimented with following parameter settings.</p>
          <p>SNUMedinfo1: SCDM Type C (mu=500, λT=0.85, λO=0.00, λU=0.00, λO_SC=0.10, λU_SC=0.05)
SNUMedinfo2: SCDM Type C (mu=500, λT=0.70, λO=0.00, λU=0.00, λO_SC=0.20, λU_SC=0.10)
SNUMedinfo3: SCDM Type C (mu=500, λT=0.70, λO=0.10, λU=0.05, λO_SC=0.10, λU_SC=0.05)
SNUMedinfo4: SCDM Type D (mu=500, λT=0.85, λO=0.00, λU=0.00, λO_SC=0.10, λU_SC=0.05)
SNUMedinfo5: SDM (mu=500, λT=0.85, λO=0.00, λU=0.00, λO_SC=0.10, λU_SC=0.05)
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Task 3b Phase B – Ideal answer generation</title>
        <p>In Task 3b Phase B, we participated only at the ideal answer generation subtask. We
reformulated this task as, among relevant lists of passages given1, selecting most
appropriate ones. We experimented with following heuristic method to select m passages
and combine them to form the ideal answer.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Identifying keyword terms and rank passages based on the number of unique keywords it contain</title>
        <p>Firstly, candidate passages are ranked based on number of keywords. Parameter minDF
represents minimum proportion of passages that keyword term should occur. If there
are 20 relevant passages given, and minDF is set to 0.5, then any terms occurring ≥ 10
passages are considered as keywords. With identified keywords list, we rank passages
based on the number of unique keywords each passage contains.</p>
        <p>Then, passages from top ranked ones are included for answer generation. Parameter
minUnseen represents minimum proportion of new tokens that does not exist in the
previously selected passages. We check proportion of tokens in the passage that does
not occur in the previously selected passages, and if it is ≥ minUnseen threshold,
second-ranked passage is selected. If proportions of newly found tokens are below
minUnseen threshold, that passage is abandoned, and we check next rank passage. This
process is repeated until m passage is selected. We intend to enhance
comprehensiveness of answer text by increasing the diversity of tokens.
1 We used gold relevant text snippets provided by the BioASQ.</p>
        <p>In this method, our intention was enhancing comprehensiveness of answer text by
increasing the diversity of tokens.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results &amp; Discussion</title>
      <p>At the moment of writing this paper, the final evaluation results are not available yet.
So we report tentative evaluation results currently available for us.
3.1</p>
      <sec id="sec-3-1">
        <title>Task 3b Phase A – Document retrieval</title>
        <p>There were five distinct batches within this task.
Generally, SDM and SCDM showed better performance compared to the baseline QL
method. But compared to the previous year, limit of returned document per query is
decreased from 100 to 10. We presume that the evaluation scores become more volatile
because of that.
We submitted five runs trying different parameter values, but according to the
automatic evaluation score (Rouge-2 and Rouge-SU4) evaluation, performance change
seems not very meaningful.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <article-title>Classification and retrieval of biomedical literatures: Snumedinfo at clef qa track bioasq 2014</article-title>
          .
          <source>Proceedings of Question Answering Lab at CLEF</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and San Juan, E.,
          <source>CLEF 2015 Labs and Workshops. 2015, CEUR Workshop Proceedings (CEUR-WS.org).</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Strohman</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , et al.
          <article-title>Indri: A language model-based search engine for complex queries</article-title>
          .
          <source>in Proceedings of the International Conference on Intelligent Analysis</source>
          .
          <year>2005</year>
          .
          <article-title>McLean, VA</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>J. Lafferty</surname>
          </string-name>
          ,
          <article-title>A study of smoothing methods for language models applied to Ad Hoc information retrieval</article-title>
          ,
          <source>in Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          .
          <year>2001</year>
          , ACM: New Orleans, Louisiana, USA. p.
          <fpage>334</fpage>
          -
          <lpage>342</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.,
          <article-title>Semantic concept-enriched dependence model for medical information retrieval</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          .
          <volume>47</volume>
          : p.
          <fpage>18</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Metzler</surname>
            , D. and
            <given-names>W.B.</given-names>
          </string-name>
          <string-name>
            <surname>Croft</surname>
          </string-name>
          ,
          <article-title>A Markov random field model for term dependencies</article-title>
          ,
          <source>in Proceedings of the 28th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          .
          <year>2005</year>
          , ACM: Salvador, Brazil. p.
          <fpage>472</fpage>
          -
          <lpage>479</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>