<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Improving Search Results for Medical Experts and Laypersons</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Karin Friberg Heppin</string-name>
          <email>karin.friberg.heppin@svenska.gu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anni Järvelin</string-name>
          <email>anni.jarvelin@svenska.gu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Gothenburg, Språkbanken, Department of Swedish</institution>
          ,
          <addr-line>Box 200, 405 30 Gothenburg</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In a domain such as medicine, it is important that individuals' information needs are met with information on a suitable level of difficulty and expertise. This paper focuses on facilitating medical information access through reformulating queries and re-ranking result lists utilizing features typical for the language written for professionals or for laypersons. The aim is to produce result lists where the ranking is better suited for the expertise level of the user. We will explore the possibility of using features such as trigger phrases for query reformulation and document length, average word length or compound ratio for re-ranking. The Swedish medical IR test collection, MedEval, from Språkbanken, University of Gothenburg, will be used to find features specific for professional language and lay language and to study the effectiveness of these features in reformulating queries and re-ranking search results based on the target group. The test collection contains 42,250 documents from the medical corpus MedLex , collected from all types of written medical information found in electronic format, except patient records. The collection contains 62 topics. In total, 7,044 documents have been assessed both for relevance to these topics and for target group. Our experiments will be based on earlier explorative studies on medical expert and lay language where some features were identified. It was found that documents written for professionals tended to have more tokens per document, longer words, and more compounds than lay documents The assessed documents were run through a perl program which counted the frequencies of occurring multiword units (MWUs). The frequencies of individual MWUs were much higher in the expert documents. For the doctors, medical phrases dominate the MWUs while the patients' documents mostly contained general language units. The most frequent patient MWUs from the medical domain, were more frequent in the doctor documents and could therefore not be said to be typical for patient documents. Many frequently recurring MWUs are not specific for any topic. However, they may be seen as trigger phrases indicating target group. Such trigger phrases will be used in the reformulation of the queries. We believe that the phrases typical for a user group can be used to reformulate queries and that the likewise typical features can be useful for calculating Cross-Language Evaluation of Methods, Applications, and Resources for eHealth Document Analysis CLEFeHealth 2012 Workshop, edited by Hanna Suominen © CLEF 2012</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>target group scores for medical documents and further for improving the
ranking of the documents to better match the expertise of the user.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>In a domain such as medicine, it is important that individuals’ information needs are
met with information on a suitable level of difficulty and expertise. Laypersons and
experts might need information, about the same topics, but not in the same form.
Search engines do a fairly good job of finding documents to satisfy the range of
information needs people may have for a query, but do less well in discerning
individuals’ search goals.1</p>
      <p>This paper focuses on the need to facilitate medical information access through
reformulating queries and re-ranking result lists utilizing features typical for the
language of the documents for professionals or for laypersons. The aim is to produce
result lists where the ranking is better suited for the expertise level of the user. We
will explore the possibility of using features such as trigger phrases for query
reformulation and document length, average word length or compound ratio for re-ranking.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Materials and Methods</title>
      <p>The Swedish medical IR test collection, MedEval, from Språkbanken, University of
Gothenburg will be used to find features specific for professional language and lay
language and to study the effectiveness of these features in reformulating queries and
re-ranking search results based on the target group.2</p>
      <p>The test collection contains 42,250 documents from the medical corpus MedLex,3
collected from scientific journals, health care communicators, newspapers, patient
FAQs etc., in short, all types of medical information found in electronic format,
except patient records. The collection contains 62 topics. In total, 7,044 documents have
been assessed both for relevance to these topics and for target group: medical
professionals or laypersons.4 3,272 documents were assessed to have medical professionals
as target group; 4,334 documents to have laypersons as target group; and 562 of these
were assigned to both categories. For a document to be assigned to both categories it
must have appeared in the pool of more than one topic and further have been assigned
different target groups for the different topics.</p>
      <p>In addition to the assessed documents, the target group of the documents can in
many cases be reliably identified from the document source: scientific journals are
written for medical professionals, while health care counselling web sites and
discussion forums are typically for laypersons. Thus more documents can be added to the
sets of documents used for extracting the target group features.</p>
      <p>To test the reliability and effectiveness of target group features in re-ranking results
for the correct target group, a set of 20 topics from the MedEval test collection will be
used. The effect of the re-ranking based on the target group features will be tested for
both target groups using both medical and common vocabulary in the queries.
3</p>
    </sec>
    <sec id="sec-4">
      <title>Preliminary Results and Discussion</title>
      <p>Our experiments will be based on the explorative studies on medical expert and lay
language described in Friberg Heppin4 where some promising features were
identified. It was found that documents written for professionals tended to have more
tokens per document, longer words, and more compounds than lay documents (see
Table 1).</p>
      <p>The assessed documents were run separately through MWT (multi-word term), a perl
program which counts the frequencies of occurring multiword units using both lexical
and statistical calculation.5 The version used here was later enhanced and provided to
us by Jussi Karlgren (personal communication).</p>
      <p>There were differences in both types and frequencies of multiword units (MWUs)
found in the documents for the two target groups. The frequencies of individual
MWUs were much higher in the doctor documents. For the experts, medical phrases
dominate the MWUs while the patients’ documents contain general language units.
The most frequent patient MWU which could be said to be medical, att drabbas av
‘to be afflicted by’, has no more than 95 occurrences in the patient documents while it
has 255 occurrences in the doctor documents. Another of the most frequent patient
phrases när det gäller ‘when it comes to’, is also more common in the doctor
documents, 574 for doctors and 102 for patients. Thus these cannot be said to be MWUs
typical for the patient documents.</p>
      <p>Many frequently recurring MWUs are not specific for any topic. However, some of
them may be seen as trigger phrases indicating target group. Potential trigger phrases
for doctors and patients can be seen in Table 2 and 3. Such trigger phrases will be
used in the reformulation of the queries, for example by adding or reducing ranking
scores of documents where the phrases are present.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>Our preliminary studies have identified a few target group specific features and
trigger phrases in professional and lay medical documents. These features and phrases
were extracted from a relatively small number of documents and could be improved
by using a larger number of documents. We believe that these phrases can be used to
reformulate queries and that these features can be useful for calculating target group
scores for medical documents and further for improving the ranking of the documents
to better match the expertise level of the user.</p>
      <p>Acknowledgements
We would like to thank the Centre of Language Technology at the University of Gothenburg
for providing funding for this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Jaime</given-names>
            <surname>Teevan</surname>
          </string-name>
          , Susan T. Dumais, and
          <string-name>
            <given-names>Eric</given-names>
            <surname>Horvitz</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Potential for personalization</article-title>
          .
          <source>ACM Trans. Comput</source>
          .-Hum. Interact.,
          <volume>17</volume>
          (
          <issue>1</issue>
          ):4:
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          :
          <fpage>31</fpage>
          ,
          <string-name>
            <surname>April</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Karin</given-names>
            <surname>Friberg Heppin</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>MedEval - A Swedish Medical Test Collection with Doctors and Patients User Groups</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          ,
          <volume>2</volume>
          (
          <issue>Suppl 3</issue>
          ):
          <fpage>S4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Dimitrios</given-names>
            <surname>Kokkinakis</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>MEDLEX: Technical report</article-title>
          .
          <source>Technical report</source>
          , Department of Swedish, University of Gothenburg.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Karin</given-names>
            <surname>Friberg Heppin</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Resolving Power of Search Keys in MedEval a Swedish Medical Test Collection with User Groups: Doctors and Patients</article-title>
          .
          <source>Ph.D. thesis</source>
          , University of Gothenburg
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. John S. Justeson and
          <string-name>
            <surname>Slava M John S. Justeson and Slava M. Katz</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Technical terminology: Some linguistic properties and an algorithm for indentification in text</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>9</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>