<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Initial Results in the Development of SCAN</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A Swedish Clinical Abbreviation Normalizer</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Niklas Isenius</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sumithra Velupillai</string-name>
          <email>sumithra@dsv.su.se</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Kvist</string-name>
          <email>maria.kvist@karolinska.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Clinical Immunology &amp; Transfusion Medicine, Karolinska University Hospital</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Computer &amp; Systems Sciences, Stockholm University</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <abstract>
        <p>Abbreviations are common in clinical documentation, as this type of text is written under time-pressure and serves mostly for internal communication. This study attempts to apply and extend existing rule-based algorithms that have been developed for English and Swedish abbreviation detection, in order to create an abbreviation detection algorithm for Swedish clinical texts that can identify and suggest definitions for abbreviations and acronyms. This can be used as a pre-processing step for further information extraction and text mining models, as well as for readability solutions. Through a literature review, a number of heuristics were defined for automatic abbreviation detection. These were used in the construction of the Swedish Clinical Abbreviation Normalizer (SCAN). The heuristics were: a) freely available external resources: a dictionary of general Swedish, a dictionary of medical terms and a dictionary of known Swedish medical abbreviations, b) maximum word lengths (from three to eight characters), and c) heuristics for handling common patterns such as hyphenation. For each token in the text, the algorithm checks whether it is a known word in one of the lexicons, and whether it fulfills the criteria for word length and the created heuristics. The final algorithm was evaluated on a set of 300 Swedish clinical notes from an emergency department at the Karolinska University Hospital, Stockholm. These notes were annotated for abbreviations, a total of 2,050 tokens. This set was annotated by a physician accustomed to reading and writing medical records. The algorithm was tested in different variants, where the word lists were modified, heuristics adapted to characteristics found in the texts, and different combinations of word lengths. The best performing version of the algorithm achieved an F-Measure score of 79%, with 76% recall and 81% precision, which is a considerable improvement over the baseline where each token was only matched against the word lists (51% F-measure, 87% recall, 36% precision). Not surprisingly, precision results are higher when the maximum word length is set to the lowest (three), and recall results higher when it is set to the highest (eight). Algorithms for rule-based systems, mainly developed for English, can be successfully adapted for abbreviation detection in Swedish medical records. System performance relies heavily on the quality of the external resources, as well as on the created heuristics. In order to improve results, part-of-speech in-</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>formation and/or local context is needed for disambiguation. In the case of
Swedish, compounding also needs to be handled.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>Abbreviations are common in clinical documentation, as this type of text is written
under time-pressure and serves mostly for internal communication. For text mining
and information extraction techniques, as well as for readability for e.g. patients and
caregivers from other specialties, it would be beneficial to automatically identify and
expand abbreviations to their full-form counterparts. However, most research on
abbreviation detection and expansion has been performed on English texts. This study
attempts to apply and extend existing rule-based algorithms that have been developed
for English and Swedisha abbreviation detection,1-5 in order to create an abbreviation
detection algorithm for Swedish clinical texts that can identify and suggest definitions
for abbreviations and acronyms. This can be used as a pre-processing step for further
information extraction and text mining models, as well as for readability solutions.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Materials and Methods</title>
      <p>Through a literature review, a number of heuristics were defined for automatic
abbreviationb detection and were used for the construction of the Swedish Clinical
Abbreviation Normalizer (SCAN). These heuristics were:
1. freely available external resources:
─ a dictionary of general Swedish,c
─ a dictionary of medical terms,d and
─ a dictionary of known Swedish medical abbreviations,6
2. maximum word lengths (from three to eight characters), and
3. heuristics for handling common patterns such as hyphenation.</p>
      <p>For each token in the text, the algorithm checks whether it is a known word in one of
the lexicons, and whether it fulfils the criteria for word length and the created
heuristics. The final algorithm was evaluated on a set of 300 Swedish clinical notes from an
emergency department at the Karolinska University Hospital, Stockholm. e These
notes were annotated for abbreviations, resulting in a total of 2,050. This set was
annotated by a physician accustomed to reading and writing medical records.
a This study is applied on Swedish biomedical texts, not on clinical documentation.
b In the study presented here, acronyms are included in the definition of abbreviations.
c http://runeberg.org/words/ and http://g3.spraakdata.gu.se/saob/
d http://www.fass.se
e Ethical approval is granted by the Regional Ethical Review Board in Stockholm
(Etikprövningsnämnden I Stockholm), permission number 2009/1742-31/5</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>The algorithm was tested in different variants, where the word lists were modified,
heuristics adapted to characteristics found in the texts, and different combinations of
word lengths. The best performing version of the algorithm achieved an F-Measure
score of 79%, with 76% recall and 81% precision, which is a considerable
improvement over the baseline where each token was only matched against the word lists
(51% F-measure, 87% recall, 36% precision), see Table 1. Not surprisingly, precision
results are higher when the maximum word length is set to the lowest (three), and
recall results higher when it is set to the highest (eight).
In this study, a rule-based abbreviation system tailored for Swedish medical records is
presented. The system relies on lexicons and heuristics and the overall results are
encouraging. Through an error analysis, different types of common errors have been
identified, such as ambiguous words (e.g. hö which is a valid word (hay) but also a
common abbreviation for höger (right)) and abbreviations within compounds, e.g.
lungrtg (x-ray of lungs, x-ray abbreviated rtg). Compared to similar research for
English (e.g. Xu et al.1), results are lower. Relying on word lengths and external lexicons
limits coverage. In order to improve results, part-of-speech information and/or local
context is needed for disambiguation, for instance, which would probably improve
precision results. In the case of Swedish, compounding also needs to be handled.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>Algorithms for rule-based systems, mainly developed for English, can be successfully
adapted for abbreviation detection in Swedish medical records. System performance
relies heavily on the quality of the external resources, as well as on the created
heuristics. However, promising results are obtained without extensive tailoring of previous
algorithms developed for other languages. In the case of Swedish, language-specific
properties need to be addressed. Future work involves expanding the abbreviations to
their definitions, where emphasis should be put on improving precision rather than
recall, as precision would be more important for this task (an erroneous expansion
would create unfortunate misunderstandings). In the current system definitions are
provided if they exist in the list of known medical abbreviations but no evaluation has
yet been performed on this part. Detecting abbreviations automatically in medical
records has great importance for information access from this type of text. In one
study, extraction of disorders were hampered by abbreviations, as 14% of disorders in
clinical text were written as abbreviations and was not found when matched to
SNOMED CT7. If these could be correctly expanded to their full-length counterpart,
automatic disorder extraction would of course also improve. With rule-based
algorithms, very little language specific tailoring is needed, at least in the case of English
versus Swedish. However, for better performance, some issues need to be handled,
e.g. compound splitting and disambiguation.</p>
      <p>Acknowledgements
The authors wish to thank the anonymous and known reviewers for their comments
and feedback.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stetson</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <year>2007</year>
          .
          <article-title>A Study of Abbreviations in Clinical Notes</article-title>
          ,
          <source>in AMIA Annual Symposium Proceedings</source>
          <year>2007</year>
          , pp.
          <fpage>821</fpage>
          -
          <lpage>825</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Yeates</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>1999</year>
          .
          <article-title>Automatic Extraction of Acronyms from Text</article-title>
          ,
          <source>in Proceedings of the Third New Zealand Computer Science Research Students' Conference</source>
          , pp.
          <fpage>117</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Byrd</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <year>2001</year>
          .
          <article-title>Hybrid text mining for finding abbreviations and their definitions</article-title>
          ,
          <source>in Proceedings of Empirical Methods in Natural Language Processing</source>
          <year>2001</year>
          , pp.
          <fpage>126</fpage>
          -
          <lpage>133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Larkey</surname>
            ,
            <given-names>L.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogilvie</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Price</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamilio</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2000</year>
          .
          <source>Acrophile: An Automated Acronym Extractor and Server</source>
          ,
          <source>in Proceedings of the Fifth ACM Conference on Digital Libraries</source>
          , pp.
          <fpage>205</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dannélls</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <year>2003</year>
          .
          <string-name>
            <given-names>Acronym</given-names>
            <surname>Recognition</surname>
          </string-name>
          .
          <source>Master Thesis</source>
          . Department of Linguistics, Gö- teborg University, Sweden
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cederblom</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>2005</year>
          .
          <article-title>Medicinska förkortningar och akronymer</article-title>
          .
          <source>Lund: Studentlitteratur AB.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Skeppstedt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kvist</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalianis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Rule-based Entity Recognition and Coverage of SNOMED CT in Swedish Clinical Text</article-title>
          .
          <source>In Proc. LREC</source>
          <year>2012</year>
          , Istanbul, Turkey.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>