<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Preventing Adverse Drug Events by Extracting Information from Drug Fact Sheets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefania Rubrichi</string-name>
          <email>stefania.rubrichi@unipv.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alex Spengler</string-name>
          <email>alex.spengler@lip6.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Gallinari</string-name>
          <email>patrick.gallinari@lip6.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvana Quaglini</string-name>
          <email>silvana.quaglini@unipv.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Laboratoire d'Informatique de Paris 6, Universite Pierre et Marie Curie</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Laboratory for Biomedical Informatics, Department of Computers and Systems Science, University of Pavia</institution>
          ,
          <addr-line>Pavia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Background: The increasing volume and growing complexity of drugs lead to an increased risk of prescription errors and adverse events. A correct drug choice must be modulated to acknowledge both patients' status and drug-speci c information. This information is reported in free-text on drug fact sheets. It is often overwhelming and di cult to access. There is thus a rising need for generating comprehensive and structured data that help prevent such events by improving access to fact sheet information. This work presents a machine learning based system for the automatic prediction of drug-related entities (active ingredient, interaction e ects, etc.) in textual drug fact sheets, focusing on drug interactions. Results: Our approach learns to classify this information in the structured prediction framework, comparing conditional random elds and support vector machines. Both classi ers are trained and evaluated using a corpus of 100 drug fact sheets. They have been hand-annotated with fourteen semantic labels that have been derived from a previously developed domain ontology. Our experimental results show that the two models exhibit similar overall performance. They achieve an average F1-measure of about 93 per cent, which is promising. The performance results of both models on the individual labels are also comparably good. Conclusions: We have shown that it is possible to perform the task of information extraction from drug fact sheets using supervised machine learning techniques. Although we have focused on drug interactions, the encouraging results and the adaptability of the approach we adopted means that our system has general signi cance for the extraction of detailed information on drugs (drug targets, contraindications, side e ects, etc.).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background</title>
      <p>
        The medication management process is highly
complex and involves a large number of choices from
different health care professionals. Medication errors
occur frequently among patients, at any point in the
medication administration process. The Institute of
Medicine (IOM) reports that more than 1.5 million
adverse drug events (ADEs) are preventable each
year in the US alone [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Examples of errors
include a patient receiving the wrong medication, a
medication to which they have a known allergy, or
a patient receiving an incorrect dose of medicine.
This phenomenon is aggravated by aging patients'
multi-pathologies and the ever-growing number and
complexity of drugs (e.g. drugs combining more than
one active ingredient deserve more attention for
interactions and contraindications). Physicians need
to take into account many drug-speci c and
patientspeci c characteristics and studies show that these
factors are often overlooked or recognized too late.
      </p>
      <p>
        In recent years, considerable e orts have been
made to reduce medication errors and to detect
and prevent ADEs. Computerized systems that
incorporate speci c applications in electronic medical
records or in clinical information systems support
medication ordering, dispensing and administration
functions. These systems are referred to as
Computerized Provider Order Entry { CPOE [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. To
enhance their performance, such systems have to
include rich domain knowledge. Thus, they will be
able to support clinical decision, assisting the
physician, for example, by screening orders for allergies
and drug-drug or drug-laboratory tests interactions,
generating alerts tailored to patients' characteristics.
Not much of such knowledge is available in
semistructured form, and even less in normalized,
structured form. In particular, drug-related information
is reported in free-text on fact sheet. For e
ective use, this information locked in natural language
must rst be transformed into structured data.
      </p>
      <p>In this work, we consider the problem of
automatic extraction of drug information conveyed in the
Summary of Product Characteristic (SPC), focusing
on a speci c section concerned with drug-related
interactions. Our contributions are:
1. We formulate the problem in a machine
learning framework, in which we seek to assign the
correct semantic label, such as
InteractionEffect or ActiveDrugIngredient, to each word on
a drug fact sheet. To this end, we employ two
state-of-the-art classi ers: linear-chain
conditional random elds (CRFs) and structured
support vector machines (SVMs). These
classi ers discriminate the semantic labels trough
the automatic adaptation of hundreds of
engineered text features, taking into account both
local (on a word level) and global (sentence or
fact sheet level) information.
2. We introduce a corpus of 100 interaction
sections in Italian language that have been
annotated with fourteen semantic labels, with
respect to a previously implemented ontology.
3. We apply both CRF and SVM to our data set
and evaluate their overall and individual label
performances. Both classi ers achieve an
average F1-measures of about 93%|a promising
result with regard to real-world applications.</p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <sec id="sec-2-1">
        <title>Named Entity Recognition in SPCs</title>
        <p>SPCs represent a source of information for health
professionals on how to use medicines safely and
effectively. It forms an intrinsic and integral part of
the marketing authorisation. In order to obtain an
authorization to place a medicinal product on the
market, a SPC shall be included in the application
made to the competent authority. Its content is
regulated by Article 11 of Directive 2001/83/EC. A
SPC sets out the agreed position (results of
physicochemical, biological or microbiological tests,
toxicological and pharmacological tests, clinical trials etc.)
on the medicinal product as collected during the
course of the assessment process.</p>
        <p>SPCs of specialty medicines for human use are
organized into 12 sections: name, therapeutic
categories, active ingredient, excipients, indications,
contraindication/side e ects, undesired e ects,
posology, storage precautions, warnings, interactions, use
in case of pregnancy and nursing. Access to this
comprehensive information provides a wide range of
coded data, which are then available for new or
improved clinical applications, facilitating and
improving the prescription process. It is therefore an
important step in preventing medical errors.</p>
        <p>We propose a rst approach to extracting
drugrelated interaction information reported as
freetext in SPCs, following a named entity
recognition (NER) approach. NER is an important step
of an integral information extraction task and aims
at identifying words or phrases in natural language
text that belong to certain classes of interest, and
labeling them according to their type. As an
example, consider the following sentence (translated from
Italian):
hEnoxapariniActiveDrugIngredient dosed as
a h1.0 mg/kgiPosology hsubcutaneous
injectioniIntakeRoute for hfour
dosesiPosology hdid not alter the
pharmacokineticsiInteractionE ect of
hepti batideiActiveDrugIngredient.</p>
        <p>In NER, each token is sought to be associated with a
label that indicates its appropriate domain-speci c
category.</p>
        <p>
          Typically, the rst step in most NER tasks is to
identify the named entities (labels) that are relevant
to the concepts, relations and events described in
the text. A system for NER is thus based on
speci c knowledge on the domain. Thus, as part of an
understanding of the factual information process, we
previously developed a domain ontology de ning the
entity classes, relations and attributes [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Based on
the ontology, an extensive knowledge base of
concepts is maintained.
        </p>
        <p>One of the most successful methods for
performing such labeling and segmentation tasks is that of
employing supervised machine learning techniques.
These methods automatically tune their own
parameters to maximize their performance on a set of
example texts that have been annotated by hand. The
machine then generalizes from these examples.</p>
        <p>
          We developed a framework for simultaneously
recognizing occurrences of multiple entity classes
using linear-chain CRFs [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and structured SVMs [
          <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
          ].
Both supervised machine learning approaches
predict words' labels using a large number of
descriptive characteristics (features) of the input by
assigning real-valued weight to these features. They can
be seen as a way to \capture" the hidden patterns
in both labels and features, and \learn" what would
be the likely output considering these patterns. Due
to paucity of space, however, we limit the treatment
of these subjects to a presentation of the employed
features and refer the reader to the original
publications for details on the statistical properties.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Features</title>
        <p>The feature construction process aims at capturing
the salient characteristics of each token in order to
help the system predict its semantic label. Since
statistical models such as the CRF and the SVM
crucially depend on a wise choice of these features,
their de ntion has critical impact on the overall
performance of the system.</p>
        <p>De ning features means to construct a set of
generally binary-valued feature functions f (x; t) for a
sentence x and a word position t. For example,
fenoxaparin(x; t) =
8 1 : if the word at position t
&lt;</p>
        <p>in x is enoxaparin
: 0 : otherwise
is a binary word feature which returns 1 whenever
the word at position t in sentence x is enoxaparin
and 0 otherwise.</p>
        <p>We implemented and employed a large variety of
informative features that can be derived from the
fact sheets { both locally on a token or word level
and globally on a sentence or section level.</p>
        <p>
          Before the actual feature assignment process, we
split each input sentence into tokens. We use a
simple, but robust tokenization method which considers
white-space, colon and parenthesis as token
boundaries. We then remove all punctuation with the
exception of hyphens occuring between alphanumeric
strings in a second preprocessing step. Due to some
length constraints our original database is not
properly hyphenated. To remedy this we consulted an
Italian language lexicon [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We now describe the
features we used in our experiments.
        </p>
        <sec id="sec-2-2-1">
          <title>Word and Neighbouring Word Features</title>
          <p>Each word in the stream of tokens has been
converted into a binary feature. Moreover, we equally
created features for the words preceding or
following the current position t in a sentence x,
modeling local context. Consider the excerpt from
page 2, for instance. Apart from fenoxaparin(x; t)
which is 1 only for t = 1, we also have a feature
fenoxaparin, -1(x; t) which is 1 whenever the
preceding word is enoxaparin, here for t = 2. In our
experiments, we report results for context sizes 0 (no
context), 3=3 and 7=7.</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Orthographical Features</title>
          <p>Besides word features, we added orthographical
features that indicate whether a token consists of digits.
This is useful for identifying Posology entities.</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>Word Substring Features</title>
          <p>Some substrings can provide good clues for
classifying named entities. In particular, we identi ed a set
of words which occur frequently with the same label;
for example Italian words which start with \e et-"
(e ect) are usually Interaction E ects, those
starting with \mg-" (mg) have usually been tagged as
Posology, and so on.</p>
        </sec>
        <sec id="sec-2-2-4">
          <title>Punctuation Features</title>
          <p>Also notable are features which characterize
interesting punctuation in sentences. After browsing our
corpora we found that colons and brackets may be
helpful. Given a medication, colons are usually
preceded by the interacting substance and followed by
the explanation of the speci c interaction e ects.
Round brackets show extra information. For each
token, the punctuation features test if it is preceded
or followed by a colon or a parenthesis. All
punctuation features have been used in conjunction with a
context window of sentence length.</p>
        </sec>
        <sec id="sec-2-2-5">
          <title>Active Ingredient Dictionary Feature</title>
          <p>
            Finally, we added an examplary feature carrying
domain-speci c knowledge. Farmadati Italia [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]
database provides a complete archive of active
ingredients. The dictionary feature is based on these
database entries and tells us whether a token is an
active ingredient or not.
          </p>
          <p>Model
CRF
SVM</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Experiments</title>
        <sec id="sec-2-3-1">
          <title>Data Collection</title>
          <p>
            The goal of this work lies in extracting information
from SPCs, with a focus on drug-related
interactions. We created a corpus which consists of 100
manually annotated interaction sections of specialty
medicines for human use. They have been extracted
from SPCs selected uniformly at random from the
Farmadati Italia database. We used the BDF
(Bancadati Federfarma) software [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] for an exploratory
data analysis and for exporting the SPCs to a text
le.
          </p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Ontology-based Annotation Process</title>
          <p>
            Semantic annotation is used to establish links
between the tokens in the SPCs and their semantic
descriptions or concept classes. For a reliable
annotation, the semantic descriptions must be well
dened and easy to understand by the domain expert
who annotates the text. We therefore annotated the
text with respect to a previously developed
ontologybased model of drug information as conveyed in the
SPCs [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ], which speci es the classes of concept (i.e.
concepts representing drug characteristics), the
relationships that bind them and other distinctions that
are relevant for modeling the application area. The
annotation process was performed by a biomedical
engineer with domain knowledge. A review of the
data has been used to validate and, when necessary,
correct the annotations.
          </p>
          <p>Leveraging the established ontology, we mapped
its elements to the SPCs' text content. We
accurately inspected all the corpus lines distinguishing
the di erent senses with respect to the ontology; we
then annotated each word in the extracted SPC
interaction sections with the corresponding class in the
ontology. Active ingredients have also been
automatically extracted using the Farmadati database.</p>
        </sec>
        <sec id="sec-2-3-3">
          <title>Experimetal Setup</title>
          <p>We randomly split the 100 interaction sections into
two sets; one for training which consists of 60
sections and one for testing which contains 40 sections.
In total, there are 840 input sentences for training
and 413 input sentences for testing.</p>
          <p>
            We measure and evaluate the performance of our
models based on precision (P), recall (R) and
F1measure (F1) [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]. We report results for the two
classi ers in terms of overall and individual label
performance. When dealing with multi-label
classication and imbalanced labels, the performance on
the individual labels can essentially be aggregated
into overall performance results in two
complementary ways: either we compute their arithmetic mean,
giving equal weight to each of the labels
(macroaveraged); or we compute the mean by weighting
each label by the number of times they occur in the
data set (micro-averaged).
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion</title>
      <p>Table 1 presents a summary of key performance
gures for both CRF and SVM. Overally, our
experiments show that the two classi ers, with carefully
designed features, can identify information related
to drug interactions with very high accuracy (about
93%). There is no clear superiority of one model
over the other. Although the data might contain
noise inherent to manual annotation, the learning
algorithms reach high performance. This high
recognition performance can be attributed to the latent
structural regularities of natural language text and
the regularity of the appearance of groups of named
entities in the investigated paragraphs. Approaching
the problem of information extraction from SPCs in
the described machine learning approach is
promising.</p>
      <p>Table 2 shows the performance of CRF and SVM
on the individual labels, employing all available
features and a context of size -7/7. The labels
OtherSubstance and DiagnosticTest are most di cult to
extract, which is probably due to the tiny number
of examples available. Rare labels as AgeClass,
IntakeRoute exhibit better performance, they may in
fact pro t from a precise de nition, contributing to
high performance.</p>
      <p>Moreover, we investigate the in uence of the
word neighbouring features with regard to overall
performance. Using local word context appears to
be useful for determining the semantic labels. The
larger the context window size, the better and the
more precise the results. Table 3 illustrates the
performance of both classi ers for varying context
window sizes. We followed an additive strategy: starting
with no word neighbouring features (i.e. a window
size of 0), we increased the window size piecemeal,
measuring the performance of the resulting
classiers at each step. The initial classi ers didn't use
a word neighbouring feature set. The addition of
neighbouring words in the window -3/3 as features
improves the F1-measure by about 3 4%
(microaveraged), and 7 9% (macro-averaged).
Incrementing the context window to a size of -7/7 gives rise
to an improvement on all metrics, boosting
microaveraged F1 by about 6% and the macro-averaged
F1 by about 7%. Non-zero context window sizes
hence provide an important bene t with respect to
the overall classi cation performance. An analysis
of the di erent performance increments for contexts
of size -3/3 and -7/7 will be left to future work.</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>We have presented a framework for simultaneously
recognizing occurrences of multiple entity classes in
textual drug fact sheets, using supervised machine
learning techniques. We compared the performance
of two state-of-the-art discriminative classi ers with
carefully engineered features. Our empirical
evaluation shows that the two classi ers exhibit similar
overall performance, achieving high overall accuracy.
Although we have focused on drug interactions, the
encouraging results and the adaptability of adopted
approach show that our system is signi cant for the
extraction of detailed information about drugs (drug
targets, contraindications, side e ects).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Institute of Medicine (Ed):
          <source>Preventing Medication Errors. Washington: The National Academics Press</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Sittig</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stead</surname>
            <given-names>W</given-names>
          </string-name>
          :
          <article-title>Computer-based physician order entry: the state of the art</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Rubrichi</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leonardi</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quaglini</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>A Drug Ontology as a basis for safe therapeutic decision</article-title>
          .
          <source>In Proceedings of the Second National Conference of Bioengineering</source>
          ,
          <year>Patron 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>La erty</surname>
            <given-names>JD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCallum</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            <given-names>FCN</given-names>
          </string-name>
          :
          <article-title>Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data</article-title>
          .
          <source>In Proceedings of the International Conference on Machine Learning</source>
          , Volume
          <volume>18</volume>
          2001:
          <volume>282</volume>
          {
          <fpage>289</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Tsochantaridis</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joachims</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hofmann</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altun</surname>
            <given-names>Y</given-names>
          </string-name>
          :
          <article-title>Large Margin Methods for Structured and Interdependent Output Variables</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <year>2005</year>
          ,
          <volume>6</volume>
          :
          <fpage>1453</fpage>
          {
          <fpage>1484</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bordes</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Usunier</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bottou</surname>
            <given-names>L</given-names>
          </string-name>
          :
          <article-title>Sequence Labelling SVMs Trained in One Pass</article-title>
          .
          <source>In ECML PKDD 2008</source>
          ,
          <year>Springer 2008</year>
          :
          <volume>146</volume>
          {
          <fpage>161</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Zanchetta</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baroni</surname>
            <given-names>M</given-names>
          </string-name>
          :
          <article-title>Morph-it!: a Free CorpusBased Morphological Resource for the Italian Language</article-title>
          .
          <source>In Proceedings of the Corpus Linguistics</source>
          <year>2005</year>
          conference
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. http://www.farmadati.it/.</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. http://www.farmadati.it/navigate.aspx?id=
          <fpage>17</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Van Rijsbergen</surname>
            <given-names>CJ</given-names>
          </string-name>
          :
          <article-title>Information Retrieval</article-title>
          . Department of Computer Science, University of Glasgow 1979.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>