<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using the Micropublications ontology and the Open Annotation Data Model to represent evidence within a drug-drug interaction knowledge base</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jodi Schneider</string-name>
          <email>jodi.schneider@inria.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Ciccarese</string-name>
          <email>paolo.ciccarese@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tim Clark</string-name>
          <email>clark@harvard.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Richard D. Boyce</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <country>INRIA Sophia Antipolis France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Massachusetts General Hospital and Harvard Medical School</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Pittsburgh</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Semantic web technologies can support the rapid and transparent validation of scienti c claims by interconnecting the assumptions and evidence used to support or challenge assertions. One important application domain is medication safety, where more e cient acquisition, representation, and synthesis of evidence about potential drug-drug interactions is needed. Potential drug-drug interactions (PDDIs), de ned as two or more drugs for which an interaction is known to be possible, are a signi cant source of preventable drug-related harm. The combination of poor quality evidence on PDDIs, and a general lack of PDDI knowledge by prescribers, results in many thousands of preventable medication errors each year. While many sources of PDDI evidence exist to help improve prescriber knowledge, they are not concordant in their coverage, accuracy, and agreement. The goal of this project is to research and develop core components of a new model that supports more e cient acquisition, representation, and synthesis of evidence about potential drug-drug interactions. Two Semantic Web models|the Micropublications Ontology and the Open Annotation Data Model|have great potential to provide linkages from PDDI assertions to their supporting evidence: statements in source documents that mention data, materials, and methods. In this paper, we describe the context and goals of our work, propose competency questions for a dynamic PDDI evidence base, outline our new knowledge representation model for PDDIs, and discuss the challenges and potential of our approach.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>drug-drug interactions</kwd>
        <kwd>evidence bases</kwd>
        <kwd>Micropublications</kwd>
        <kwd>Open Annotation Data Model</kwd>
        <kwd>knowledge bases</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>tinually growing and changing, as new scienti c studies are completed and new
documents are published. The state of current knowledge in any given domain
can be di cult for any one individual to fully grasp, because bits of knowledge
are updated at frequent intervals.</p>
      <p>In the biosciences, this problem has taken on particular importance, due
to an exponential growth in the aggregate publication rate. Manually curated
databases are used to record certain types of knowledge. To update and maintain
these databases, curators must make knowledge-intensive decisions, identifying
the best available evidence in the current scienti c literature. Maintaining such
databases is challenging because there is limited tracking of the source
information.</p>
      <p>In an ongoing project, we are experimenting with using the Micropublications
Ontology4 [Clark2014] and the Open Annotation Data Model5 [W3C2013] to
create an audit trail between assertions, evidence, and source documents, so
that assertions and evidence can be agged for update in exible and intelligent
ways. Updates may be needed when the underlying sources change, when a
particular method for establishing an assertion is discredited, etc. Our goal is
to provide better linkages between an assertion recorded in a knowledge base
and its supporting evidence (i.e., data, materials, and methods) found in source
documents.</p>
      <p>In the remainder of the paper, we describe the competency questions for
our evidence base and the new evidence model that we are creating, which
combines the Micropublication Ontology and the Open Annotation Data Model,
and adapts them to the existing evidence modeling of the Drug Interaction
Knowledge Base6 [Boyce2007,Boyce2009]. We then re ect on how the new model
performs for our goal of creating an audit trail between assertions, evidence, and
source documents.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Context and goals</title>
      <p>Our work is in the context of a larger project on organizing and synthesizing
scienti c evidence from the biomedical literature on potential drug-drug
interactions. Potential drug-drug interactions (PDDIs), de ned as two or more drugs
for which an interaction is known to be possible, are a signi cant source of
preventable drug-related harm (i.e., adverse drug events, or ADEs). The
combination of poor quality evidence on PDDIs, and a general lack of PDDI
knowledge by prescribers, results in many thousands of preventable medication
errors each year. While many sources of PDDI evidence exist to help improve
prescriber knowledge, they are not concordant in their coverage [Saverno2011],
accuracy [Wang2010], and agreement [Abarca2003]. Di culties with
synthesizing evidence, and gaps in the scienti c knowledge of PDDI clinical relevance,
underlie such disagreement.</p>
      <sec id="sec-2-1">
        <title>4 http://purl.org/mp/ 5 http://www.openannotation.org/spec/core/ 6 http://purl.net/net/drug-interaction-knowledge-base/</title>
        <p>To address these problems, our research group is studying the potential
benet of applying recent developments from the Semantic Web community on
scienti c discourse modeling and open annotation. The goal is to develop core
components of a new PDDI knowledge representation model that will support a more
e cient acquisition, representation, and synthesis of PDDI evidence. The desired
knowledge representation will provide better linkages between PDDI assertions
and their supporting evidence, by directly connecting to annotated section(s) of
relevant source documents.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <p>Our new approach will draw upon the current version (1.2) of the Drug
Interaction Knowledge Base [Boyce2007,Boyce2009], the Open Annotation Data
Model [W3C2013], and the Micropublications Ontology [Clark2014].</p>
      <p>The Drug Interaction Knowledge Base (DIKB) is a static, manually
constructed evidence base that indexes assertions and evidence of PDDI for over 60
drugs. Its taxonomy of assertion types and evidence types [Boyce2014] is a
starting point for the new knowledge base. The current version of the DIKB
implements a version of the SWAN semantic discourse ontology [Ciccarese2008] to
represent evidence relations. Speci cally, the knowledge base uses
swanco:citesAsSupportingEvidence and swanco:citesAsRefutingEvidence to link to an entire
source document as a supporting or refuting citation. At the time the DIKB
1.2 was constructed (2007{2009), annotation methodologies were less well
developed. Consequently, version 1.2 of the DIKB stores quotes as textual strings
manually copied from source documents. The text has been enriched with
metadata about the source section, but it is non-trivial to return to the appropriate
segment of the text from this information.</p>
      <p>Our use of the Open Annotation Data Model (OA) re ects a change in the
state of the art. OA is an \an interoperable framework for creating associations
between related resources, annotations, using a methodology that conforms to
the Architecture of the World Wide Web"7. In particular, OA allows an evidence
database to provide explicit connections from quotes to their source documents.
For example, as shown in Figure 1, an OA resource can be used to quote a speci c
part of a drug product label (also known as a summary of product characteristics)
to indicate evidence that escitalopram inhibits CYP2D6. In general, OA enables
queryable links between selections from source documents (as target) to the
instances of data, methods, and materials (as body) that we want to model to
support drug interaction knowledge base use cases.</p>
      <p>Similarly, the Micropublications Ontology improves the depth with which
evidence can be represented and queried. The most important feature of the
Micropublications model, in our view, is its ability to represent the data, methods,
and materials that act as support for a claim, and to transitively close chains</p>
      <sec id="sec-3-1">
        <title>7 http://www.openannotation.org/spec/core/</title>
        <p>of claims8 and citations across the literature to their fundamental supporting
evidence. A mp:Micropublication mp:argues a mp:Claim based on connecting
any number of mp:Representations. The whole Micropublication is a
Representation, as are Data and Methods (including Materials and Procedures), whether
textual or pictoral. A mp:Representation may mp:support or mp:challenge any
other mp:Representation, making the evidence explicit and queryable.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Competency Questions</title>
      <p>To design an appropriate enhancement of the DIKB model with
Micropublications and the Annotation Ontology, we need to understand what sorts of
questions experts would like to retrieve about the PDDIs. The competency questions
below were elicited from experienced editors of clinically oriented drug
compendia during the process of developing DIKB 1.2. Most fall into three categories:
nding assertions and evidence; assessing the evidence; and enabling updates. A
second area of interest is statistical information about the evidence base which
is useful for various analytics related to knowledge base maintainance.
4.1</p>
      <p>Finding assertions and evidence
1. Finding assertions:
(a) List all assertions that are not supported by evidence
(b) Which assertions are supported (or refuted) by just one type of evidence?
(c) Which assertions have evidence from source X (e.g., product labeling)
(d) Which assertions have both evidence for and evidence against from a
single source X?
2. Finding evidence:
(a) List all evidence for or against assertion X (by evidence type, drug, drug
pair, transporter, metabolic enzyme, etc.)
(b) What is the in vitro evidence for assertion X? the in vivo evidence?
(c) List all evidence that has been agged as rejected from entry into the
the knowledge base
(d) Which single evidence items act as support or rebuttal for multiple
assertions of type X (e.g., substrate of assertions)?
4.2</p>
      <p>Assessing the evidence:
1. Understanding evidence coming from a given study:
(a) What data, methods, materials, are reported in evidence item X?
(b) Which evidence items are related to and follow-up on evidence item X?
(c) Which research group conducted the study used for evidence item X?
(d) Are the evidence use assumptions for evidence item X concordant? unique?
non-ambiguous?
8 `Assertion' in DIKB terminology corresponds to a `Claim' in the Micropublications
model; this variation in terms is because the term `claim' is used in a di erent sense
in medical billing.
2. Verifying plausibility of an evidence item:
(a) Has evidence item X been rejected for assertion Y? If so, why and by
whom?
(b) Which other assertions are being supported/challenged by this evidence
item?
(c) What are the assumptions required for use of this evidence item to
support/refute assertion X?
3. Checking assertions about pharmacokinetic parameters (i.e., area
under the concentration time curve (AUC))
(a) How many pharmacokinetic studies used for evidence items in the DIKB
could be used to support or refute an assertion about pharmacokinetic
paramater X (e.g., `X increases AUC')?
(b) How many pharmacokinetic studies in the DIKB used for evidence items
for assertion X are based on data from the product label?
(c) What is the result of averaging (or applying some other statistical
operation) to the values for pharmacokinetic parameter X across all relevant
studies used for evidence items?
4. Checking for di erences in the product labeling:
(a) Are there di erences in the evidence items that were identi ed across
di erent versions of product labeling for the same drug?
(b) What version of product labeling was used for evidence item X? Original
manufacturer or repackager? Most current label or outdated? Is the drug
on market in country X or not? American or country X?
4.3</p>
      <p>Supporting updates to evidence and assertions
1. Changing status of redundant and refuted evidence:
(a) Remove a older version of a redundant evidence item
(b) Change the modality of a supporting evidence item to be a refuting
evidence item
2. Updating when key sources change:
(a) Get all assertions that are supported by evidence items identi ed from
an FDA guidance or other source document just released as an updated
version.
4.4</p>
      <p>Understanding the evidence base</p>
      <sec id="sec-4-1">
        <title>1. Statistical information about the evidence base:</title>
        <p>(a) Number of assertions in the system
(b) Number of evidence items for and against each assertion type
(c) Show the distribution of the levels of evidence for various assertion types
(e.g., pharmacokinetic assertions)</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5 Modeling evidence about drug-drug interactions</title>
      <p>The Micropublications ontology is used to structure the evidence relating to
data, methods, and materials, and the overall indication that evidence mp:supports
or mp:challenges a mp:Claim. We qualify Claims (C1 in the gure) by reusing
identi ers from DRON10 [Hanna2013] and the Protein Ontology11 [Natale2011].
9 http://dbmi-icode-01.dbmi.pitt.edu/dikb-evidence/escitalopram_does_not_
inhibit_cyp2d6.html
10 http://purl.obolibrary.org/obo/dron.owl
11 http://pir.georgetown.edu/pro/
The new model reuses the DIKB evidence taxonomy12 to provide epistemic
quali cation (SQ2, SQ5, SQ6 in the gure) to statements (S1, S2, and S3 in the
gure), data (D1 in the gure), methods (Me1 in the gure), and materials (not
shown in this example). The Open Annotation Data Model (previously shown in
Figure 1) is used to link quotes taken from source documents back to their
originating information artifacts. The approach to modeling other DIKB assertions
would be similar to this example.
6
6.1</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <sec id="sec-6-1">
        <title>Expected Bene ts</title>
        <p>Certain bene ts accrue from upgrading from the current DIKB. Many of the
competency questions (Section 4) are not supported in the DIKB 1.2. The new
model is designed to support these and additional questions relevant in the
domain. Visual inspection of the model suggests that we will be able to answer
some competency questions quite naturally. In particular, nding the assertions
that are not supported by evidence already in the evidence base, the evidence
that should be checked most thoroughly (e.g. evidence that by itself supports
multiple assertions), and the data, methods, and materials associated with a
given evidence item as described in source documents.</p>
        <p>Further, as a Linked Data resource, our new knowledge base will also enable
innovative queries using knowledge from other sources about tagged entities (i.e.,
drugs and proteins) represented in the evidence base. Unlike the current DIKB,
we will be able to render annotations in their original context. We also expect to
be able to support distributed community annotation/curation, since MP and
OA take account of provenance, and since OA is being increasingly adopted by
a variety of annotation tools.
6.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Modeling challenges</title>
        <p>Our project does raise certain modeling challenges. To date, MP has not been
used to represent both unstructured claims and the related logical sentences.
Figure 1 shows the assertion escitalopram does not inhibit CYP2D6 as unstructured
text. However, the DIKB requires that 1) assertions about PDDIs be formulated
by experts prior to collecting evidence, and 2) that the assertions be represented
both as unstructured statements and sentences in a logical formalism. Careful
thought is being put into how to properly accommodate this use case. Such
challenges are to be expected since MP is a relatively new ontology and since this is
a new application of it.</p>
        <p>Another challenge is to ensure that, as the evidence base scales, competency
questions can be answered e ciently. To address this, we building the model
using an iterative design-and-test approach. In this process, e cient querying is
a key requirement.
12 http://bioportal.bioontology.org/ontologies/DIKB
For enabling synthesis over the PDDI information, the model is not the only
concern. Applying this model will require integration work. One challenge is
inherent to scholarly documents: the existing evidence items within the DIKB
refer to many data, materials, and methods that exist only in PDF documents
accessible only through proprietary portals or academic library systems.
Consequently, resolving annotations requires a method for pointing to proprietary
oa:target s.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusions &amp; Future Work</title>
      <p>13 http://dailymed.nlm.nih.gov/dailymed/about.cfm
of the art from scienti c documents. The knowledge representations we are now
creating will be bene cial for integrating PDDI evidence, and we hope they will
inspire an increased use of linked data for evidence synthesis in other domains.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work was carried out during the tenure of an ERCIM \Alain Bensoussan"
Fellowship Programme. The research leading to these results has received
funding from the European Union Seventh Framework Programme (FP7/2007-2013)
under grant agreement no 246016, and a grant from the National Library of
Medicine (1R01LM011838-01). We thank Carol Collins, Lisa Hines, and John R
Horn for serving on the Evidence Panel of \Addressing PDDI Evidence Gaps",
and for contributing to the competency questions presented here.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Abarca2003] Abarca, Jacob, Daniel C. Malone, Edward P. Armstrong,
          <string-name>
            <given-names>Amy J.</given-names>
            <surname>Grizzle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Philip D.</given-names>
            <surname>Hansten</surname>
          </string-name>
          ,
          <string-name>
            <surname>Robin C. Van Bergen</surname>
          </string-name>
          , and Richard B. Lipton. \
          <article-title>Concordance of severity ratings provided in four drug interaction compendia</article-title>
          .
          <source>" Journal of the American Pharmacists Association</source>
          <volume>44</volume>
          ;
          <issue>2</issue>
          (
          <year>2003</year>
          ):
          <volume>136</volume>
          {
          <fpage>141</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Boyce2014]
          <string-name>
            <surname>Boyce</surname>
            , R.D. \
            <given-names>A Draft</given-names>
          </string-name>
          <string-name>
            <surname>Evidence</surname>
          </string-name>
          <article-title>Taxonomy and Inclusion Criteria for the Drug Interaction Knowledge Base."</article-title>
          <source>August 9</source>
          ,
          <year>2014</year>
          , url: http://purl.net/net/druginteraction
          <article-title>-knowledge-base/evidence-types-and-inclusion-criteria</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Boyce2007] Boyce,
          <string-name>
            <given-names>Richard D.</given-names>
            ,
            <surname>Carol</surname>
          </string-name>
          <string-name>
            <surname>Collins</surname>
          </string-name>
          , John Horn, and Ira Kalet. \
          <article-title>Modeling Drug Mechanism Knowledge Using Evidence</article-title>
          and
          <string-name>
            <given-names>Truth</given-names>
            <surname>Maintenance</surname>
          </string-name>
          .
          <source>" IEEE Transactions on Information Technology in Biomedicine 11;4</source>
          (
          <year>2007</year>
          ):
          <volume>386</volume>
          {
          <fpage>397</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Boyce2009] Boyce,
          <string-name>
            <given-names>Richard D.</given-names>
            ,
            <surname>Carol</surname>
          </string-name>
          <string-name>
            <surname>Collins</surname>
          </string-name>
          , John Horn, and Ira Kalet. \
          <article-title>Computing with evidence: Part I: A drug-mechanism evidence taxonomy oriented toward con dence assignment</article-title>
          .
          <source>" Journal of Biomedical Informatics</source>
          <volume>42</volume>
          ;
          <issue>6</issue>
          (
          <year>2009</year>
          ):
          <volume>979</volume>
          {
          <fpage>989</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Ciccarese2008] Ciccarese,
          <string-name>
            <given-names>Paolo N.</given-names>
            ,
            <surname>Elizabeth</surname>
          </string-name>
          <string-name>
            <surname>Wu</surname>
          </string-name>
          , Gwen Wong, Marco Ocana, June Kinoshita, Alan Ruttenberg, and Tim Clark. \
          <article-title>The SWAN biomedical discourse ontology</article-title>
          .
          <source>" Journal of Biomedical Informatics</source>
          <volume>41</volume>
          ;
          <issue>5</issue>
          (
          <year>2008</year>
          ):
          <volume>739</volume>
          {
          <fpage>751</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Ciccarese2014] Ciccarese,
          <string-name>
            <given-names>Paolo N.</given-names>
            ,
            <surname>Marco</surname>
          </string-name>
          <string-name>
            <surname>Ocana</surname>
          </string-name>
          , and Tim Clark. \
          <article-title>Open semantic annotation of scienti c publications using DOMEO."</article-title>
          <source>Journal of Biomedical Semantics Apr</source>
          <volume>24</volume>
          ;
          <issue>3</issue>
          (
          <year>2012</year>
          ):
          <issue>Suppl 1</issue>
          :
          <fpage>S1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Clark2014] Clark, Tim,
          <string-name>
            <given-names>Paolo N.</given-names>
            <surname>Ciccarese</surname>
          </string-name>
          , and
          <string-name>
            <surname>Carole</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Goble</surname>
          </string-name>
          . \
          <article-title>Micropublications: a semantic model for claims, evidence, arguments and annotations in biomedical communications</article-title>
          .
          <source>" Journal of Biomedical Semantics</source>
          <volume>5</volume>
          ;
          <fpage>28</fpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Hanna2013] Hanna, Josh, Eric Joseph, Mathias Brochhausen, and
          <string-name>
            <surname>William</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hogan</surname>
          </string-name>
          . \
          <article-title>Building a drug ontology based on RxNorm and other sources</article-title>
          .
          <source>" Journal of Biomedical Semantics</source>
          <volume>4</volume>
          (
          <year>2013</year>
          ):
          <volume>44</volume>
          {
          <fpage>52</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Natale2011] Natale,
          <string-name>
            <given-names>Darren A.</given-names>
            ,
            <surname>Cecilia</surname>
          </string-name>
          <string-name>
            <given-names>N.</given-names>
            <surname>Arighi</surname>
          </string-name>
          , Winona C.
          <article-title>Barker, Judith A</article-title>
          .
          <string-name>
            <surname>Blake</surname>
            ,
            <given-names>Carol J.</given-names>
          </string-name>
          <string-name>
            <surname>Bult</surname>
            , Michael Caudy,
            <given-names>Harold J.</given-names>
          </string-name>
          <string-name>
            <surname>Drabkin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter D'Eustachio</surname>
            ,
            <given-names>Alexei V.</given-names>
          </string-name>
          <string-name>
            <surname>Evsikov</surname>
          </string-name>
          , Hongzhan Huang, Jules Nchoutmboube,
          <string-name>
            <surname>Natalia</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>Barry</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>Jian</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <surname>Cathy H. Wu</surname>
          </string-name>
          . \
          <article-title>The Protein Ontology: a structured representation of protein forms and complexes." Nucleic acids research 39, no</article-title>
          .
          <issue>suppl 1</issue>
          (
          <year>2011</year>
          )
          <article-title>: D539{ D545.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>