<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Text Mining for Drug Development: Gathering Insights to Support Decision Making</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sherri Matis-Mitchell</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Consultant</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>DataStar Insights Oxford PA</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>USA Sherrimatismitchell@gmail.com</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>- Drug discovery in Pharma R&amp;D is an information driven process requiring many disparate bits of data from many different sources, both structured and unstructured. Text mining is the key methodology used to extract entities and relationships from unstructured text in the quest for the knowledge needed to bring a safe and effective drug to market and beyond. Much of the insight needed in early drug research to identify drug target to disease relationships and progress a potential drug target, comes from published literature and internal reports. Later stage drug development requires many additional sources of information including case reports, clinical trials, competitive intelligence and other diverse sources. In this publication, I will present 4 different use cases on how text mining is used to drive decision making in drug discovery and development and also how it can be used to identify patient insights from sources such as social media</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Keywords—Text Mining; Drug Discovery; Pharma R&amp;D;
Social media; drug safety; patient journey;</p>
    </sec>
    <sec id="sec-2">
      <title>I. INTRO TO DRUG DISCOVERY</title>
    </sec>
    <sec id="sec-3">
      <title>Drug discovery began with the use of medical plants</title>
      <p>to treat illness. Later drug discovery began by extracting the
pharmacologically active compound from nature products.
Accompanying the genomic revolution, modern drug
discovery methods shifted to identification of disease
associated genes as potential drug targets, followed by
discovery of a compound or biologic that would interact
favorably with the target to treat disease. Because of the
variability of the human population, a drug can vary in
efficacy and safety so extensive preclinical and clinical testing
is require by the regulatory agencies before a drug is launched
on to market. The modern drug discovery process takes place
over an average of 10.7 years and at a cost of about 2.6 billion
dollars1,2. To ensure drug safety, constant post launch,
monitoring of a drug or biologic is required. Finally, drugs can
and do fail at any stage along the process costing millions or
billion or in lost revenue and in some cases, causing harm or
even death.</p>
      <p>Drug discovery requires a lot of information to
succeed and asking the right question can go very far in
ensuring a quicker and surer outcome. By providing quicker
understanding of disease and find better disease targets, and
finding the right compound, and, identifying potential risk
earlier in the process to help make it more efficient and
shorten the timeline to a safe medicine. We need to select the
right dose to maximize efficacy and minimize risk and
develop meaningful trials in the right patient populations to
ensure success Finally, even after the drug is launched,
companies need to monitor for reports of adverse events that
arise following drug treatment but also need to understand
what patients are saying to minimize risk, understand
competing therapies, and alleviate any new issues arising.
Because much of this information is present in written reports,
published literature or study reports, text mining can help
wade thru the pages and help to uncover facts and
relationships in unstructured text3</p>
    </sec>
    <sec id="sec-4">
      <title>II. USE CASES FOR TEXT MINING</title>
      <sec id="sec-4-1">
        <title>A. Text Mining in Early Drug Discovery</title>
        <p>Many diseases like diabetes or cancer arise from a complex
series of events involving multiple genes and pathways but
others, including many rare diseases are associated with a
single gene or even a single mutation in that gene. With the
help of semantically enriched taxonomies of genes, diseases
and drugs and biologics, text mining can identify both old and
new relationships between genes and diseases, genes and
drugs and drugs and disease4. New drug to disease
relationships can represent potential repositioning
opportunities. Most drug projects are designed to cure diseases
that affect large populations but the rare disease patient
community is empowered and are pushing for answers, and
initiatives like the Orphan drug act offer incentives to address
these rare diseases. Understanding these less complicated but
rare diseases that often arise from a single gene mutation can
shed light on more complex diseases and text mining has great
utility in this use case5.</p>
      </sec>
      <sec id="sec-4-2">
        <title>B. Mining for Preclinical and Clinical Drug Safety</title>
        <p>The ultimate goal of preclinical testing is to accurately model
the drug’s safety in animals to predict what will happen in
humans. The risk for adverse events can vary across different
therapeutic areas and some drug classes inherently have
liabilities for certain adverse events6. The tolerance of side
effect can also vary across therapeutic areas. Because only a
fraction of drugs succeeds, there is a large amount of data
from failed compound in unstructured internal reports and
study documents as well as published literature.</p>
        <p>In a small number of cases, unsafe medicines have been
progressed into human trials and beyond due to lack of a
preclinical safety “signal” in and much is being done to
prevent this. In the 1990’s, a number of drugs were found to
cause a life threatening cardiac arrhythmia caused by QT
prolongation and were withdrawn from the market.7. This has
led to testing of all drugs for this liability. One example of
how text mining can benefit is in the building of a reference
compound set for evaluation of QT prolongation. In 2015, The
HESI Pro-Arrhythmia Working Group published on using text
mining to identify both human and non-rodent animal studies
that assessed QT signal concordance between species and
identified drugs that prolonged the QT interval.8 In this work,
text mining was essential to identifying compound to
biological effect to species relationships in the published
literature for expert review.</p>
      </sec>
      <sec id="sec-4-3">
        <title>C. The Role of Text Mining in Drug Submission</title>
        <p>The submission of a new drug to the FDA or requires
proof that the medicine is safe and effective as demonstrated
by non-clinical testing and clinical trials. The submission
package can contain thousands of pages of written material.
Text mining can support this process in a number of ways and
can positively impact the process by saving the time of project
teams and potential reviewers.</p>
        <p>A real world example of how text mining can impact the
submission process will be discussed and while the example
has been stripped of proprietary details, it should still
demonstrate a tangible value. In this case, the team was filing
for a waiver for additional safety studies and based on their
knowledge of the drug’s pharmacology and that of other drugs
in the same class, the team felt this was warranted but still
needed to convince the regulatory agency. A keyword based
literature search found 3000 full text documents that needed
further review to identify the smaller set of documents
relevant to the specific question. The completed review and
summary report was due in 7 months and an outside vendor
quoted a figure of 9 months and 180,000$ to complete the
review. Text mining of the full text documents and subsequent
review was then completed in 2 months saving 5 months’ time
and 180,000$. While the monetary impact of text mining and
other informatics processes R&amp;D processes can be hard to
quantitate, this example demonstrates a clear value of text
mining.</p>
      </sec>
      <sec id="sec-4-4">
        <title>D. Post-launch, Mining Social Media for Patient Insights.</title>
        <p>Social media is a largely untapped source of information
and insights for pharma, on therapy efficacy and safety,
patient journey, unmet need, and customer reputation. When a
patient receives a life altering disease diagnosis and
subsequent treatment, many will turn to social media for
support and additional information. As an industry, pharma
has been using social media to communicate with patients via
channels like Twitter, but this has largely been driven by the
commercial to inform the public. While pharmaceutical
companies are using social media to provide product
information and promotional materials, they should also be
using it to better understand patients' needs and experiences,
and to provide additional education, particularly to those with
chronic illnesses. One emerging trend is to use text mining to
analyze sentiments, identify adverse events and glean insights
from social media. Social media conversations also can inform
R&amp;D and pharmacovigilance efforts.10 Social media is here to
stay and pharma should be responsibly engaging in it to get in
touch with patients. When companies engage, and have the
right tools and analytics capabilities technologies like text
mining they can gain valuable insight into what patients are
saying and use those insights to make better treatments that
improve the quality of lives.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>J. A DiMasi</surname>
          </string-name>
          ,, H. G Grabowski,
          <article-title>Tufts CSDD briefing on R&amp;D cost study</article-title>
          ,
          <year>2014</year>
          http://csdd.tufts.edu/news/complete_story/pr_tufts_csdd_
          <year>2014</year>
          <article-title>_cost_stu dy</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          , C. Cheng, J.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Qin</surname>
          </string-name>
          ,
          <article-title>Developing new drugs from annals of Chinese medicine</article-title>
          ,
          <source>Acta Pharmaceutica Sinica B</source>
          <year>2012</year>
          ;
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>McEntire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Szalkowski</surname>
          </string-name>
          , J. Butler, MS Kuo,
          <string-name>
            <given-names>M</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Freeman</surname>
          </string-name>
          ,
          <string-name>
            <surname>S McQuay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <surname>M McGlashen</surname>
          </string-name>
          ,
          <string-name>
            <surname>WD Cornell</surname>
          </string-name>
          , JJ Xu,
          <article-title>Application of an automated natural language processing (NLP) workflow to enable federated search of external biomedical content in drug discovery and development</article-title>
          .
          <source>Drug Discovery Today</source>
          <year>2016</year>
          ,
          <volume>21</volume>
          (
          <issue>5</issue>
          ) :
          <fpage>826</fpage>
          -
          <lpage>835</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Oellrich</surname>
          </string-name>
          <string-name>
            <surname>A</surname>
          </string-name>
          ,
          <article-title>Hoehndorf Text-mining solutions for biomedical research: enabling integrative biology</article-title>
          .
          <source>Nat Rev Genet</source>
          .
          <year>2012</year>
          Dec;
          <volume>13</volume>
          (
          <issue>12</issue>
          ):
          <fpage>829</fpage>
          -
          <lpage>39</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D</given-names>
            <surname>Sardana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <surname>M Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>RC</given-names>
            <surname>Gudivada</surname>
          </string-name>
          ,
          <string-name>
            <surname>L Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>AG</given-names>
            <surname>Jegga</surname>
          </string-name>
          .
          <article-title>Drug repositioning for orphan diseases</article-title>
          .
          <source>Brief Bioinform</source>
          (
          <year>2011</year>
          )
          <volume>12</volume>
          (
          <issue>4</issue>
          ):
          <fpage>346</fpage>
          -
          <lpage>356</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Cronin</surname>
            <given-names>MT1</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaworska</surname>
            <given-names>JS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walker</surname>
            <given-names>JD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Comber</surname>
            <given-names>MH</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watts</surname>
            <given-names>CD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worth</surname>
            <given-names>AP</given-names>
          </string-name>
          .”
          <article-title>Use of QSARs in international decision-making frameworks to predict health effects of chemical substances</article-title>
          .
          <source>Environ Health Perspect</source>
          .
          <year>2003</year>
          Aug;
          <volume>111</volume>
          (
          <issue>10</issue>
          ):
          <fpage>1391</fpage>
          -
          <lpage>1401</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>CE</given-names>
            <surname>Pollard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N Abi</given-names>
            <surname>Gerges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>MH</given-names>
            <surname>Bridgland-Taylor</surname>
          </string-name>
          , A Easter, TG Hammond, and
          <string-name>
            <surname>J-P Valentin</surname>
          </string-name>
          “
          <article-title>An introduction to QT interval prolongation and non-clinical approaches to assessing and reducing risk”</article-title>
          .
          <source>Br J Pharmacol</source>
          .
          <year>2010</year>
          Jan;
          <volume>159</volume>
          (
          <issue>1</issue>
          ):
          <fpage>12</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>HM</given-names>
            <surname>Vargas</surname>
          </string-name>
          , AS Bass,
          <string-name>
            <given-names>J</given-names>
            <surname>Koerner</surname>
          </string-name>
          , S Matis-Mitchell, MK Pugsley,
          <string-name>
            <given-names>M</given-names>
            <surname>Skinner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Burnham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Bridgland-Taylor</surname>
          </string-name>
          , S Pettit, JP Valentin “
          <article-title>Evaluation of drug-induced QT interval prolongation in animal and human studies: a literature review of concordance”</article-title>
          .
          <source>Br J Pharmacol</source>
          .
          <year>2015</year>
          Aug;
          <volume>172</volume>
          (
          <issue>16</issue>
          ):
          <fpage>4002</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Larkin</surname>
          </string-name>
          , “
          <article-title>Social media for Pharma- an Experts View”</article-title>
          <year>2014</year>
          https://www.elsevier.com/connect/social
          <article-title>-media-for-pharma-an-expertsview</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A</given-names>
            <surname>Sarker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R</given-names>
            <surname>Ginn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Nikfarjam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Jayaraman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Upadhaya</surname>
          </string-name>
          , G Gonzalez.
          <article-title>“Utilizing social media data for pharmacovigilance: A review</article-title>
          .”
          <string-name>
            <given-names>J Biomed</given-names>
            <surname>Inform</surname>
          </string-name>
          .
          <year>2015</year>
          Apr;
          <volume>54</volume>
          :
          <fpage>202</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>