<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detection of Adverse Drug Reaction from Twitter Data</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Bioengineering, University of Pittsburgh</institution>
          ,
          <addr-line>Pittsburgh, PA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Biomedical Informatics, University of Pittsburgh</institution>
          ,
          <addr-line>Pittsburgh, PA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Fuchiang Tsui</institution>
          ,
          <addr-line>PhD</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Intelligent Systems Program, University of Pittsburgh</institution>
          ,
          <addr-line>Pittsburgh, PA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In response to the challenges set forth by the 2nd Social Media Mining for Health Applications Shared Task 2017, we describe a framework to automatically detect Twitter tweets with mentioned adverse drug reactions (ADRs). We used a dataset with 10822 annotated tweets provided by the event organizers to develop a framework comprised of a preprocessing module, 6 feature extraction modules, and one predictive modeling module with 5 models (K2 Bayesian network, naïve Bayes, decision tree, random forest, and support vector machines). To evaluate our framework, we employed a blind test dataset with 9961 tweets provided by the event organizer. The area under the ROC curve (AUC) from the K2 Bayesian network model was 0.74 (95% C.I. 0,721-0.759) and F-Measure ranged from 0.339 to 0.342. The described framework in this paper demonstrated a potential public health surveillance tool for ADR surveillance from Twitter tweets.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>
        The training dataset provided by SMM4H 2017 originally comprised of total 15,667 tweet identifiers with annotated
ADR results from two batches: each with 10,822 and 4845 identifiers respectively; after employing the python script
provided by the event organizer, we only retrieved 10,281 tweets for the training dataset. The blind test dataset
released 7 days prior to the competition deadline comprised of 9,961 tweets. All the datasets were annotated and
provided by the event organizer.
Figure 1 summarizes our proposed framework, which contained a data preprocessing module, and 6 feature extraction
modules, and one predictive modeling module. We first preprocessed each tweet in the training dataset in 7 steps: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
lowercase text conversion, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) sentence partition, (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) stop words removal, (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) tokenization, (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) spelling correction, (
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
lemmatization and (
        <xref ref-type="bibr" rid="ref7">7</xref>
        ) part-of-speech (POS) tagging. In the stop words removal step, we removed commonly used
words that are typically ignored by a search engine, such as “the”, “a”, “at”, etc. In the (Twitter-specific) tokenization
step, we parsed Twitter-specific tokens such as retweet flag (RT), @USERNAME, Hashtag (#), URL and emoticon,
and used them as features for each tweet. We used the Natural Language Processing Toolkit6 (NLTK) to perform the
data preprocessing.
      </p>
    </sec>
    <sec id="sec-3">
      <title>II.2 Feature Modules</title>
      <sec id="sec-3-1">
        <title>II.2.1 N-grams</title>
        <p>We employed six feature extraction modules: N-grams, term frequency-inverse document frequency (TF-IDF), word
embeddings, and sentiment statistics.</p>
        <p>We used unigram (single word), bi-gram (two words) and tri-gram (three words) as features for predictive modeling.
For example, single words like “sick” and “weight” belong to unigram and “Cipro pills” containing two words is a
bigram.</p>
        <p>II.2.2 Term Frequency-Inverse Document Frequency
TF-IDF score measures how important a term, e.g., word, is to a document (tweet) in a corpus (training dataset). The
following equation defines TF-IDF.</p>
        <p>, ,  =  ,  ∗ log</p>
        <p>+ 1
 ,  + 1
where TF(t,d) denotes the term frequency (number of times that a term t appears in a tweet d), |D| is the total number
of tweets, and DF(t, D) is the document frequency (number of tweets that contains the term t in all tweets D). Instead
of using words in the above equation, we used hashed terms for computational efficiency. We used the vector of the
hashed terms from each tweet’s as a feature vector.</p>
        <p>II.2.3 Word Embeddings
We employed the Apache Spark Word2Vec7 to obtain a word embeddings matrix from the training dataset. We used
Skip-gram model with softmax activation function in a neuron. The Skip-gram model learns word vector
representations from sentences by maximizing the average log-likelihood of a word wt in a sentence based on
conditional independence assumption as shown below:</p>
        <p>log  234 2 ,
259 4576
where k is the size of a pre-defined training window, T is the total number of words within a sentence, and the
conditional probability p(&lt;|4) is determined based on the softmax model.</p>
        <p>For each tweet, we computed an averaged word vector based on the word embeddings matrix, which becomes a feature
vector.</p>
        <p>II.2.4 Sentiment statistics
We employed Stanford NLP toolkit8,9 to annotate each sentence's sentiment level as one of 5 levels in a tweet: very
negative, negative, neutral, positive or very positive. From the annotation, we generated 6 sentiment features for each
tweet: 1) the count of negative sentences annotated with negative and very negative sentiment levels, 2) the ratio
between the numbers of negative sentences and all sentences, 3) the count of sentences annotated with neutral
sentiment level, 4) the ratio between the numbers of neutral sentences and all sentences, 5) the count of positive
sentences annotated with positive and very positive sentiment levels, 6) the ratio between the numbers of positive
sentences and all sentences.</p>
        <p>II.2.5 cTAKES findings</p>
      </sec>
      <sec id="sec-3-2">
        <title>II.2.6 ADR tags</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>II.3 Predictive Modeling</title>
      <p>We employed cTAKES10, a medical natural language processing application, to extract symptoms and findings from
clinical narratives. We further customized cTAKES with 2017 US Edition of SNOMED CT11 as a clinical-term
lookup dictionary. We identified 10 annotation types such as drug, disorder, finding, procedure, lab, etc.
We employed ADRMine12, a social-media-specific named entity recognition (NER) application, to identify ADR
mentions in a tweet. Any non-negative ADR mentions in the training dataset were included in our feature set.
We used five machine learning algorithms to build predictive models and validated the models through nested 10-fold
cross validation. All the features were obtained from the feature modules stated in Section II.2. For each fold, we
applied information gain first followed by correlation-based feature selection. We then used K2, Naïve Bayes (NB),
decision tree (DT), random forest (RF), and SVM algorithms to build predictive models.</p>
      <p>We identified three different methods to determine probability thresholds. The first method is cross-validation
threshold, which used the average threshold that optimized F1-score during cross-validation testing. The second
method is prevalence threshold, which calculated the threshold that would produce the same class prevalence observed
in the training set. The third method is training threshold, which calculated the threshold that optimized F1-score in
the entire training dataset.</p>
      <p>The evaluation metrics for predictive models are F-measure (also known as F1-measure), precision, recall, and the
area under the ROC curve (AUC).</p>
      <p>III.</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>Compared to other models, the K2 model had the highest average AUC (0.844) and F1-score (0.469) across 10 test
folds within the 10-fold cross validation (CV). Table 1 lists predictive results from the 5 algorithms within the
10fold CV. DT had the best recall (0.641) and SVM (0.62) had the best precision.
0.581 0.269
The final K2 model comprised 26 nodes from unigram (9 nodes), bi-gram (2 nodes), tri-gram (1 node), cTAKES (8
nodes), word embeddings (1 node), and ADR tags (5 nodes).</p>
      <p>Based on Table 1, we used K2 model for the final test (blind) dataset provided by the organizer released a week before
the predictive result submission. Table 2 summarizes the predictive performance in the test dataset. The AUC of the
K2 model was 0.74 (95% C.I. 0,721-0.759). The K2 model with the training threshold method had the best F-measure
(0.342) and the prevalence threshold method had the best recall (0.394).</p>
      <p>Algorithm
K2-CV</p>
      <p>AUC
0.74
K2-Prevalence 0.74
K2-Training</p>
      <p>0.74
IV.</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>F-Measure Precision Recall
0.341
0.339
0.342
0.333
0.298
0.336
0.35
0.394
0.348
In this study, we built and evaluated a framework to identify individual tweets with mentioned ADR. The framework
comprises of preprocessing module, feature extraction modules, and predictive modeling module. The results
demonstrated a potential public health surveillance tool for ADR surveillance from Twitter tweets.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. U.S. Food and
          <string-name>
            <given-names>Drug</given-names>
            <surname>Administration</surname>
          </string-name>
          .
          <article-title>Costs associated with ADRs</article-title>
          . https://www.fda.gov/drugs/developmentapprovalprocess/developmentresources/druginteractionslabeling/ucm11 4949.htm.
          <source>Accessed October 5</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Impicciatore</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choonara</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clarkson</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Provasi</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pandolfini</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bonati</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>Incidence of adverse drug reactions in paediatric in/out-patients: a systematic review and meta-analysis of prospective studies</article-title>
          .
          <source>Br J Clin Pharmacol</source>
          .
          <year>2001</year>
          ;
          <volume>52</volume>
          (
          <issue>1</issue>
          ):
          <fpage>77</fpage>
          -
          <lpage>83</lpage>
          . doi:
          <volume>10</volume>
          .1046/j.0306-
          <fpage>5251</fpage>
          .
          <year>2001</year>
          .
          <volume>01407</volume>
          .x.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Sultana</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cutroneo</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trifirò</surname>
            <given-names>G.</given-names>
          </string-name>
          <article-title>Clinical and economic burden of adverse drug reactions</article-title>
          .
          <source>J Pharmacol Pharmacother</source>
          .
          <year>2013</year>
          ;
          <volume>4</volume>
          (
          <issue>Suppl 1</issue>
          ):
          <fpage>S73</fpage>
          -
          <lpage>7</lpage>
          . doi:
          <volume>10</volume>
          .4103/
          <fpage>0976</fpage>
          -
          <lpage>500X</lpage>
          .
          <fpage>120957</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bootman</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnson</surname>
            <given-names>JA</given-names>
          </string-name>
          .
          <article-title>Drug-related morbidity and mortality: A cost-of-illness model</article-title>
          .
          <source>Arch Intern Med</source>
          .
          <year>1995</year>
          ;
          <volume>155</volume>
          (
          <issue>18</issue>
          ):
          <fpage>1949</fpage>
          -
          <lpage>1956</lpage>
          . doi:
          <volume>10</volume>
          .1001/archinte.
          <year>1995</year>
          .
          <volume>00430180043006</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Naaman</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boase</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lai</surname>
            <given-names>C</given-names>
          </string-name>
          -H.
          <article-title>Is it really about me?</article-title>
          <source>In: Proceedings of the 2010 ACM Conference on Computer Supported Cooperative Work - CSCW '10. ;</source>
          <year>2010</year>
          :189. doi:
          <volume>10</volume>
          .1145/1718918.1718953.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bird</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loper</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Natural</surname>
          </string-name>
          <article-title>Language Processing with Python</article-title>
          . Vol
          <volume>43</volume>
          .;
          <year>2009</year>
          . doi:
          <volume>10</volume>
          .1097/
          <fpage>00004770</fpage>
          - 200204000-00018.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Meng</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bradley</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yavuz</surname>
            <given-names>B</given-names>
          </string-name>
          , et al.
          <source>[seminal] MLlib: Machine Learning in Apache Spark. J Mach Learn Res</source>
          .
          <year>2016</year>
          ;
          <volume>17</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          . http://arxiv.org/abs/1505.06807.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Manning</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Surdeanu</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bauer</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Finkel</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bethard</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McClosky D. The Stanford CoreNLP Natural Language Processing</surname>
          </string-name>
          <article-title>Toolkit</article-title>
          .
          <source>In: Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations. ;</source>
          <year>2014</year>
          :
          <fpage>55</fpage>
          -
          <lpage>60</lpage>
          . doi:
          <volume>10</volume>
          .3115/v1/
          <fpage>P14</fpage>
          -5010.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Socher</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perelygin</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Recursive deep models for semantic compositionality over a sentiment treebank</article-title>
          .
          <source>Proc …</source>
          .
          <year>2013</year>
          :
          <fpage>1631</fpage>
          -
          <lpage>1642</lpage>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0073791</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Savova</surname>
            <given-names>GK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masanz</surname>
            <given-names>JJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogren</surname>
            <given-names>P V</given-names>
          </string-name>
          , et al.
          <article-title>Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications</article-title>
          .
          <source>J Am Med Inf Assoc</source>
          .
          <year>2010</year>
          ;
          <volume>17</volume>
          (
          <issue>5</issue>
          ):
          <fpage>507</fpage>
          -
          <lpage>513</lpage>
          . doi:
          <volume>10</volume>
          .1136/jamia.
          <year>2009</year>
          .
          <volume>001560</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>NIH-NLM. SNOMED Clinical</surname>
          </string-name>
          <article-title>Terms® (SNOMED CT®)</article-title>
          .
          <source>NIH-US Natl Libr Med</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Nikfarjam</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarker</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Connor</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ginn</surname>
            <given-names>R</given-names>
          </string-name>
          , Gonzalez G.
          <article-title>Pharmacovigilance from social media: Mining adverse drug reaction mentions using sequence labeling with word embedding cluster features</article-title>
          .
          <source>J Am Med Informatics Assoc</source>
          .
          <year>2015</year>
          ;
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <fpage>671</fpage>
          -
          <lpage>681</lpage>
          . doi:
          <volume>10</volume>
          .1093/jamia/ocu041.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>