<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>InfyNLP at SMM4H Task 2: Stacked Ensemble of Shallow Convolutional Neural Networks for Identifying Personal Medication Intake from Twitter</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Indian Institute of Technology</institution>
          <addr-line>Guwahati, Assam</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Infosys Limited</institution>
          ,
          <addr-line>Palo Alto, CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Jasper Friedrichs</institution>
          ,
          <addr-line>MS</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes Infosys's participation in the “2nd Social Media Mining for Health Applications Shared Task at AMIA, 2017, Task 2”. Mining social media messages for health and drug related information has received significant interest in pharmacovigilance research. This task targets at developing automated classification models for identifying tweets containing descriptions of personal intake of medicines. Towards this objective we train a stacked ensemble of shallow convolutional neural network (CNN) models on an annotated dataset provided by the organizers. We use random search for tuning the hyper-parameters of the CNN and submit an ensemble of best models for the prediction task. Our system secured first place among 9 teams, with a micro-averaged F-score of 0.693.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Methods</title>
      <p>Deep learning systems have recently shown to achieve top results in shared tasks related to natural language
processing on tweets4. Historically, ensemble learning has proved to be very effective in most of the machine
learning tasks including the famous winning solution of the Netflix Prize5. Ensemble models can offer diversity over
training data splits, random initialization of the same model or model architectures, and a combination of multiple
average or low performing learners to produce a robust and high-performing learning model. A convolutional neural
network (CNN)6 is a deep learning architecture, that has shown strong performance on sentence-level text
classification. Even fairly simple CNNs evaluate at a level of or even better than more complex deep learning
architectures7. Therefore, we designed and implemented a stacked ensemble of shallow convolutional neural
networks (Figure 1) for solving the classification task presented in this paper. The main intuition behind developing
such an ensemble was to take the best of all worlds. Next, we explain stacked ensemble of CNNs.
In order to get the best results from any classification model, hyperparameter tuning is a key step and CNNs are no
exception. While the existing literature offers guidance on practical design decisions, identifying the best
hyperparameters of a CNN requires experimentation. This requires evaluating trained models on a cross-validation
dataset and choosing the best hyperparameters manually that produce best results. Automated hyperparameter
searching methods like grid search, random search, and Bayesian optimization methods are also popularly used. In
our presented system we use random search8, to explore the hyperparameters of a shallow CNN architecture and
form an ensemble of the best models, which we refer to as a stacked ensemble. Next, we share the detailed output
and analysis of our experiments.</p>
      <sec id="sec-1-1">
        <title>Dataset and Data Preprocessing</title>
        <p>The organizers provided 8000 annotated tweets as a training dataset and 2260 additional tweets as development
dataset. We collected the tweets using the script
provided along with the dataset, by querying Twitter’s Class 1 Class 2 Class 3 Total
API. However, we could not collect all the tweets as Train 1847 3027 4789 9663
some of them were not available at the moment when
we executed our collection process. Later, the Test 1731 2697 3085 7513
organizers also shared the test dataset, that was used
for calculating the final scores of the submitted Table 1. Shared task data distribution. Class 1, 2
models. A distribution of tweets provided for each and 3 represent personal medication intake, possible
class and the mapping of each class is shown in Table medication intake, and no medication intake,
1. It is to be noted over here that for training our respectively.
models, we combine the training and development dataset
provided and treat it as our training dataset, therefore learning
our models using 9663 tweets with 5-fold cross validation.</p>
        <p>We use Spacy4 for all our data preprocessing and cleaning
activities. We do not remove stopwords. Each document in our
training and test dataset is converted to a fixed size document
of 47 words/tokens. We use two pre-trained word embeddings
godin9 and shin10, shared by the authors. Each of these
embeddings are of 400 dimensions. Each word in the input
tweet is represented by its corresponding embedding vector,
when present in the vocabulary of the model.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Hyperparameters for the CNNs</title>
        <p>We use Xavier weight initialization scheme11, for initializing
the weights of the CNNs. Adam with two annealing restarts has
been shown to work faster and perform better than SGD in
other NLP tasks12. Therefore, we use the same as our
optimization algorithm. We use five filters with varying filter
sizes in the convolution layer and use dropout during the
training process. The models are implemented using
TensorFlow5. The entire ranges of the hyperparameters that we
give to our random search procedure is shown in Table 2. The
word embedding model to be used during training is also
treated as a hyperparameter.</p>
        <p>Hyperparameter
adam_b2
n_dense_output
keep_prob
(dropout)
batch_size
learning_rate
word_embedding
n_filters
filter_sizes</p>
        <p>Range
An ensemble of five CNNs is trained during 5-fold cross-validation training performed on our combined training
dataset along with random search on the hyperparameter ranges. We train 99 such ensembles. The performance of
the top 20 ensemble on the training data (blue) and on the test data (red) is shown in Figure 2. The models are
arranged in the order of their decreasing training performance. We create stacked ensembles from these ensembles
by taking top K ensemble models. We show the performances for such top K stacked ensembles (brown), as well.
The detailed performances on the evaluation metrics of Top 3, Top 10 and Top 20 stacked ensembles are shown in</p>
        <sec id="sec-1-2-1">
          <title>4 https://spacy.io/</title>
        </sec>
        <sec id="sec-1-2-2">
          <title>5 https://www.tensorflow.org</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Conclusion and Future Work</title>
      <p>By participating in this shared task we showed the generic effectiveness of CNNs and ensembles on identification of
personal medication intake from Twitter posts. Our proposed architecture of stacked ensemble of shallow CNNs,
out-performed other models submitted in the task. This provided an empirical evaluation of our initial aim of
combining ensembles with CNNs along with training the models using random search on the hyperparameters. In
the future, we plan to work more on hyperparameter tuning using random search and various other search
procedures and analyze their effectiveness. Instead of using pre-trained word embeddings it would also be
interesting to look at the performance of our models by training word and phrase embeddings on a domain specific
dataset of tweets. We would also like to formalize the architecture of stacked ensembles of CNNs and compare our
models with an exhaustive set of other deep learning as well as traditional machine learning models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Hrmark</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Grootheest</surname>
            <given-names>AC</given-names>
          </string-name>
          .
          <article-title>Pharmacovigilance: methods, recent developments and future perspectives</article-title>
          .
          <source>Euro-pean journal of clinical pharmacology</source>
          .
          <source>2008 Aug</source>
          <volume>1</volume>
          ;
          <issue>64</issue>
          (
          <issue>8</issue>
          ):
          <fpage>743</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Klein</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarker</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rouhizadeh</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Connor</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Detecting Personal Medication Intake in Twitter: An Annotated Corpus and Baseline Classification System</article-title>
          .
          <source>BioNLP</source>
          <year>2017</year>
          .
          <year>2017</year>
          :
          <fpage>136</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Sarker</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikfarjam</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Social media mining shared task workshop</article-title>
          .
          <source>In Biocomputing 2016: Proceedings of the Pacific Symposium</source>
          <year>2016</year>
          (pp.
          <fpage>581</fpage>
          -
          <lpage>592</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Rosenthal</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farra</surname>
            <given-names>N</given-names>
          </string-name>
          , Nakov P. SemEval
          <article-title>-2017 task 4: Sentiment analysis in Twitter</article-title>
          .
          <source>In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017) 2017</source>
          (pp.
          <fpage>502</fpage>
          -
          <lpage>518</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bell</surname>
            <given-names>RM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koren</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volinsky</surname>
            <given-names>C</given-names>
          </string-name>
          .
          <article-title>All together now: A perspective on the netflix prize</article-title>
          .
          <source>Chance. 2010 Jan</source>
          <volume>1</volume>
          ;
          <issue>23</issue>
          (
          <issue>1</issue>
          ):
          <fpage>24</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kim</surname>
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Convolutional neural networks for sentence classification</article-title>
          .
          <source>arXiv preprint arXiv:1408.5882. 2014 Aug</source>
          <volume>25</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Le</surname>
            <given-names>HT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cerisara</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <article-title>Denis A. Do Convolutional Networks need to be Deep for Text Classification?</article-title>
          .
          <source>arXiv preprint arXiv:1707.04108. 2017 Jul</source>
          <volume>13</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Bergstra</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Random search for hyper-parameter optimization</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          .
          <year>2012</year>
          ;
          <volume>13</volume>
          (Feb):
          <fpage>281</fpage>
          -
          <lpage>305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Godin</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vandersmissen</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Neve</surname>
            <given-names>W</given-names>
          </string-name>
          , Van de Walle R.
          <article-title>Multimedia lab@ acl w-nut ner shared task: named entity recognition for twitter microposts using distributed word representations</article-title>
          .
          <source>ACL-IJCNLP. 2015 Jul</source>
          <volume>31</volume>
          ;
          <year>2015</year>
          :
          <fpage>146</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Shin</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            <given-names>JD</given-names>
          </string-name>
          .
          <article-title>Lexicon integrated cnn models with attention for sentiment analysis</article-title>
          .
          <source>arXiv preprint arXiv:1610.06272</source>
          . 2016 Oct 20.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Glorot</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            <given-names>Y. Understanding</given-names>
          </string-name>
          <article-title>the difficulty of training deep feedforward neural networks</article-title>
          .
          <source>In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics 2010 Mar</source>
          <volume>31</volume>
          (pp.
          <fpage>249</fpage>
          -
          <lpage>256</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Denkowski</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neubig</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Stronger Baselines for Trustable Results in Neural Machine Translation</article-title>
          .
          <source>arXiv preprint arXiv:1706.09733. 2017 Jun</source>
          <volume>29</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>