<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ensemble of Convolutional Neural Networks for Medicine Intake Recognition in Twitter</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Future Technologies, University of Turku</institution>
          ,
          <addr-line>Turku</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Nursing Science, University of Turku</institution>
          ,
          <addr-line>Turku</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Turku Centre for Computer Science</institution>
          ,
          <addr-line>Turku</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Turku Graduate School, University of Turku</institution>
          ,
          <addr-line>Turku</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present the results from our participation in the 2nd Social Media Mining for Health Applications Shared Task Task 2. The goal of this task is to develop systems capable of recognizing mentions of medication intake in Twitter. Our best performing classification system is an ensemble of neural networks with features generated by word- and character-level convolutional neural network channels and a condensed weighted bag-of-words representation. A relatively strong performance is achieved, with an F-score of 66.3 according to the official evaluation, resulting in the 5th place in the shared task with performance close to the best systems created by other participating teams.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>The organizers provided a training dataset with manually assigned labels (intake, possible intake, non-intake) for
each tweet. The intake class is defined as clearly expressing a personal intake of medication, whereas the possible
intake is more ambiguous, yet still suggesting an intake by the tweet writer. The non-intake class includes the rest
of the tweets, all of which include a mention of a drug, but refer to an intake by another person or discuss the drugs
in general. Approximately 50% of the data belongs to the non-intake class, whereas the intake and possible intake
classes constitute 19% and 31% of the data, respectively.</p>
      <p>Due to the data sharing restrictions of Twitter, the organizers only provided the IDs of the tweets instead of their actual
content. Since we started investigating this task much later than the data was released, we were only able to obtain
the contents for 7444 tweets out of the total 8000 annotated tweets as some of the content had been already removed
by the Twitter users, i.e. we had 7% less training data than teams who were involved in the shared task since the very
*: these authors have contributed equally.
beginning. The organizers also provided a separate development dataset, which consists of 2260 annotated tweets, of
which we also lost roughly the same proportion. In the Results and Discussion section we evaluate the impact of the
lost data in more detail.</p>
    </sec>
    <sec id="sec-3">
      <title>Method</title>
      <p>For our baseline approach we form term frequency–inverse document frequency (TF–IDF) weighted sparse
bag-ofwords (BOW) representations for all given tweets1. These representations are not only constructed for single tokens,
but also token bigrams, trigrams and character n-grams of length 1 to 4. These representations are then fed as features
to a linear SVM classifier2. The regularization parameter is selected to optimize the micro-averaged F-score of intake
and possible intake classes, the official evaluation metric, on the development set. For the final submission in the
shared task we merge the training and development sets and train the system on the combined dataset.
As the sparse representations are not able to generalize well to unseen vocabulary, we also test various NN approaches
on this task. The final system is based on an ensemble of convolutional neural networks (CNNs)3 and utilizes word
and character information.</p>
      <p>Each tweet is represented as two separate sequences: words and characters, both of which are processed with separate
convolutional channels. Each element in these sequences is represented with a latent feature vector, i.e. an embedding.
The word embeddings are initialized using word2vec4 trained with approximately 1 billion drug related tweets as
provided by Sarker and Gonzalez5. We also tested GloVe vectors trained on 2B general domain tweets6, but these
experiments resulted in a decreased performance. The character embeddings are initialized randomly, but the network
is allowed to backpropagate to both word and character embeddings.</p>
      <p>The convolutional kernels are applied on the aforementioned two sequences using sliding windows. The outputs are
subsequently max-pooled and concatenated. The concatenated vectors are further fed through two densely connected
layers, the latter having the output dimensionality corresponding to the number of labels in the data set with softmax
activation.</p>
      <p>In addition to the convolutional layers we utilize the same TF–IDF weighted sparse vector representations as in the
baseline method. As these representations have dimensionality in the order of hundreds of thousands, we first densify
the representations to 4000 dimensional vectors using truncated singular-value decomposition (SVD)7. These vectors
are concatenated alongside the CNN outputs. This dimensionality reduction is performed mainly due to computational
reasons since the approach was prototyped on a consumer grade GPU with limited amount of memory. Projecting the
sparse vectors to 4K dimensions preserves 74% of the variance in the data and may have caused a minor performance
loss.</p>
      <p>The network is trained on the official training data using the Nadam optimization algorithm. The network is regularized
with dropout rate of 0.2 after the first dense layer, no explicit regularization is applied on the convolutional part of the
network. The training is stopped once the performance on the development set is no longer improving, measured with
the official evaluation metric. Table 1 shows the comprehensive list of used hyperparameters.</p>
    </sec>
    <sec id="sec-4">
      <title>Hyper-parameter</title>
      <p>Character embedding dimensionality
Word embedding dimensionality
Character CNN, number of filters per window size
Character CNN, window sizes
Word CNN, number of filters per window size
Word CNN, window sizes
Dimensionality of first dense layer
Dropout rate
Activation functions</p>
      <p>
        Optimal value
25
400
50
[
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2,3,4,5</xref>
        ]
200
[
        <xref ref-type="bibr" rid="ref2 ref4">2,4</xref>
        ]
400
0.2
tanh
      </p>
      <p>Tested Values
[25,50,75,100]</p>
      <p>pre-trained
[50,100,150,200]</p>
      <p>
        [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2,3,4,5</xref>
        ]
[100,200,300]
any subset of [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2,3,4,5</xref>
        ]
[100,200,300,400,500]
      </p>
      <p>[0,0.2,0.5]
[ReLU, tanh, sigmoid]</p>
      <p>Training the network on this dataset resulted in relatively large variance in the measured performance, caused by the
random initialization of the weights. Thus we stabilize the system by training 15 networks, all identical apart from the
initial (random) weights. We then select the optimal subset of these networks, as measured on the development set,
for the final system where the final predictions are created by summing the confidences of all selected networks and
choosing the label with the highest overall confidence. The final system included a subset of 6 neural networks out of
the 15. We note that this approach may potentially overfit on the development set.</p>
      <p>Other NN architectures experimented and tested during this shared task include various versions of BiLSTM and
attention based networks8, 9 but none of these experiments resulted in better performance than the CNN architecture
described in detail. However, due to the time limits of the shared task, we cannot reject the possibility of these
approaches being competitive as well.</p>
      <p>We also experimented with a way of (pre-)tuning the utilized word embeddings to this specific classification task in
an attempt to give the word-level CNN a better starting point for the training. This was done using the principles
underlying the random indexing (RI)10 method. Unique index vectors are first assigned to each of the three classes
(intake, possible intake and non-intake), and empty context vectors are assigned to each word in the data set. When
traversing the training set, each word, in each tweet, have the index vector associated with the tweet’s class added to
their context vector. After training, the resulting word context vectors are normalized to unit length and summed with
the corresponding word embeddings/vectors generated using word2vec5 (also normalized to unit length). To make the
signal provided by the RI approach have a modest impact on the conjoint vectors, these vectors are first multiplied with
a weight of 0.3. However, the described approach did not seem to result in a positive performance impact, compared
to using the original word2vec generated embeddings.</p>
      <p>We also tested the potential benefits of including part-of-speech (POS) tags, which were produced using the Twitter
NLP toolkit11. The sequences of POS tags were treated in similar fashion to the word and character sequences.
Although the benefits of POS tagging are intuitive as for instance verbs in past tense are twice as common in the intake
class as in the possible intake, we did not see any increase in the performance when POS tags were utilized.</p>
    </sec>
    <sec id="sec-5">
      <title>Results and Discussion</title>
      <p>We measure the performance of our systems using micro-averaged F-score of intake and possible intake classes,
following the official evaluation, and conduct all our experiments on the official development set. However, the
reported results are not directly comparable with other systems as we only had access to a subset of the original data
(see the Data section). The results on the test set are as reported by the organizers and thus comparable to other
systems.</p>
      <p>The overall performance of our baseline (i.e. SVM) and CNN-based systems are relatively strong, resulting in
Fscores of 69.6 and 72.7 on the development set respectively (see Table 2). We suspect the main advantage of the CNN
approach to be the generalizability of the word embeddings, which leads to the 3.1pp improvement in F-score. We
also briefly tested a nonlinear multilayer perceptron with the same BOW features as used in the SVM, which led to a
slight improvement over using the linear SVM model, but was not able to outperform our CNN-based system. Thus
the model complexity alone does not explain the performance difference between the SVM and CNN approaches.
An unexpected observation is that the intake class seems to be harder to predict than the possible intake, although
eyeballing the data suggests otherwise and the annotation guidelines provide more precise definition for the intake class.
Also for the CNN-based system it seems that the precision and recall are rather well balanced, thus no performance
improvements could have been gained through further fine tuning of these metrics.</p>
      <p>The test set results follow the same patterns as the development set evaluation: SVM and CNN systems reach F-scores
of 64.2 and 66.3, respectively. Thus it seems that either the test set is somewhat harder than the development set or
both of the systems are overfitting equally on the development set, even though the implemented ensemble system
with CNNs could have caused greater overfitting. According to the official evaluation, our best system loses to the
winning system by 3pp in F-score, the difference being roughly the same in both precision and recall. This places our
system in the 5th position in the shared task.</p>
      <p>By inspecting the confusion matrices it can be concluded that our classifiers tend to confuse intake class with both
SVM
CNN</p>
      <sec id="sec-5-1">
        <title>InfyNLP</title>
      </sec>
      <sec id="sec-5-2">
        <title>Intake</title>
        <p>Possible Intake
Overall
Intake
Possible Intake
Overall
Overall
possible intake and non-intake classes equally often, whereas the possible intake is more often confused with the
non-intake class.</p>
        <p>As we only had access to a partial training data, we try to estimate how much the performance of the systems could
have been improved with additional data. To accomplish this, we train the CNN-based system with different subsets
of the training data, starting from 5K training examples and incrementally increasing the size in steps of 300 up to the
whole training data available to us. After every increment we evaluate the system’s performance on the development
set. To reduce the variance caused by different initial random weights, we train 5 networks with each subset of
the training data, and calculate the mean performance for each subset. Fitting a linear regression on the resulting
measurements shows that in this region, the learning curve is fairly linear and decent performance improvements can
be gained by adding more training data. Assuming a performance increase equal to the slope of the fitted regression
line, having the full training dataset would have increased our performance by 0.7pp in F-score, placing our system
close to the top 3 teams in the shared task.</p>
        <p>As most approaches we tested, as well as the systems created by other teams, resulted roughly in the same performance
level, we wanted to assess what would be a theoretical performance limit for this task. To this end, we manually
annotated a random subset of 100 tweets from the development set and evaluated the annotations against the gold
standard. Surprisingly our manual annotations reached only an F-score of 59.3, notably lower than the developed
systems or what the official inter-annotator agreement would suggest12. This indicates that the task is complex even
for humans and deep understanding of the annotation guidelines is required for high quality annotations.
We have shown that strong results in detecting tweets describing personal medication intake can be achieved using
convolutional neural networks and word embeddings. However, more traditional methods relying on bag-of-words
features and linear classifiers also result in competitive performance. Considering that such a system can be
implemented in less than an hour with the existing tools and libraries, and is easily interpretable, the simpler methods may
be a more practical choice in many use cases.</p>
        <p>Since the amount of training data for this task is fairly limited, we plan to explore various approaches for pretraining
NN classifiers as a future task. The goal here is to find a suitable proxy task related to the domain for initializing
the network before the actual training. Such a task could be, for instance, sentiment detection as many of the tweets
expressing drug intake also express a certain sentiment about the condition of the user or the effects of the drug.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements References</title>
      <p>This work was supported by ATT Tieto ka¨ytto¨o¨n grant and Tekes – Ra¨a¨ta¨li project (no. 644/31/2015). Computational
resources were provided by CSC - IT Center For Science Ltd., Espoo, Finland.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Salton</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckley</surname>
            <given-names>C</given-names>
          </string-name>
          .
          <article-title>Term-weighting approaches in automatic text retrieval</article-title>
          .
          <source>Information processing &amp; management.</source>
          <year>1988</year>
          ;
          <volume>24</volume>
          (
          <issue>5</issue>
          ):
          <fpage>513</fpage>
          -
          <lpage>523</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Joachims</surname>
            <given-names>T.</given-names>
          </string-name>
          <article-title>Making large-scale SVM learning practical</article-title>
          . In: Scho¨lkopf
          <string-name>
            <given-names>B</given-names>
            ,
            <surname>Burges</surname>
          </string-name>
          <string-name>
            <given-names>C</given-names>
            ,
            <surname>Smola</surname>
          </string-name>
          <string-name>
            <surname>A</surname>
          </string-name>
          , editors.
          <source>Advances in Kernel Methods - Support Vector Learning</source>
          . Cambridge, MA: MIT Press;
          <year>1999</year>
          . p.
          <fpage>169</fpage>
          -
          <lpage>184</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>LeCun</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            <given-names>Y</given-names>
          </string-name>
          , et al.
          <article-title>Convolutional networks for images, speech, and time series. The handbook of brain theory and neural networks</article-title>
          .
          <year>1995</year>
          ;
          <volume>3361</volume>
          (
          <issue>10</issue>
          ):
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Mikolov</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            <given-names>GS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          ;
          <year>2013</year>
          . p.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Sarker</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            <given-names>G.</given-names>
          </string-name>
          <article-title>A corpus for mining drug-related knowledge from Twitter chatter: language models and their utilities</article-title>
          .
          <source>Data in Brief</source>
          .
          <year>2017</year>
          ;
          <volume>10</volume>
          (
          <string-name>
            <surname>Supplement</surname>
            <given-names>C</given-names>
          </string-name>
          ):
          <fpage>122</fpage>
          -
          <lpage>131</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Pennington</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            <given-names>C</given-names>
          </string-name>
          .
          <article-title>Glove: global vectors for word representation</article-title>
          .
          <source>In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          ;
          <year>2014</year>
          . p.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Halko</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinsson</surname>
            <given-names>PG</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tropp</surname>
            <given-names>JA</given-names>
          </string-name>
          .
          <article-title>Finding structure with randomness: stochastic algorithms for constructing approximate matrix decompositions</article-title>
          .
          <year>2009</year>
          ;.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hochreiter</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation</source>
          .
          <year>1997</year>
          ;
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Luong</surname>
            <given-names>MT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pham</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            <given-names>CD</given-names>
          </string-name>
          .
          <article-title>Effective approaches to attention-based neural machine translation</article-title>
          .
          <source>arXiv preprint arXiv:150804025</source>
          .
          <year>2015</year>
          ;.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kanerva</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kristofersson</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holst</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Random indexing of text samples for latent semantic analysis</article-title>
          .
          <source>In: Proceedings of 22nd Annual Conference of the Cognitive Science Society</source>
          . Philadelphia, PA, USA;
          <year>2000</year>
          . p.
          <fpage>1036</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Owoputi</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Connor</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyer</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gimpel</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schneider</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            <given-names>NA</given-names>
          </string-name>
          .
          <article-title>Improved part-of-speech tagging for online conversational text with word clusters</article-title>
          . Association for Computational Linguistics;
          <year>2013</year>
          . .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Klein</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarker</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rouhizadeh</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Connor</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Detecting personal medication intake in Twitter: an annotated corpus and baseline classification system</article-title>
          .
          <source>In: BioNLP 2017</source>
          . Vancouver, Canada,: Association for Computational Linguistics;
          <year>2017</year>
          . p.
          <fpage>136</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>