<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Twitter goes to the Doctor: Detecting Medical Tweets using Machine Learning and BERT</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kevin Roitero</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristian Bozzato</string-name>
          <email>bozzato.cristian@spes.uniud.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincenzo Della Mea</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Mizzaro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Serra</string-name>
          <email>serra.giuseppe@uniud.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Udine</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>We propose an effective model based on BERT to classify tweets as medical and non-medical. We experimentally validate the proposed model on more than 14k tweets, reaching accuracy levels of 0.93.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Twitter is a social media platform where million of users discuss and write on
a daily bases about multiple topics. This large user base and the presence of
available APIs makes Twitter a useful data repository for researchers. A research
area that develops around Twitter consists in the categorisation of tweets, which
allows to identify their topic [
        <xref ref-type="bibr" rid="ref10 ref2 ref5 ref8">2, 8, 10, 5</xref>
        ].
      </p>
      <p>
        In this paper we propose to use machine learning models and in particular
BERT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] embeddings and MetaMap [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to classify tweets as belonging to the
medical or non-medical domain. We experimentally evaluate our approach on a
dataset of more than 14k tweets, which we release to the research community.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>We collected the data used for the experiments as follows: we manually
selected profiles of some sources (i.e., news websites, blogs, etc.) which publish
articles that are categorised by the editors such that they include a “health /
medical” category (or related ones). We considered the following sources of
information: IFLScience, CNN, NBC News, PBS, USA Today, and BBC News
(Science section). For such sources, we considered their official Twitter profile,
and we considered only the tweets that included a full statement / article and an
URL linking to the original domain; by exploring the categories on the original
domain we where then able to discriminate between medical and non medical
tweets. Let us make this process clear with an example. Let us suppose the
IFLScience Twitter account publishes tweets with their respective URLs in the
form iflscience.com/[topic]/[article url]; IFLS uses as “health-and-medicine” as
topic to identify medical related articles; thus, we consider such articles as
belonging to the medical domain, and the others to do not. We adopt a similar
approach for the other data sources.</p>
      <p>By using such scraping strategy, we collected 14,582 tweets, 2095 labelled
as being medical and the difference labelled as being non-medical. From such</p>
      <p>Roitero et al.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <p>data, we randomly extracted 500 medical and non-medical tweets as being our
test set. The data used to conduct all the experiments can be downloaded at:
https://github.com/KevinRoitero/twitterGoesToTheDoctor.</p>
      <p>
        In this work we process the text of the tweets and we use it as a feature to predict
the probability of a tweet as being part of the medical or non-medical domain. We
consider the following machine learning models: Logistic Regression [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which
fits the data to a regression using a logistic function; Random Forest [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], an
ensemble model based on decision trees classifiers; Naive Bayes [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a
probabilistic algorithm; and Support Vector Machines (SVM) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which places the target
classes in a multidimensional space and separate them with an hyper-plane.
      </p>
      <p>
        We feed such algorithms with two kind of features: starting from the text
of the tweets, we extract BERT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] embeddings, and MetaMap [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]terms. We
consider both the cases of using BERT embeddings and MetaMap terms alone, or
combining them together. BERT is an algorithm which computes the embeddings
of a text considering the words in relation to all the other words in a context
(e.g., a sentence). We use it to extract the embedding vectors of the text of the
tweets in our dataset. Metamap is a well known tool for extraction of medical
concepts from text. We use it to recognise, in Tweets, the presence of concepts
belonging to one of the 127 MetaMap semantic types, hot-encoded. When we
use the two techniques together we simply append the Metamap terms to the
BERT embeddings.
      </p>
      <p>BERT can be used both to extract embeddings or as a stand-alone
classification algorithm; we consider both cases. When we use it as stand-alone classifier,
we start from the pre-trained model released by Google, and we perform 2 epochs
of fine-tuning on our training data.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results and Conclusion</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lang</surname>
            ,
            <given-names>F.M.:</given-names>
          </string-name>
          <article-title>An overview of MetaMap: historical perspective and recent advances</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <fpage>229</fpage>
          -
          <lpage>236</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Cotelo</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz</surname>
            ,
            <given-names>F.L.</given-names>
          </string-name>
          , Enr´ıquez,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Troyano</surname>
          </string-name>
          , J.:
          <article-title>Tweet categorization by combining content and structural knowledge</article-title>
          .
          <source>Information Fusion</source>
          <volume>31</volume>
          ,
          <fpage>54</fpage>
          -
          <lpage>64</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Liaw</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiener</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Classification and regression by randomforest</article-title>
          .
          <source>R news 2(3)</source>
          ,
          <fpage>18</fpage>
          -
          <lpage>22</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Quercia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Askham</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crowcroft</surname>
          </string-name>
          , J.:
          <article-title>Tweetlda: supervised topic classification and link prediction in twitter</article-title>
          .
          <source>In: Proceedings of the 4th Annual ACM Web Science Conference</source>
          . pp.
          <fpage>247</fpage>
          -
          <lpage>250</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Rish</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , et al.:
          <article-title>An empirical study of the naive bayes classifier</article-title>
          .
          <source>In: IJCAI 2001 workshop on empirical methods in artificial intelligence</source>
          .
          <source>vol. 3</source>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Suykens</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vandewalle</surname>
          </string-name>
          , J.:
          <article-title>Least squares support vector machine classifiers</article-title>
          .
          <source>Neural processing letters 9(3)</source>
          ,
          <fpage>293</fpage>
          -
          <lpage>300</lpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Tare</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gohokar</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sable</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paratwar</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wajgi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>: Multi-class tweet categorization using map reduce paradigm</article-title>
          .
          <source>International Journal of Computer Trends and Technology (IJCTT) 9</source>
          (
          <issue>2</issue>
          ),
          <fpage>78</fpage>
          -
          <lpage>81</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Wright</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          :
          <article-title>Logistic regression</article-title>
          . In: R.Yarnold,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Grimm</surname>
          </string-name>
          , L.G. (eds.) Reading and understanding multivariate statistics, p.
          <fpage>217</fpage>
          -
          <lpage>244</lpage>
          . American Psychological Association (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>An unsupervised framework of exploring events on twitter: Filtering, extraction and categorization</article-title>
          .
          <source>In: Twenty-Ninth AAAI Conference on Artificial Intelligence</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>