<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Smoking Cessation Causes: Contrasting Evidence from Social Media and Online Forums</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Foaad Farooghian</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mourad Oussalah</string-name>
          <email>Mourad.Oussalah@ee.oulu.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Oulu, Centre for Ubiquitous Computing</institution>
          ,
          <addr-line>Computer Science, PO Box 4500, 90014 Oulu</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>iHR, Faraday Wharf</institution>
          ,
          <addr-line>Innovation Birmingham Campus Holt Street, Birmingham Science Park Aston, Birmingham, B7 4BB</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <fpage>34</fpage>
      <lpage>40</lpage>
      <abstract>
        <p>An automated-based system has been developed in order to gather stories from iCanQuit. Similarly, using Twitter Streaming API, geolocated tweets within UK region have been collected during a three months period, and those related to smoking cessation, according to the semantic text matching, are extracted and stored in a database. An automated classi er has been employed to identify the four most dominant categories in Health eld for each document of blogs or Twitter dataset. A total of 880 stories from iCanQuit and 22155 relevant tweets have been collected and indexed in SQLite database. The automated classier highlighted four categories: Weight Loss, Mental Health, Addiction, Support Group. The analysis surprisingly reveals that both blogs and Twitter datasets agree that the dominant source of smoking cessation is related to weight loss (shape body appearance), or body-look, while the support group, which includes any clinician supports, plays little impact on the smokers' quit motivation, a result that maybe precious for health authorities.</p>
      </abstract>
      <kwd-group>
        <kwd>blogs</kwd>
        <kwd>Twitter</kwd>
        <kwd>health</kwd>
        <kwd>smoking</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Tobacco remains one of the prime causes of death worldwide causing more than
5 million deaths annually in 2012 and expected to reach 10 million by 2030
according to world health organization [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] in additional to tremendous
economical costs. This motivates the growing smoking cessation initiative programs.
In the age of E-generation, Web-based smoking cessation programs have been
found to raise smokers' consciousness about quitting and encourage them to
take necessary actions [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Indeed, online social networks o ered a new way to
interact with smokers and in uence their behaviour where social network
intervention may work through multiple mechanisms, including social support,
information transfer, social in uence, modelling, and the transmission of social
norms [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Several individuals, especially teenagers, are more vulnerable into
following their mates' routes and reasoning rather than bothering to discuss
detailed personal experiences with clinical experts. Besides, it is commonly
acknowledged that many individuals feel more comfortable when interacting via
text or online messages instead of physical one-to-one meetings. For instance,
several mobile applications have been promoted for this purpose (e.g., I QUIT,
Quint O Meter, Quit With Me, UbQUITous), that would ease the interaction
of smokers willing to quit [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], although their compliance with clinician
guidelines is debatable. Therefore, there is a recognized need for research on the use
of social media to promote health behaviors and social support [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Platforms
like QuitNet (www.quitnet.com) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] have attracted millions of smokers
seeking help, share of stories, alternative remedies, among others, mainly for the
purpose of smoking cessation. Similarly, iCanQuit (www.icanquit.com.au), is
another popular online service provided by Australia cancer institute NSW since
2010 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. It is characterized by its easy access and well-structured documents,
with many classes as opposed to QuitNet. The platform also o ers the possibility
of retrieving past data from users and clinicians. Finally, the expansion of the
social media tools, e.g., Twitter, Facebook, Flickr provides the analyst with a huge
amount of data related to daily behaviour of the individuals /smokers, which can
be e ciently employed for behavioural analysis of the smoker (s), and thereby,
prescribe appropriate remedies accordingly. Therefore, the analysis of blogs and
other social network data related to smoking cessation opens a new door for
researchers to understand the reasons that motivate people to smoke. This also
enables scientists and/or clinicians to identify new symptoms or e ects during
smoking cessation which could result in possibly more e cient treatments. This
paper presents a blog and a social media related study that focused on symptoms
and/or origin of smoking behaviour/cessation by investigating data issued from
both Twitter Streaming API and the specialized smoking quit platform
IcanQuit. In both cases, parsers and crawler were employed to collect the messages.
A textual analysis of the messages is carried out in order to identify the cause
(s) of the smoking behaviour using a classi cation like strategy. Comparison of
the classi cation results in both sources (Twitter and blogs) is carried out in
order to identify key milestones.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <sec id="sec-2-1">
        <title>Data collection</title>
        <p>
          Two sources of information have been employed for data collection and
analysis. The rst one consists of dataset gathered from IcanQuit blogs. Interestingly,
this website has already some predisposition of smoking quiet stories with
respect to a set of prede ned categories: health, money, tness, family and others.
In contrast to data that can be retrieved from QuitNet or other blog sources, the
stories in IcanQuit are well-structured texts with a title and sometimes (sub)
sections. A total of 880 stories have been collected. Fig. 1 provides an example of
the website con guration as well as an example of database outputted by
iCanQuit query related cessation story where the relevant links are therefore stored
in a SQLite database. The second source of information uses Twitter Streaming
API [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to collect geolocated tweets in UK region for a speci ed time interval.
Next, the collected Twitter database has been ltered using smoking and
quieting related terms in the text message eld. More speci cally, Twitter dataset
has been collected over the period August - November 2014, and then ltered
such that any relevant tweet contains at least one \smoking" related terms and
one \quiet" related terms. The former includes terms like: smok, cigarette, cig,
tobacco, cannabis, pipe, shisha, waterpipe, hookah, and weed (inspiring from
urban dictionary). While quiet related terms comprise both \quiet" equivalent
terms and commonly employed smoking cessation products, e.g., quiet,
cessation, cess, stop, nicotin, patch, lozeng, chantix, counselling, therapy. A total of
3245567 tweets have been collected among which 22155 have been found to t
the smoking cessation context using the above methodology. These tweets are
then stored in a SQLite database.
        </p>
        <p>
          Unlike structured story documents from iCanQuit platform, Twitter dataset,
due to the size restriction (140 characters at most), are dominantly noisy with a
lot of slang words, abbreviations, spam and links, which require a special
methodology to capture relevant information. Inspired from our previous work [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
special consideration has been given to text pre-processing of textual Tweet
messages. Especially, Appache Nutch Crawler [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] was employed in order to crawl the
link extracted from iCanQuit website. The extracted text, usually constituted
of user's experience and story, is parsed using open source Stanford Parser and
then indexed using Apache Lucene [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] in order to bene t from its highly
scalable implementations and advanced search capabilities. The created index les
are stored in SQL like database that eases the compatibility with other software
resources.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Evidence classi cation</title>
        <p>
          The text tokens outputted from previous stage using either iCanQuit or
Twitter are classi ed according to a set of prede ned classes. For this purpose, the
UClassier API [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] was employed. This implements an improved semi-supervised
nave Bayes classi er with an extensive training corpus issued from yahoo open
directory project (http://www.dmoz.org). We restricted to health topics of the
directory because of its relevance to smoking cessation. Especially, we con ne
our analysis to the four classes which exhibit high score for most of the
input dataset. This corresponds to: Mental Health, Weight Loss, Addiction and
Support Group. Especially, a given document is deemed to belong to a
speci c category if the associated classi cation score according to UClassi er is the
highest among other categories and is greater than some pre-de ned threshold
(a 10%. threshold is found to give satisfactory results). The use of threshold is
motivated by the existence of documents related to smoking cessation but they
do not contain any argument that would allow the system or, even any expert
who reads the document, to identify any possible causes for smoking cessation.
On the other hand, the analysis of dataset involves two main phases. The rst
phase examines the relationship between di erent classes based on the
Uclassier classi cation by retrieving the score attached to each document. Next, the
score of each class is recorded, and the correlation between each pair of the four
aforementioned classes is examined.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Data Correlation Analysis</title>
        <p>
          A nal stage consists in a statistical analysis of evidence issued from the
classi ers. For this purpose, the correlation matrix among the various classes is
calculated for both iCanQuit and Twitter datasets. On the other hand, the
evaluation of the extent to which the two datasets support the same evidence is also
quanti ed using statistical testing. More speci cally, Cramer's phi correlation
test [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is employed to evaluate the (global) correlation between Twitter and
iCanQuit dataset, while scatter graphs and Pearson's product correlation coe
cient is employed to quantify the correlation among any pair of categories using
either Twitter or iCanQuit dataset. The evaluation of the result of the
automated classi cation against manually labelled classi cation is performed using
groups of random selection of dataset and accuracy classi cation metric.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussions</title>
      <p>
        Given that not all collected dataset contain enough clues to enable the system
to generate the cause of the smoking cessation, it is worth pointing out the
proportion of data where such evidence is occurring. For this purpose, Table 1
(see rst column) shows the number of blogs and tweets falling in each of the
four categories. The results in terms of total number of blogs and tweets in the
four categories indicate that 53:7% ((273 + 86 + 72 + 42)=880) of retrieved blogs
and only 17% ((1730 + 903 + 906 + 234)=22155) of retrieved tweets are classi ed
to one of the four categories. Trivially, as expected, larger this to occur more
often in Twitter dataset because of wording size. Assuming the outcomes of blogs
and Twitter dataset lie on two random variables, Spearman's rank correlation
coe cient can be used to quantify their correlations. In this respect, a very
strong correlation is observed = 0:93 with p &lt; 0:07. Table 1 indicates a strong
dominance of Weight Loss or, by abuse, body look as the main catalyst for change
of smoking behaviour including smoking cessation. This is also a well-known
fact in tobacco community as the e ect of nicotine increases metabolic rate in
the human body, which, in turn, suppresses appetite and yields weight loss.
While smoking cessation is likely to induce a gain of weight [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In order to
investigate the relationships between pairs of categories while accounting for the
sensitivity of the scoring function of the classi er, we created nine equal partition
of the (normalized unit interval) of the scoring function in [0.1 1], and record
the number of blogs (resp. Tweets) whole classi cation score of the underlying
category fails in the given subdivision. This also allows us to use the Pearson
product moment correlation coe cient to evaluate the correlation between any
pair of categories. In this course, Table 1 also records the correlation matrix
for both blog and Twitter dataset (the latter is in bold). The result indicates a
moderate correlation between categories Mental Health and Weight Loss as well
as between Support Group and Weight Loss while the correlation among other
categories is quite weak to negligible. In both scenarios, this demonstrates the
prevalence of body shape argument (weight loss) for clinical analysis as well as
the importance to ease the e ects of mental health di culty.
      </p>
      <p>Category
# blogs vs #Tweets Weight Loss Addictions Support Groups Mental Health
Weight Loss
273 vs 1730
Addictions
72 vs 906
Support Groups</p>
      <p>42 vs 243
Mental Health
86 vs 903</p>
      <p>On the other hand, it is also worth pointing out that there is a substantial
amount of dataset that cannot be classi ed to any of the aforementioned
categories, where about 46% of blogs and 82% of Twitter messages have not been
classi ed. This is mainly motivated by the nature of the dataset where many
of the messages and stories are too short or contain no useful information that
would allow any classi cation system to yield a speci c category. For instance,
messages \Yes I tried to quit cig", \this is a good story of smoking cessation
attempt" or \how was your day after rst quit?" provide little information for
any external user to output any tangible conclusion regarding the cause of
cessation. Similarly, there are a large number of tweets which are mainly generated
by commercial organizations to promote some speci c smoking cessation
products, e-cigarette, nicotine brands, etc. Some stories are also found to be related
to the subject through their title but the story itself is very short and can be
restricted to a web link only. Finally, in order to compare the accuracy of the
global category classi cation of Table 1 against manual check constituted by an
independent expert, we selected a set of random samples from the dataset and
computed the accuracy in the following way. For each category, we decomposed
the associated dataset into three (almost) equal groups. For instance, the 273
blogs corresponding to Weight Loss category are split into three equal groups of
91 blogs each. For each group, we randomly selected 10 documents, which are
manually checked, and then its accuracy is computed. This process of randomly
selecting 10 documents from each group, and then calculating the associated
accuracy, is repeated for each category. Table 2 summarizes the accuracy of these
subgroups when using blogs and Twitter dataset. Results highlighted in Table 2
demonstrate a good accuracy of the automated classi cation system when
compared to a manual expert-based classi cation. We also acknowledge a slightly
decreasing accuracy in case of Twitter dataset. This can be explained by the
di culty of the tweet message to be classi ed in one of the speci ed category
because of the size restriction which renders any manual classi cation rather a
di cult task and sometimes very subjective.</p>
      <p>G.1
Weight Loss 93%
Mental Health 95%</p>
      <p>Addiction 92%
Support Group 96%</p>
      <p>Blogs dataset</p>
      <p>G.2
96%
94%
98%
99%</p>
      <p>G.3
95%
94%
97%
95%</p>
      <p>Twitter dataset
G.1 G.2 G.3
87% 81% 83%
78% 82% 87%
84% 85% 80%
77% 76% 83%</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgment</title>
      <p>This work is supported by University of Birmingham School of Electronics,
Electrical and Computer Engineering as well as EPSRC GAP Project, which are
grateful for their nancial help in conducting this research, when the rst
author was with Birmingham University.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Apache</given-names>
            <surname>Nutch</surname>
          </string-name>
          <article-title>Crawler</article-title>
          . http://nutch.apache.org/.
          <source>Accessed: April</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Quitnet</surname>
          </string-name>
          , HELP FAQs. http://www.quitnet.com/help/helpfaq.jtml?
          <source>SubjectID=925. Accessed: April 9</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Twitter Streaming API. https://dev.twitter.com/streaming/overview. Accessed: April
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <article-title>UClassi er API</article-title>
          . http://www.uclassify.com/About.aspx.
          <source>Accessed: April</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. WHO,
          <article-title>Tobacco fact sheet N339</article-title>
          . http://www.who.int/mediacentre/factsheets/ fs339/en/index.html.
          <source>Accessed: August</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>B.</given-names>
            <surname>Borrelli</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Mermelstein</surname>
          </string-name>
          .
          <article-title>The role of weight concern and self-e cacy in smoking cessation and weight gain among smokers in a clinic-based cessation program</article-title>
          .
          <source>Addictive Behaviors</source>
          ,
          <volume>23</volume>
          (
          <issue>5</issue>
          ):
          <volume>609</volume>
          {
          <fpage>622</fpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>N. K.</given-names>
            <surname>Cobb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Graham</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. B.</given-names>
            <surname>Abrams</surname>
          </string-name>
          .
          <article-title>Social network structure of a large online community for smoking cessation</article-title>
          .
          <source>American Journal of Public Health</source>
          ,
          <volume>100</volume>
          (
          <issue>7</issue>
          ):
          <volume>1282</volume>
          {
          <fpage>1289</fpage>
          , 07
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. C. Esco ery, L.
          <string-name>
            <surname>McCormick</surname>
            , and
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Bateman</surname>
          </string-name>
          .
          <article-title>Development and process evaluation of a web-based smoking cessation program for college smokers: innovative tool for education</article-title>
          .
          <source>Patient Education and Counseling</source>
          ,
          <volume>53</volume>
          (
          <issue>2</issue>
          ):
          <volume>217</volume>
          {
          <fpage>225</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>A. M. Jacobs</surname>
            ,
            <given-names>O. C.</given-names>
          </string-name>
          <string-name>
            <surname>Cobb</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Abroms</surname>
            , and
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Graham</surname>
          </string-name>
          .
          <article-title>Facebook apps for smoking cessation: A review of content and adherence to evidence-based guidelines</article-title>
          .
          <source>J Med Internet Res</source>
          ,
          <volume>16</volume>
          (
          <issue>9</issue>
          ):e205,
          <year>Sep 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>M. McCandless</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Hatcher</surname>
            , and
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Gospodnetic</surname>
          </string-name>
          . Lucene in Action,
          <source>Second Edition: Covers Apache Lucene</source>
          <volume>3</volume>
          .0. Manning Publications Co.,
          <string-name>
            <surname>Greenwich</surname>
            ,
            <given-names>CT</given-names>
          </string-name>
          , USA,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>K. M. McElwaine</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Freund</surname>
            ,
            <given-names>E. M.</given-names>
          </string-name>
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Slattery</surname>
            ,
            <given-names>P. M.</given-names>
          </string-name>
          <string-name>
            <surname>Wye</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lecathelinais</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Bartlem</surname>
            ,
            <given-names>K. E.</given-names>
          </string-name>
          <string-name>
            <surname>Gillham</surname>
            , and
            <given-names>J. H.</given-names>
          </string-name>
          <string-name>
            <surname>Wiggers</surname>
          </string-name>
          .
          <article-title>Clinician assessment, advice and referral for multiple health risk behaviors: Prevalence and predictors of delivery by primary health care nurses and allied health professionals</article-title>
          .
          <source>Patient Education and Counseling</source>
          ,
          <volume>94</volume>
          (
          <issue>2</issue>
          ):
          <volume>193</volume>
          {
          <fpage>201</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>M. Oussalah</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Bhat</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Challis</surname>
            , and
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Schnier</surname>
          </string-name>
          .
          <article-title>A software architecture for twitter collection, search and geolocation services</article-title>
          .
          <source>Knowl.-Based Syst.</source>
          ,
          <volume>37</volume>
          :
          <fpage>105</fpage>
          {
          <fpage>120</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Sheskin</surname>
          </string-name>
          .
          <article-title>Handbook of Parametric and Nonparametric Statistical Procedures</article-title>
          .
          <source>Chapman &amp; Hall/CRC, 4 edition</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>J. L. Westmaas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Bontemps-Jones</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Bauer</surname>
          </string-name>
          .
          <article-title>Social support in smoking cessation: Reconciling theory and evidence</article-title>
          .
          <source>Nicotine &amp; Tobacco Research</source>
          ,
          <volume>12</volume>
          (
          <issue>7</issue>
          ):
          <fpage>695</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>