<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Microposts</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Sentiment Analysis of Wimbledon Tweets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Priyanka Sinha</string-name>
          <email>priyanka27.s@tcs.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anirban Dutta Choudhury</string-name>
          <email>@tcs.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amit Kumar Agrawal</string-name>
          <email>amitk.agrawal@tcs.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Tata Consultancy Services</institution>
          ,
          <addr-line>Limited, Ecospace 1B, New Town, Rajarhat, Kolkata 700156</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tata Consultancy Services</institution>
          ,
          <addr-line>Limited, Ecospace 1B, New Town, Rajarhat, Kolkata 700156, India, anirban.duttachoudhury</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>4</volume>
      <abstract>
        <p>Annotating videos in the absence of textual metadata is a major challenge as it involves complex image and video analytics, which is often error prone. However, if the video is a live coverage of an event, time correlated textual feed about the same event can act as a valuable source of aid for such annotation. Popular real time microblog streams like Twitter feeds can be an ideal source of such textual information. In this paper we explore the possibility of such correlation with the sentiment analysis of a set of tweets of the Roger Federer and Novak “Nole” Djokovic semi finals match at Wimbledon 2012.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Twitter</kwd>
        <kwd>Wimbeldon</kwd>
        <kwd>Sentiment Analysis</kwd>
        <kwd>TV</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>One way to understand the sentiment of people viewing or
experiencing the event is to analyze the video feed from TV
or web hosting sites like Youtube. It is a challenging
computationally hard problem especially without any helping
text. Sentiment annotation in videos at finer granularity is
not a much explored area. Sentiment annotations of a live
Copyright c 2014 held by author(s)/owner(s); copying permitted
only for private and academic purposes.</p>
      <p>
        Published as part of the #Microposts2014 Workshop proceedings,
available online as CEUR Vol-1141 (http://ceur-ws.org/Vol-1141)
video can be leveraged to enable targeted advertisement.
However, there is problem with annotation quality due to
the fact that manual video annotation is tedious and time
consuming process whereas automated supervised video
annotation is very limited in its coverage (i.e. incomplete and
sometimes wrong). [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is an example of how human
crowdsourcing is one of the possible ways to annotate videos.
Based on our experiments detailed in ”Our Approach”
section, we observe that there exists a correlation between
sentiment analysis on tweets and live coverage of video in real
time of a popular event, i.e., the Wimbledon semi final match
between Roger Federer and Novak Djokovic. This
correlation is observed on this particular event and the confirmation
of generality of the phenomenon is part of our future work.
We have been successful in doing text analysis on
microposts for sentiment analysis of a live event with respect to
participants in that event in real time which has not been
done before. We are proposing a novel approach for
automatic sentiment annotation of live coverage of videos related
to events affecting mass at large such as politics, natural
calamities, sports etc. For a widely well known event with
a large number of stakeholders, it is generally seen that the
traffic on Twitter is huge with lots of people tweeting about
it.
2. RELATED WORK
[
        <xref ref-type="bibr" rid="ref6 ref8">8, 6</xref>
        ] have demonstrated that twitter based sentiment
analysis can be used for closely predicting political election
results. However the approach is limited in temporal
correlation because the political event (i.e., gold standard) is
covered using various news flashes. It is not as fine grained and
accurate as capturing the video of the unfolding of political
events as they appear on either TV or web.
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is similar to our work in that they do discover named
entities in tweets and micro-events for live events, but it is
a different text mining task than sentiment analysis of the
event.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3. OUR APPROACH</title>
      <p>
        Tweets [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] have a maximum length of 140 characters. ”Come
on Federer! #Wimbledon” is an example tweet. RT is an
acronym for retweet. @ is used to mention a twitter user
name. # is used to represent a hashtag. http://bit.ly/9K4n9p
is a shortened URL linking to external content.
Figure 1 depicts the steps we take to find the correlation
between manually annotated sentiments of video and tweets.
We set up one linux desktop to capture tweets using [
        <xref ref-type="bibr" rid="ref3 ref4">4, 3</xref>
        ].
Since Twitter uses OAuth 2.0 as authentication to connect
to its API, we created an application and generated valid
oauth token and secret online for use directly without the
OAuth handshake. We grabbed the tweets for ”Wimbeldon”
keyword during the live telecast. In order to capture the
TV video of the live telecast of wimbledon semi final match
between ”Roger Federer” and ”Novak Djokovic”, we tuned
the Tata Sky set top box to Star Sports and attached a usb
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to it connecting to a linux laptop. We used mencoder
on the linux laptop with settings of aac for audio, h.264 for
video and mp4 as the mux. The tweets and the video capture
were started almost simultaneously thereby synchronizing
the starting timestamps.
      </p>
      <p>Three people manually annotated the video in two column
format where the first column was ”time in seconds since
start of match” and second column was either ”1” for positive
sentiment for Roger or ”2” for positive sentiment for Novak.
Majority voting was taken to create the ground truth. We
use supervised text classifiers such as Naive Bayes on tweets
for sentiment polarity detection. We trained the classifier
using part of the tweets which we manually annotated.
Finally we used the sentiments derived from analysis on tweets
and compared them to the ground truth. We found that the
tweet sentiments were correlated with the video, and the
time lag between video telecast and tweet was negligible.
4. CONCLUSION AND FUTURE WORK
When game sentiment is towards a particular player,
advertisements endorsed by that player can be shown. We can
also split the game into parts and get real time
summarization of the game sentiment upto that point or within a
time span.If intensity of sentiment is used then we can
detect peaks of sentiments towards players as well and can tag
best moments in the game as well.</p>
      <p>Futuristic applications include allowing V-Commerce on live
events like popular fashion shows where positive sentiment
towards a participant of a video can inform the backend to
adjust load towards possible increase in volume of incoming
purchases.</p>
      <p>Manual annotation to obtain the gold standard from video
is a tedious task but gives an accurate understanding of the
events sentiment. Video emotional analysis can be used to
augment this process.</p>
      <p>We understand that since this analysis has been done on one
event, similar analysis on more popular real life events would
help. Future work would also involve better techniques of
sentiment analysis taking into account the short and noisy
nature of tweets. Identification and treatment of languages
other than english would help for certain events such as the
Japanese tweets for Fukushima earthquake. Based on the
correlation that we find between tweets and live events in the
form of video, we are motivated enough to create a system
where automatic annotation of live coverage of an event will
take place using sentiment derived from tweets for the same
event.</p>
    </sec>
    <sec id="sec-3">
      <title>5. ACKNOWLEDGEMENTS</title>
      <p>We would like to thank Chirabrata Bhaumik and Avik Ghose
in helping us set up the video capture of the Wimbledon
semi-finals which gave us our gold standard. We would also
like to thank the workshop reviewers for their comments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] Dc 60. http://easycap.blogspot.in/p/easycap-dc60.
          <fpage>html</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Twitter</surname>
          </string-name>
          . https://dev.twitter.com/docs/platform-objects/tweets.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] Twitter streaming api</article-title>
          . https://dev.twitter.com/docs/api/1.1/post/statuses/filter.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <fpage>Twitter4j</fpage>
          . http://twitter4j.org/en/index.html.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Breslin</surname>
          </string-name>
          .
          <article-title>Extracting semantic entities and events from sports tweets</article-title>
          . In Making Sense of Microposts':
          <article-title>Big Things Come in Small Packages: co-located with the 8th Extended Semantic Web Conference</article-title>
          , ESWC2011, Heraklion, Crete, May
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Rao</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          .
          <article-title>Analyzing stock market movements using twitter sentiment analysis</article-title>
          .
          <source>In Proceedings of the 2012 International Conference on Advances in Social Networks Analysis and Mining (ASONAM</source>
          <year>2012</year>
          ),
          <source>ASONAM '12</source>
          , pages
          <fpage>119</fpage>
          -
          <lpage>123</lpage>
          , Washington, DC, USA,
          <year>2012</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Vondrick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Patterson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Ramanan</surname>
          </string-name>
          .
          <article-title>Efficiently scaling up crowdsourced video annotation</article-title>
          .
          <source>International Journal of Computer Vision</source>
          ,
          <volume>101</volume>
          (
          <issue>1</issue>
          ):
          <fpage>184</fpage>
          -
          <lpage>204</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Can</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kazemzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          .
          <article-title>A system for real-time twitter sentiment analysis of 2012 u.s. presidential election cycle</article-title>
          .
          <source>In Proceedings of the ACL 2012 System Demonstrations, ACL '12</source>
          , pages
          <fpage>115</fpage>
          -
          <lpage>120</lpage>
          , Stroudsburg, PA, USA,
          <year>2012</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>