<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cultural micro-blog Contextualization 2016 Workshop Overview: data and pilot tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Liana Ermakova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Josiane Mothe</string-name>
          <email>josiane.mothe@irit.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philippe Mulhem</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jian-Yun Nie</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eric SanJuan</string-name>
          <email>eric.sanjuan@univ-avignon.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IRIT, UMR5505 CNRS, ESPE, Universite de Toulouse</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LIA</institution>
          ,
          <addr-line>Universite d'Avignon</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LIG, Universite de Grenoble</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Microblog Contextualization based on Wikipedia</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>RALI, Universite de Montreal</institution>
          ,
          <addr-line>Quebec</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>CLEF Cultural micro-blog Contextualization Workshop is aiming at providing the research community with data sets to gather, organize and deliver relevant social data related to events generating a large number of micro-blog posts and web documents. It is also devoted to discussing tasks to be run from this data set and that could serve applications.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        This initial task aimed at generating a short summary providing the background
information of a tweet to help a user to understand it following[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Given a
microblog announcing some cultural event, participants have to provide a short
5 All resources are available online:http://cmc.talne.eu
summary extracted from Wikipedia that provides -extensive -background about
this event. The summary must contain some context information about the event
in order to help answering questions of the form "what is this tweet about?"
using a recent cleaned dump of Wikipedia. The context should take the form of
a readable summary, not exceeding 500 words, composed of passages from the
provided Wikipedia corpus.
      </p>
      <p>Any open access resource could be used in addition to the data provided to
participants subject to describing it and providing a valid URL.
2.1</p>
      <sec id="sec-1-1">
        <title>Datasets</title>
        <p>A restricted set of public micro-blogs in English were collected from a set of
public on Twitter, all related to the keyword festival. The micro-blogs are in
UTF8 csv format with various elds. In this task, the tweets do not contain
URL. The other suggested tasks would use additional information.</p>
        <p>Unlike tweets, Wikipedia is under Creative Common license, and its contents
can be used to contextualize tweets or to build complex queries referring to
Wikipedia entities. We extracted from Wikipedia an average of 10 million XML
documents per year since 2012 in the four main twitter languages:- en, es, fr and
pt. -These documents reproduce in an easy-to-use XML structure the contents
of the main Wikipedia pages: title, abstract, section and subsections as well as
Wikipedia internal links. Other contents such as images, footnotes and external
links are stripped out in order to obtain a corpus easy to process by standard
NLP tools. By comparing contents over the years, it is possible to detect long
term trends
2.2</p>
      </sec>
      <sec id="sec-1-2">
        <title>Evaluation</title>
        <p>
          Following [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] the summaries would be evaluated according to informativeness
and readability.
        </p>
        <p>Informativeness is the way they overlap with relevant passages (number of
them, vocabulary and bi-grams included or missing). For each tweet, all passages
from all participants will be merged and displayed to the assessor in alphabetical
order. Therefore, each passages informativeness will be evaluated independently
from others, even in the same summary. Assessors will only have to provide a
binary judgment on whether the passage is worth appearing in a summary on
the topic, or not.</p>
        <p>Readability can only be accurately assessed by humans. A small panel of
scholars in humanities will have to evaluate readability for a pool of summaries
using on an online web interface. Each summary consists of a set of passages
and for each passage, assessors will have to tick four kinds of check boxes:
{ Syntax (S): tick the box if the passage contains a syntactic problem (bad
segmentation for example),
{ Anaphora (A): tick the box if the passage contains an unsolved anaphora,
{ Redundancy (R): tick the box if the passage contains redundant information,
i.e. information that has already been given in a previous passage,
{ Trash (T): tick the box if the passage does not make any sense in its context
(i.e. after reading the previous passages). These passages must then be
considered as trashed, and the readability of following passages must be assessed
as if these passages were not present.
3</p>
        <p>Cultural MicroBlog Search based on WikiPedia entities
Given a cultural entity as a set of WikiPedia pages (typically a set of places to
visit, artists to see on stage, festivals of interest etc.), the proposed task would
be to provide a double extensive summary of relevant microblogs from insiders
and outsiders. This task will involved two sub-tasks:
Task 2a: Retrieval of relevant microblogs for an entity (described by its wikipedia
page)
Task 2b: Summarization of the most informative tweets (and comparison to
manually built summaries)
3.1</p>
      </sec>
      <sec id="sec-1-3">
        <title>Micro-blog collection</title>
        <p>The document collection is provided to registered participants by ANR GAFES6
project and consists in a pool of more than 50M unique micro-blogs from di erent
sources with their meta-information as well as ground truth for the evaluation.</p>
        <p>The micro-blog collection contains among other sources, all public posts on
Twitter using the keyword festival since June 2015. These micro-blogs are
collected using private archive service based on streaming API7. The average of
unique micro-blog posts (i.e. without re-twitts) between June and September is
2; 616; 008 per month.The total number of collected micro-blog posts after one
year (from May 2015 to May 2016) is 50, 490, 815 (24, 684, 975 without re-posts).</p>
        <p>These micro-blog posts are available online on a relational database with
associated elds, among them 12 are listed in Table 1. The \Comments" row in
Table 1 gives some gures about the existing corpus.</p>
        <p>Because of privacy issues, they cannot be publicly released but can be
analyzed inside the organization that purchases these archives and among
collaborators under privacy agreement. CLEF 2016 CMC Workshop provided this
opportunity to share this data among academic participants. These archives can
be indexed, analyzed and general results acquired from them can be published
without restriction.
3.2</p>
      </sec>
      <sec id="sec-1-4">
        <title>Linked web pages</title>
        <p>66% of the collected micro-blog posts contain Twitter t.co compressed URLs.
Sometimes these URLs refer to other online services like adf.ly, cur.lv, dlvr.it,
6 http://anr-gafes.univ-avignon.fr/demo.html
7 https://dev.twitter.com/streaming/public</p>
        <p>Name
text
from user</p>
        <p>id
iso language code</p>
        <p>source
&lt;geo type,
geo coordinates 0,
geo coordinates 1&gt;</p>
        <p>Description Comments
text of the twitt 99% of the twitts contain a non empty text</p>
        <p>66% contain an external compressed URL
author of twitt (string) 62; 105 organizations among 11; 928; 952 users.
unique id of micro-blog total so far: 50; 490; 815 posts.
encoding of the twitt the most frequent tags: en (57%), es (15%),</p>
        <p>fr (6%) and pt (5%).
interface used for posting the twitt frequent tags: Twitter Web Client (16%)</p>
        <p>iPhone and Twitterfeed clients (11% each).
geolocalization</p>
        <p>triplet valued in 2:3% of the twitts.
ow.ly, thenews.uni.me and twrr.co.vu that hide the real URL. We used the
spider mode to get the real URL, this process can require several DNS requests.
The number of unique uncompressed urls collected in one year is 11; 580; 788
from 641; 042 distinct domains. Most frequent domains are: twitter.com (23%),
www.facebook.com (5.7%), www.instagram.com (5.0%), www.youtube.com (4.5%),
item.ticketcamp.net (1.1%) and g1.globo.com (1%)
4</p>
        <p>
          TimeLine illustration based on Microblogs
The goal of this task is to link the events of a given festival program to related
microblog posts. Such information is useful for attendees of festivals, for people
that are interested in knowing what happens in a festival, and for organizers to
get feedback[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>Microblog posts are provided with their timestamps, which are crucial as
a basis for the requested linking. However, such timestamps must be use with
care: they do not necessarily give accurate enough information (for instance in
the case of parallel sessions), or might even generate perturbations (microblogs
about one event may be posted before, during, or after the actual event).</p>
        <p>Participants would be required to provide, for each event of the program, the
10 best tweets based on their relevance and diversity. In this task, diversity is a
must because retrieving several times the same post is not bene cial in our case.
4.1</p>
      </sec>
      <sec id="sec-1-5">
        <title>Data</title>
        <p>Participants for this task would use a subset of the microblogs collection,
matching the months the targeted festivals were organized at (July and December
2015).</p>
        <p>In its tentative form, Festival programmes are provided in French: Two
French music festivals have been selected: the festival des vieilles charrues and
the transmusicales de Rennes. The timelines provided are selected subset of each
festival program: the organizers selected a subset of the whole festival program
(for each stage and time, list of artists playing).</p>
        <p>The participants would be free to use any additional data to provide results:
social (popularity, ) or not (knowledge bases, ); it should be described in the
related paper and speci ed when submitting the runs.
4.2
The evaluation would be carried out on selected parts of the program chosen by
the task organizers depending on the number of relevant tweets per event. The
evaluation measures planned would be recall/precision based. Several types of
runs will be proposed: time-only, content-only, time&amp;content.
5</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Conclusion</title>
      <p>Cultural Microblog Contextualization CLEF 2016 WorkShop aims at developing
processing methods for social media mining. Our focus is around festivals that
are organized or that have a large presence on social media. Micro-blogs linked
to an event make a dense, rich but very noisy corpus. Content is often imprecise,
duplicate or non-informative.</p>
      <p>We also envisage to provide an extra corpus of Images related to cultural
festivals in the world. This access would allow researchers in IR and NLP to
experiment a broad variety of multilingual microblog search techniques (WikiPedia
entity search, and automatic summarization). Extensive textual references would
be provided by scholars in humanities involved in the ANR GAFES project.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Heijnen</surname>
            , J., de Reuver,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouwman</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warnier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horlings</surname>
          </string-name>
          , H. :
          <article-title>Social Media Data Relevant for Measuring Key Performance Indicators? A Content Analysis Approach</article-title>
          . In Co-created E ective, Agile, and
          <source>Trusted eServices, Lecture Notes in Business Information Processing</source>
          , Vol.
          <volume>155</volume>
          , Springer Berlin Heidelberg,
          <volume>74</volume>
          {
          <fpage>84</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Rui</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Whinston</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Designing a Social-broadcasting-based Business Intelligence System</article-title>
          ,
          <source>ACM Trans. Manage</source>
          . Inf. Syst.,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <volume>2</volume>
          (
          <issue>4</issue>
          ):1{
          <fpage>19</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheung</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Understanding Twitter Usage: What Drive People Continue to twitt</article-title>
          .,
          <source>PACIS</source>
          ,
          <volume>92</volume>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>SanJuan</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P</given-names>
          </string-name>
          , Moriceau,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Tannier</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          ,
          <article-title>Overview of the INEX 2010 Question Answering Track (QA@INEX)</article-title>
          , in: S. Geva,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schenkel</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Trotman (Eds.),
          <source>INEX</source>
          , Vol.
          <volume>6932</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2010</year>
          , pp.
          <volume>269</volume>
          {
          <fpage>281</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moriceau</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tannier</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>SanJuan</surname>
          </string-name>
          , E. :
          <article-title>INEX Tweet Contextualization task: Evaluation, results and lesson learned in Information Processing</article-title>
          &amp; Management, in press,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Leskovec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Backstrom</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleinberg</surname>
          </string-name>
          , J. :
          <article-title>Meme-tracking and the dynamics of the news cycle</article-title>
          .
          <source>In Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining (KDD '09)</source>
          . ACM, New York, NY, USA,
          <volume>497</volume>
          {
          <fpage>506</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>