<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>FootballWhispers: Transfer Rumour Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Neil Ireson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Ciravegna</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Named Entity Linking &amp; Rumour Detection</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>She eld University</institution>
          ,
          <addr-line>She eld</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Social media has been shown to have potential to predict various real world events, such as movements in the stock market and the outcomes of political elections. In this paper we present the Football Whispers (FW), a website dedicated to fans discussing transfer rumours. The unique selling point of the site is that it provides a crowdsourced assessment of those rumours, measuring the relative likelihood of a player's movements from social media chatter. This talk will focus on the rumour identi cation process, highlighting the role of open knowledge graphs and linked data to augment a domain knowledge-based to enable e ective Named Entity Linking in noisy, informal social media messages.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>2
46,631 name variations; with 17,489 (42%) players sharing a last name and 785
(2%) have identical names. In order to increase the number of alternative names
the Opta entities are mapped to Wikidata and DBpedia entities. Wikidata
contains 210,375 players and 30,710 teams, although a large number of the players
are inactive, in order to identify potentially ambiguity it is necessary to include
all names which may be mentioned. In addition to players, 11,439 managers,
pundits, referees, etc. are also extracted. The talk will describe the entity
mapping process, which considers the similarity of entities' available features, e.g.
string similarity of names and numerical distance of dates, with the importance
of a features being weighted according to the degree of variation in its values.
The mapping process was also applied to DBpedia, which contains 126,790
players; this resulted in a slight increase in alternative names, primarily due to the
DBpedia extraction of nicknames. In total 32,754 (80%) of Opta players are
mapped, and for these players name variations are tripled to 95,535.</p>
      <p>The team and player names are then used to generate a Deterministic Finite
Automata to e ciently extract candidate entity mentions from the message text.
The talk will describe how the contextual disambiguation processes are used
to link mentions to an entity instance, where other entity candidates in the
message provide the context. A name which occurs frequently in small number
of (expected) contexts (e.g. player name mentioned only with their team) is
deemed to maintain its meaning outside those contexts, while a name which
occurs in multiple (unexpected) contexts is deemed too ambiguous to be used
for entity linking when not contextualised. In addition, the message language is
also considered, as names can be ambiguous within a given language context.</p>
      <p>Evidence for a rumour is given by a message containing player and team
entities, and at least one transfer term. The talk will brie y outline the four
determinants of the veracity of a rumour: consensus (amount of evidence),
recency/constancy (evidence time decay), authority (evidence sources) and
coherence/consistency (evidence is not contradictory).
3</p>
    </sec>
    <sec id="sec-2">
      <title>Football Whispers</title>
      <p>In order to select tweets, which belong to the football domain, team names are
used to lter the messages, this results in between 1-2 million tweets per day, and
despite only English team names being used in the lter almost 56% of messages
are in other languages. The rumour detection system processes the messages
in real-time and the resultant rumours and likelihoods are validated by FW
experts before appearing on the website. The talk will present the evaluation
of the rumour detection and show how the use of knowledge graph (KG) data
has led to signi cantly increased player NEL performance, and identi cation
of the rumours concerning actual 2016 football transfers. Primarily this is due
to the availability of multilingual data in the KGs and the use of language
agnostic statistical disambiguation techniques in NEL. The success of FW (which
currently has two million users) has now led to the developed of Sports Whispers
and the application of this approach to other sporting domains.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>