<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>New online tools and digital environments for translation into emoji Johanna Monti</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Johanna Monti</string-name>
          <email>jmonti@unior.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Benjamin EPFL Lausanne</string-name>
          <email>martin@kamusiproject.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Switzerland</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sina Mansour EPFL Lausanne</string-name>
          <email>mansour@ee.sharif.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Switzerland</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Francesca Chiusaroli University of Macerata Italy</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>L'Orientale University Naples</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Emojitalianobot and EmojiWorldBot are two new online tools and digital environments for translation into emoji on Telegram, the popular instant messaging platform. Emojitalianobot is the first open and free Emoji-Italian and Emoji-English translation bot based on Unicode descriptions. The bot was designed to support the translation of Pinocchio into emoji carried out by the followers of the "Scritture brevi" blog on Twitter and contains a glossary with all the uses of emojis in the translation of the famous Italian novel. EmojiWorldBot, an off-spring project of Emojitalianobot, is a multilingual dictionary that uses Emoji as a pivot language from dozens of different languages. Currently the emoji-word and word-emoji functions are available for 72 languages imported from the Unicode tables and provide users with an easy search capability to map words in each of these languages to emojis, and vice versa. This paper presents the projects, the background and the main characteristics of these applications.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Emojitalianobot e
EmojiWorldBot sono due applicazioni
online per la traduzione in e da emoji
su Telegram, la popolare piattaforma
di messaggistica istantanea.
Emojitalianobot `e il primo bot aperto e gratuito
di traduzione che contiene i dizionari
Emoji-Italiano ed Emoji-Inglese basati
sule descrizioni Unicode. Il bot `e stato
ideato per coadiuvare la traduzione di
Pinocchio in emoji su Twitter da parte
dei follower del blog Scritture brevi e
contiene pertanto anche il glossario con
tutti gli usi degli emoji nella traduzione
del celebre romanzo per ragazzi.
EmojiWorldBot, epigono di Emojitalianobot,
`e un dizionario multilingue che usa gli
emoji come lingua pivot tra dozzine
di lingue differenti. Attualmente le
funzioni emoji-parola e parola-emoji
sono disponibili per 72 lingue
importate dalle tabelle Unicode e forniscono
agli utenti delle semplici funzioni di
ricerca per trovare le corrispondenze in
emoji delle parole e viceversa per
ciascuna di queste lingue. Questo
contributo presenta i progetti, il background
e le principali caratteristiche di queste
applicazioni.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        Emojitalianobot 1 and EmojiWorldBot 2 are two
new translation bots3 into and from emoji.
These two bots were designed starting from the
hypothesis of setting up an emoji multilingual
dictionary and translator through a process of
selection and assessment of conventional
semantic values. Translation cases may show
how images can convey common and universal
meanings, beyond specific peculiarities, so as
they can stand as models in the perspective of
an interlanguage
        <xref ref-type="bibr" rid="ref4">(Chiusaroli, 2015)</xref>
        . The two
1https://telegram.me/emojitalianobot/
2https://telegram.me/emojiworldbot
3Computer programmes that carry out repetitive
tasks and in their more sophisticated form can also
simulate human behaviours.
bots ease the use of emojis but also collect,
refine and make available valuable linguistic data
by means of crowdsourcing and gamification
approaches.
      </p>
      <p>This contribution presents the
state-of-theart concerning the use of crowdsourcing and
gamification approaches to linguistics in
section 2, the Emojitalianobot and the Pinocchio
project in section 3, the EmojiWorldBot in
section 4 and finally conclusions and future work
in section 5.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Crowdsourcing and gamification</title>
      <p>
        Crowdsourcing, i.e., the act of a company or
institution taking a function once performed
by employees and outsourcing it to an
undefined (and generally large) network of people
in the form of an open call
        <xref ref-type="bibr" rid="ref5">(Howe, 2006)</xref>
        is
becoming a widespread practice on the
Internet to develop linguistic resources
(dictionaries, glossaries, translation memories, etc.) or
services (translation, localisation, fansubbing,
etc.)
        <xref ref-type="bibr" rid="ref8 ref9">(Monti, 2012, 2014)</xref>
        . It allows the large
scale involvement of users who contribute with
their knowledge, their ideas, and their skills,
in this way performing an active role in the
achievement of a common goal.
Crowdsourcing can be used for the creation, maintenance
and sharing of lexical/terminological data such
as: i. lexical resources for online dictionaries,
e.g., Wiki platforms such as Wiktionary4 and
Omegawiki5, and recent forays by more
traditional dictionary publishing companies like
Collins, Oxford, and Macmillan; ii.
terminological resources for online terminological
databases, like TermWiki6 , the terminological
counterpart of Wiktionary or TaaS7; iii. lexical
and semantic resources for Natural Language
Processing (NLP) tasks, such as Word Sense
Disambiguation (WSA), Sentiment Analysis,
Computer Aided Translation, Machine
Translation and so on, using platforms for
distributing parts of large development projects to
professional or occasional lexicographers such as
Mechanical Turk8. To the best of our
knowledge only very few projects so far have been
tailored to mobile devices to gather linguistic
data in the field, (i) to collect dialect data as in
Dialectbot9, (ii) to document endangered
lan4https://en.wiktionary.org/
5http://www.omegawiki.org/Meta:Main_Page
6http://it.termwiki.com/
7https://term.tilde.com/
8https://www.mturk.com/mturk/welcome
9https://telegram.me/dialectbot/
guages as in Aikuma10 and Ma Iwaidja11, or
(ii) to gather grammaticality judgments
        <xref ref-type="bibr" rid="ref6">(Madnani et al., 2011)</xref>
        . The social dimension of
these types of activities is sometimes connected
and fed by social communities, where users
discuss problems, give suggestions, and exchange
ideas
        <xref ref-type="bibr" rid="ref3 ref7">(Brabham, 2012; McGonigal, 2011)</xref>
        . In
order to loyalize social communities and
improve their engagement, gamification is used
very often. The use of games is a very
effective tool for active participation since it
provides a strong motivational framework which
pushes people to act for good. Some
effective uses of games are to create new habits
or modify wrong actions. Wang et al. (2013)
list Games with a purpose (GWAPs)12 among
the different types of crowdsourcing. Some
good examples of games with a purpose in
the lexicographic field are Phrase Detectives13
and JeuxDeMots14. The main advantage of
GWAPs is their high attractiveness, because
people love playing games and it is easier to
obtain their contribution in this way in
comparison to other forms of crowdsourcing. The
difficulty in designing such games is to match
attractiveness with usefulness, i.e. an
attractive game which produces valuable data.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Emojitalianobot and the</title>
    </sec>
    <sec id="sec-5">
      <title>Pinocchio project</title>
      <p>Emojitalianobot is the first open and free
Emoji-Italian translation bot on Telegram.
It was developed to support the translation
project of Pinocchio in emoji 15 launched on
Twitter in February 2016 by F. Chiusaroli, J.
Monti and F. Sangati. The translation of the
famous children’s novel was carried out by the
followers of the Scritture brevi blog16 (by F.
Chiusaroli and F.M. Zanzotto) and the first
fifteen chapters have been translated, which
correspond to the original novel published by
Collodi in 1881. Every day tweets with sentences
taken from the novel were posted on Twitter
and the followers suggested their translations
10http://www.aikuma.org/aikuma-app.html
11https://itunes.apple.com/au/app/
ma-iwaidja/id557824618?mt=8</p>
      <p>12When a player without any special knowledge is
put into a gaming environment and has to make
decisions to win the game under the pressure of time or
any game mechanics’ constraints.</p>
      <p>13https://anawiki.essex.ac.uk/
phrasedetectives/
14http://www.jeuxdemots.org/jdm-accueil.php
15http://www.treccani.it/lingua_italiana/
speciali/ludolinguistica/Chiusaroli.html
16https://www.scritturebrevi.it/
in emoji; at the end of each day, the official
version of the translations was validated and
published.17 Translators used Emojitalianobot
that contains (i) the Emoji-Italian dictionary,
(ii) the Emoji-English descriptions based on
Unicode and (iii) a glossary with all the uses
of emoji in the translation of Pinocchio. The
project was associated with the Emojitalia
discussion group on Telegram, where users met to
discuss problems, solutions, suggest
improvements of the bot, in addition to the
translation choices for Pinocchio and communicate
in emoji. The Pinocchio translation project
therefore allowed to crowdsource different
linguistic data connected with the use of emojis
as actual means of communication and not just
simple graphics to express amusement or
interest. In this respect the main findings of the
project are twofold: the need to recur to
compound multi-emoji expressions in order to
express concepts which are not represented in the
current set, as well as a related simple
grammar to express syntactic relations among
emojis, past and future tenses, etc. Unlike previous
literary translation project in emojis, such as
the translations of Moby Dick or Alice in
Wonderland, this is the first attempt of a collective
shared emoji code (vocabulary and grammar)
based on a word for word translation totally in
emojis. Emojitalianobot is an ideal test bench
to experiment with new approaches like
crowdsourcing and gamification in the field of
Natural Language Processing (NLP). The Pinocchio
project, games and features available in the bot
to learn or guess the meaning of emoji are
devised indeed both to enjoy the bot while using
it and at the same time to give the
opportunity to users to develop linguistic descriptions
of emoji tailored on their actual perceptions.
The most important reward for playing with
the bot is the awareness of helping develop
a linguistic resource for one’s mother tongue,
and the pride in contributing to it.</p>
      <p>Since its release on Telegram, the project
was an instant success, becoming a viral web
phenomenon thanks to the Scritture brevi
community and the Pinocchio translation in
emojis, so that the bot has now almost 750 users.
The Pinocchio translation project in emojis
counts 611 tweets , 980 glossary entries which
correspond to 2127 words, of which 185 are
multi-emojis, i.e. compound emojis, such as
17The translation of Pinocchio in emoji can be
followed on Twitter using #emojitaliano.</p>
      <p>for the Italian word peggio (worst).
4</p>
    </sec>
    <sec id="sec-6">
      <title>EmojiWorldBot</title>
      <p>On the basis of (both linguistic and
technological) experience with Emojitalianobot, the three
Italian researchers together with Martin
Benjamin and Sina Mansour of the Kamusi Project
International18 and EPFL (Switzerland)
designed a new bot on Telegram in April 2016:
EmojiWorldBot, a multilingual dictionary that
uses Emoji as a pivot language from dozens of
different languages. Currently the emoji-word
and word-emoji functions are available for 70
languages imported from the Unicode tables 19
and provide users with an easy search
capability to map words in each of these languages to
emojis, and vice versa. Looking at the
UNICODE descriptions (see Fig. 1) it is apparent
that emojis are not annotated in a coherent
way across languages, so some languages have
more descriptions and some others, especially
underrepresented languages, have less or in the
most cases some languages are not represented
at all.</p>
      <p>Our first goal with EmojiWorldBot is
therefore to reach a uniform and comprehensive list
of tags across multiple languages with a precise
mapping between any language pair, which
may serve to bootstrap a massive multilingual
dictionary. The bot currently features:
emoji-to-word and word-to-emoji
translation for more than 70 languages
Eggs, a tagging game for people to
contribute to the expansion of these
dictionaries or the creation of new ones for any
additional language. Users can suggest
additional tags for single emojis in any
language (for example adding egg to the
tag list for in English).
18https://kamusi.org/
19http://www.unicode.org/cldr/charts/29/
annotations/
inline queries: type EmojiWorldBot and a
word, and it will suggest a set of emojis for
that word you can send in any Telegram
conversation
the possibility to add new languages.To
date 56 new languages were added, such as
Latin, Esperanto, Sardinian among
others.</p>
      <p>The basic idea of the Eggs game is to collect
new tags to associate with emojis as shown in
Fig. 2.</p>
      <p>
        With fewer than 2000 official emojis,
stretching the boundaries of their
communicative potential makes them more useful.
However, it also makes the dictionary more
essential, so that someone who receives in a chat
in any language might look to see if it
signifies something other than an eggplant. In
the future, Eggs will experiment with
multiemoji terms (METs), building on the work of
the Pinocchio translation project to Emoji, in
an effort to build a larger pictorial
vocabulary that is comprehensible across languages
        <xref ref-type="bibr" rid="ref4">(Chiusaroli, 2015)</xref>
        . A new version of the bot
is already under development. It will feature
Ducks, a second game where users are asked
to map tags from a source language (e.g.
English) to a target language (e.g. Swahili). In
the example of Figure 1, several Romanian
users would be shown the sense-specific
definition of grin from Wordnet and all of the
emojis that have been attached to that definition,
and be asked which of the options among fat,˘a
ˆıncaˆntata˘ fat,˘a and ˆıncaˆntare, if any, is a good
translation. The game would also be played
for face and grinning face. In this way, all of
the one-to-one relationships should be
discovered, and all instances of a term that does not
have a translation equivalent on the other side
will be revealed. When it is known that no
match from English exists, Ducks presents the
definition, the emojis, and the English term,
and asks the user to type in the best
equivalent in their language. This is the method
that will be most efficacious for new languages,
bypassing the need to disentangle the
manyto-many associations introduced through term
clustering in the CLDR annotations. It should
be noted that many terms will be removed
from the game cycle through comparisons with
Wordnets for available languages. For
example, самолет appears in conjunction with
English airplane in both the Emoji annotations
and the Bulgarian Wordnet that are linked to
the same English Princeton Wordnet (PWN)
sense, which gives sufficient confirmation
without needing a mass of human players. As of
this writing, the project is in the process of
importing and aligning Wordnet data for some 50
languages. In future work, terms from
Wordnet synsets will be tested against the emojis
with which they theoretically share a sense,
e.g. asking crowd members whether applies
to other members of the PWN synset for bus
(autobus, coach, jitney, motorbus, etc.), but
the mechanism for doing so has not been
finalized. In this way EmojiWorldBot employs
crowd methods as part of an arsenal intended
to conquer the walls of collecting data for
numerous diverse languages. Data validation will
be achieved via a consensus model through
which answers are accepted as correct if the
same result is provided by a threshold number
of respondents. The new version of the bot will
allow to:
add new terms to the current languages
(including the names of the countries for
national flags)
compare definitions across languages.
From the computational point of view, this
project, as the Emojitalianobot, attempts to
address the data chasm for natural language
processing for most languages by distilling
data collection to simple micro-tasks
        <xref ref-type="bibr" rid="ref1">(Benjamin, 2015)</xref>
        using techniques adapted to
leastcommon-denominator technology.
5
      </p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>
        We described the Emojitalianobot and the
EmojiWorldBot projects. Combining
crowdsourcing, gamification a nd a s martphone app
is a powerful strategy to collect, improve and
refine v aluable l inguistic d ata e asily a nd i n a
short time particularly for less-resourced
languages
        <xref ref-type="bibr" rid="ref2">(Benjamin and Radetzky, 2014)</xref>
        .These
may be the first crowdsourcing projects of this
type to use bots for linguistic data collection
and validation and are unique in their
attempts at engaging participants for different
languages.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Benjamin</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Crowdsourcing microdata for cost-effective and reliable lexicography</article-title>
          .
          <source>In Proceedings of AsiaLex 2015 Hong Kong, EPFL-CONF-215062</source>
          , pages
          <fpage>213</fpage>
          -
          <lpage>221</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Benjamin</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paula</given-names>
            <surname>Radetzky</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Multilingual lexicography with a focus on less-resourced languages: Data mining, expert input, crowdsourcing, and gamification</article-title>
          .
          <source>In 9th edition of the Language Resources and Evaluation Conference</source>
          , EPFL-CONF200375.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Daren C Brabham</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A model for leveraging online communities</article-title>
          .
          <source>The participatory cultures handbook</source>
          ,
          <volume>120</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Francesca</given-names>
            <surname>Chiusaroli</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>La scrittura in emoji tra dizionario e traduzione</article-title>
          .
          <source>CLiC it</source>
          , page
          <volume>88</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Howe</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The rise of crowdsourcing</article-title>
          .
          <source>Wired magazine</source>
          ,
          <volume>14</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Nitin</given-names>
            <surname>Madnani</surname>
          </string-name>
          , Joel Tetreault,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Chodorow</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Alla</given-names>
            <surname>Rozovskaya</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>They can help: Using crowdsourcing to improve the evaluation of grammatical error detection systems</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies: short papersVolume 2</source>
          , pages
          <fpage>508</fpage>
          -
          <lpage>513</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Jane</given-names>
            <surname>McGonigal</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Reality is broken: Why games make us better and how they can change the world</article-title>
          .
          <source>Penguin.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Johanna</given-names>
            <surname>Monti</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Translators' knowledge in the cloud: The new translation technologies</article-title>
          .
          <source>In International Symposium on Language and Communication: Research Trends and Challenges(ISLC).</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Johanna</given-names>
            <surname>Monti</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Dictionaries in the cloud: state of the art, trends and challenges</article-title>
          .
          <source>Les Cahiers du dictionnaire</source>
          , (
          <volume>6</volume>
          ):
          <fpage>95</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Aobo</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <source>Cong Duy Vu Hoang, and MinYen Kan</source>
          .
          <year>2013</year>
          .
          <article-title>Perspectives on crowdsourcing annotations for natural language processing</article-title>
          .
          <source>Language resources and evaluation</source>
          ,
          <volume>47</volume>
          (
          <issue>1</issue>
          ):
          <fpage>9</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>