<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Which user interaction for cross-language
information retrieval? design issues and reflections. J. Am. Soc. Inf. Sci. Technol.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Evaluating the Impact of Personal Dictionaries for Cross-Language Information Retrieval of Socially Annotated Images</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Diana Irina Tanase</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Epaminondas Kapetanios</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science, University of Westminster</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2006</year>
      </pub-date>
      <volume>57</volume>
      <issue>5</issue>
      <fpage>709</fpage>
      <lpage>722</lpage>
      <abstract>
        <p>These working notes focus on the users' actions in order to assist translations and on the usage of personal dictionaries (a feature which enables saving user added words). The special interest for this feature comes from a need to investigate to what extent users get actively involved in the query translation and contribute to overcoming the limitations of automatic translations. It is also our hope that by understanding the relationship between user language skills and the usage of the personal dictionary feature in the iCLIR context, we will be able to get at least a partial answer to a bigger question regarding collaborative translations in today's participatory web space. ACM Categories and Subject Descriptors: H.3.5 [Online Information Services]: Web-based services; H.3.3 [Information Search and Retrieval]: Search process; General Terms: Experimentation Free Keywords: personal dictionary, cross language information retrieval This year's iCLIR 2008 challenge was set to meet the need for a large-scale experiment, in a realistic naturally multilingual environment opened to all users in the web community. It comes as a follow-up to the iCLIR 2006 experiments that concentrated on developing a CLIR interface for Flickr (a widely used photo and video sharing service) and on devising evaluation measures to gauge the user experience. The 2006 task was completed by researchers from three universities, each investigating different aspects of the user-system interaction. The UNED group assessed how users with different language skills interacted with the system and if the CLIR facilities provided were used. The Swedish Institute of Computer Science group focused on user's perceptions namely satisfaction, completeness, and quality, while the University of Sheffield group focused both on user's behavior and search effectiveness. Their results unveiled a certain level of reticence in using the assisted query translation functionality, more so when the target language was unknown, with automatic translation being favored by users. All experiments were run with relatively small numbers of users (less than 25) and with homogeneous types of users in terms of language skills. To be able to draw broader conclusions regarding users' interaction with a CLIR system, this year's CLEF organizers made available for its participants a set of logs from a specially developed system with a CLIR front-end to the Flickr database. This system enabled monolingual and multilingual searches through a set of 180 images. The images were annotated in one or several of the following languages: English, Spanish, German, French, Dutch, and Italian, and it was promoted as a game, made available to any interested user. Hence, this year's experimental setting has a unique character by allowing researchers to learn from a larger and heterogeneous user group. In these working notes, the focus is set on the users' actions in order to assist translations and also on the usage of personal dictionaries (a feature which enables saving user added words). The special interest for this feature comes from a need to investigate further to what extent users get actively involved in the query translation and contribute to overcoming the limitations of automatic translations. Previous research in (Petrelli 2006, Oard 2008) has acknowledged that a modern iCLIR system should incorporate the necessary functionality to allow users to type their own translations. Comparative studies have been set up to gauge a user's willingness to supervise the translation step by selecting or deselecting from a list of potential translations. Results indicated that in supervised mode, when users verify and refine the translated query, the system has performed better than in delegated mode, when users do not intervene in the translation process. Though differences were not statistically significant in terms of precision and recall. It was also discovered that the supervised mode helped some of the users to reformulate their initial query based on suggested translations (Petrelli 2006). This search pattern was observed by the experiments with MIRACLE (He 2007).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>It is also our hope that by understanding the relationship between user language skills and the usage of the
personal dictionary feature in the iCLIR context, we will be able to get at least a partial answer to a bigger
question regarding collaborative translations in today’s participatory web space. If ad-hoc web communities
form to share information, communicate on events, stories, things to-do, and overall facilitate each other to
find and identify relevant resources, would it be possible to trigger such informal collaboration between
users with different language skills and facilitate the identification of relevant web resources regardless of
the document language? A starting point in investigating this question is determining to what extent users do
take a participatory role in creating personal dictionaries and what is the quality of these dictionaries. Ideally,
these customized language resources could be shared inside multilingual web communities and capture
upto-date usage of languages.</p>
      <p>Until the above hypothesis gains more weight, we will focus on describing the context and the experimental
setup (Section 2) that generated the logs, the research questions we tried to answer while analyzing and
interpreting the logs (Section 3). We will conclude with a discussion (Section 4) of the answers crystallized
from our iCLEF 2008 participation.</p>
    </sec>
    <sec id="sec-2">
      <title>2. The Flickling Game</title>
      <p>The system created for the 2008 evaluation (Flickling) implements a baseline set of functionalities for a
CLIR system. It relies on the Flickr API for the image retrieval task and on a term-to-term translation
mechanism. The latter relies mainly on a set of free dictionaries, and on a selection algorithm for “best”
target translation based on i) Flickr’s related terms for the query (often multilingual) and on ii) string
similarity between the source and the target words (Clough 2008). It is worth mentioning that the dictionaries
selected for this task have limited coverage, and uneven sizes. In other words, Flickling was equipped with a
basic set of language resources.</p>
      <p>The game entails finding images by determining the correct query terms to describe a given image and
suitable translations for these query terms when needed. The clear search task and setup of this system,
attracted around 300 users, but after filtering the data only 176 have actually played the game. The
participating users filled in a short pre-game questionnaire specifying their mother language, the active
languages (fluent writing and reading), passive languages (some level of fluency) and unknown languages.
The results of these questionnaires show that the participants had a good range of language skills from
monolingual to polyglots.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Analysis of Log File Data</title>
      <p>The log files distributed to all the participants of this year’s evaluation recorded step-by-step the user’s
actions. In our analysis we focused on extracting the entries that related to the users’ language profiles, the
interactions with the translation mechanism, the addition of new entries in the personal dictionary and on the
overall user’s results for the game.</p>
      <p>Research Questions
1. Does the degree of confidence with a language affect usage and creation of personal dictionary entries, i.e.,
do those users with little knowledge of a language make use of the personal dictionary and to which extent?
2. Does the degree of confidence with a language affect quality of personal dictionary?
3. Can it be inferred that the user’s performance in the game results improved by using the personal
dictionary and/or the assisted translation mechanism?
4. Is the personal dictionary a useful interface facility?
Based on the active and passive language skills, we created a language skill coefficient that describes the
number of active languages and the number of passive languages. For example a user that specified “EN” as
the active and “FR, DE” as passive would be assigned a coefficient of 3 (a count of the known languages
giving equal weight to active and passive languages). This generated five classes of users with the following
distribution: 37% are in class 1, respectively class 3, 20% are in class 2, 4% are in class 4, and 2% in class 5.
This coefficient is a measure of user’s degree of confidence with languages needed to play the game and will
be used to classify the answers to the research questions above.</p>
      <p>Our explorations concerning the questions above, starts with an overall look into the results of two
questionnaires posed to the user during the game whenever an image was found or its search abandoned. This
will provide more insight on the challenges of the game. Tables 1 a) and b), as well as Tables 2 a) and b)
reflect the results grouped by language coefficient and the challenges of the Flickling game.</p>
      <sec id="sec-3-1">
        <title>Found Image Questionnaire</title>
        <p>What problems did you encounter while searching for this image?
0: It was easy
1: It was hard because of the size of the image set
2: It was hard because the translations were bad
3: It was difficult to describe the image
4: It was hard because I didn’t know the language in which the image was annotated
5: It was hard because of the number of potential target languages
6: It was hard because I needed to translate the query
The results above indicate that users found the image searching process easy (Q0), and the task became
harder when the size of the image set was large (Q3) (see language coefficient 1 and 4). Another challenging
aspect can be identified as the describing of the image itself. This stands out across all types of users. It is in
essence a classic problem of information retrieval regardless of the multilingual aspect. Previous iCLEF
evaluations did emphasize that query formulation and re-formulation have a strong impact on search results.
Bad translations and not knowing the language in which the image was annotated were subsequent problems.
The actual translation process was problematic for people with one active language.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Give Up Questionnaire</title>
        <p>Why are you giving up on this image?
0: There are too many images for my search
1: The translations provided by the system are not right
2: I can't find suitable keywords for this image
3: I have difficulties with the search interface
4: I just don't know what else to do
5: Other (please, comment below)</p>
        <sec id="sec-3-2-1">
          <title>Language Skill</title>
          <p>Coefficient</p>
          <p>1
log /
lang.
coef.
log
/lang
coef
1
170
206
72
150
91
236
12
23
16
21
41
335
18
204
43
284
3
0
32
37
Q2
44.14
40.54
48.92
37.14
166
210
90
132
160
167
13
22
22
15
45
16
27
4
8</p>
          <p>Q3
3.45
4.95
5.81
14.28
2.70
13
363
11
211
19
308
5
1
30
36
The results of the give up questionnaire complement the previous results. It is apparent that regardless of the
user profile in terms of language skills, it is the problem of describing the image with the correct keywords
that determined users to abandon an image search. A second issue for all users was the set of images to
search through. Users would acknowledge not knowing what to do next, to be more of a problem than
dealing with bad translation. The difference for each group of users between the answers of each question is
only marginal, but it surprisingly reflects that users that know only one language trusted the translation and
did not point out at translations as a major problem.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Interactions – Assisted Translations and Personal Dictionary</title>
        <p>The Flickling system was equipped with a straightforward interface that allowed users to look at the list of
translated words, to add or remove words from the suggested translation, and also to add their own dictionary
entries. Table 3 a) and b) describe in terms of distributions the actions taken by each group of users.
add transuggestion
remove transuggestion
total interactions
47
329
25
197
65
262
6
2
29
35</p>
        <p>Q4
12.5
11.26
19.87
17.14
14
362
26
196
22
305
1
0
34
37</p>
        <p>Q5
3.72
11.71
6.72
2.85
0
62
39
59
14
9
4.09%
35 / 18
Based on the results from Table 3b the main type of interactions with the Flickling system is viewing the
existing translations, followed by typing new translations. The interesting aspect of the data showing in this
table is that users with fewer language skills were quite active in terms of adding new translations. The most
skilled set of users were most active in selecting or deselecting words when translations were needed.
Using the observations from “type new translation” column (normalized by number of users) we computed
the correlation between usage of the personal dictionary and language skill (-0.5946), which indicates a
decreasing linear dependency between the two, and provides a negative answer to the first research question
in this section regarding the correlation between language skills and involvement in working with the
personal dictionary. This is a slightly surprising result, but it can be explained by experimental attitude
towards adding their new dictionary entries.</p>
        <p>We have also plotted a user vs. personal dictionary graph, to identify the overall trend in the user base
(Figure 1). Out of the 94 users (~ 53%) that did use the personal dictionary, the mean number of interactions
was 18, with a median of 5. We isolated the users that have added entries to their personal dictionary more
than 18 times, but the scatter plots did not show a positive correlation between overall score, precision, time
and the usage of the personal dictionary. This result may be due to the fact that we considered the overall
personal dictionary interaction and not individual image searches.
The log recorded a total of 460 new entries to the Personal Dictionaries. By individually analyzing them we
have noticed that there are several types of new entries. The most frequent are direct translations of the
source query term, when there is no entry for it in the dictionaries. For the rest of the cases the users try to
improve the provided translations list by adding synonyms (“antena” for “dish”), plural expressions
(“anemone” for “anémona”), named entities (“London” for “Londres”), multiword expressions (“puente
torre” for “tower bridge”) or related concepts (“Africa” for “Uganda”, “alligator” for “caiman”).
Overall the contribution to the existing dictionaries was rather modest, and this can be explained by the fact
images were annotated with words that existed in the dictionary and in very few cases there was a dictionary
coverage problem, or by the fact almost 50% of the users confide completely in the automatic translation.
Also, due to an average number of just 18 entries per user, it is hard to assess an overall trend for each of the
five groups of users in terms of the quality of the personal dictionary. In the context of this game, bad
translations will quickly stand out, since judging relevance of the results set after a translation has been
inputted is a quick visual task. Intuitively, the degree of confidence with a language would positively affect
the quality of the personal dictionary.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Overall view of scores, precision, average time, and language skills</title>
        <p>We have last to answer the third research question regarding the dependency between the user’s performance
in the game results and the interactions with the personal dictionary and/or the assisted translation
mechanism. As seen in Figure 2, the language coefficient vs. distribution of translation related-actions shows
a very weak correlation (correlation = -0.07266). For the second graph the correlation coefficient points to a
medium strength between score and number of actions assisting translations (correlation = 0.3034) , while
the third graph shows a very weak link between retrieval precision and number of actions assisting
translations (correlation = 0.156).
These results support findings from previous research (Oard 2008) that acknowledged the user involvement
in the translation as a positive interaction; as with all information retrieval, the quickest way to the relevant
results is formulating a good query. This is in accordance with users’ feedback from Found Image
Questionnaire and Give Up Questionnaire that pointed at the difficulty in choosing good query terms for
characterizing the searched image.</p>
        <p>We conclude the presentation of these results, by looking at some of the answers from the overall
questionnaire that was filled in by user after searching 15 images, and was completed by 63 users.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Overall Questionnaire</title>
        <p>Out of the different questions comprised in this questionnaire, there are two main sets of questions that
regard how translation was performed. The first question refers to the most useful interface facilities. The
results show that automatic and assisted translations were perceived as equally important features. The
second question investigates how the translations decisions are made. The dominant answers refer to using
known languages or other language resources outside the game.</p>
        <p>Which interface facilities were most useful?
A. The automatic translation of query terms.
B. The possibility of improving the
translations chosen by the system.</p>
        <p>C. The additional query terms suggested by
the system ("You might also want to try
with... ").</p>
        <p>D. The assistant to select new query terms
from the set of results.</p>
        <p>How did you select the best translation for
the query terms?
A. Using my knowledge of target languages
whenever possible
B. Using additional dictionaries and other
online sources.</p>
        <p>C. I did not pay attention to the translations,
I just trusted the system
7
13
9
35
12
5
31
38
21
28
22
18
20
17
14
19
18
5
14
21
4
9
6
0
18
16</p>
        <sec id="sec-3-5-1">
          <title>Frequently</title>
        </sec>
        <sec id="sec-3-5-2">
          <title>Sometimes</title>
        </sec>
        <sec id="sec-3-5-3">
          <title>Rarely</title>
        </sec>
        <sec id="sec-3-5-4">
          <title>Never</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>Results obtained by previous iCLIR tasks at CLEF struggled to prove statistically the user’s impact on an
iCLIR system’s precision and recall. The difficulty of making such assessments lies in the complexity of the
interaction between a user and a CLIR system, and in defining suitable measures for interactive CLIR
systems in general. During the log analysis, it became apparent that finding suitable measures for computing
the different degrees of confidence in using a language is paramount for uncovering correlations. A finer
grain analysis of user actions grouped by image search might have shed more light on the relationship
between assisted translation, personal dictionary, and user’s search performance.</p>
      <p>Though we could not detect a clear link between usage of personal dictionaries and the efficiency of the
gamers search, this year’s experiment opens the path for furthering the research in using personalized
language resources in the translation process.</p>
      <p>Projects such as the Wiktionary, OmegaWiki, or Global WordNet are examples of global language resources
that could prospectively be used for deploying large-scale CLIR systems. These resources are by design
global and generic, and do not reflect the associated conceptualizations of specific groups of users. To
compensate this aspect, a user’s personal dictionary or maybe a community’s custom dictionary, can keep
translations in tune with a word’s most frequent used sense or changes of meaning.</p>
    </sec>
    <sec id="sec-5">
      <title>5. References</title>
      <p>Gonzalo, J., Clough, P., Karlgren, J. (2008) Overview of iCLEF 2008: search log analysis for Multilingual
Image Retrieval. In Borri, F., Nardi, A. and Peters, C., CLEF 2008 workshop notes.</p>
      <p>He, D. and Wang, J. (2007). Information Retrieval: Searching in the 21st Century, chapter Cross-Language
Information Retrieval. John Wiley &amp; Sons.</p>
      <p>Oard, D. W., He, D., and Wang, J. (2008). User-assisted query translation for interactive cross-language
information retrieval. Inf. Process. Manage., 44(1):181–211.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>