=Paper=
{{Paper
|id=None
|storemode=property
|title=CUES: Cognitive Usability Evaluation System
|pdfUrl=https://ceur-ws.org/Vol-909/poster5.pdf
|volume=Vol-909
|dblpUrl=https://dblp.org/rec/conf/eurohcir/PikeWDM12
}}
==CUES: Cognitive Usability Evaluation System==
CUES: Cognitive Usability Evaluation System
Matthew Pike Max L. Wilson Anna Divoli & Alyona Medelyan
Department of Computer Science Mixed Reality Lab Pingar Research
Swansea University, UK University of Nottingham, UK anna.divoli@pingar.com
matpike@gmail.com m.l.wilson@nottingham.ac.uk alyona.medelyan@pingar.com
ABSTRACT devices for measuring brain activity as users perform specific
A Cognitive Usability Evaluation System, CUES, was constructed tasks. Kitamura et al used fMRI outputs to show that after
to allow the simple integration of cognitive data from a repeating the task of learning how to use chopsticks the neural
commercialized EEG brain scanner, with other common usability activity patterns indeed indicate learning [2]. Cernea et al [1] used
measures, such as interaction logs, screen capture, and think the same EEG-device as the one used in CUES, the EPOC1. Their
aloud. CUES was iteratively evaluated with a small number of goal was to evaluate EPOC’s accuracy as it predicts users’ facial
participants to understand whether and how the visualisation of expressions (smiling, blinking) and their emotional state
EEG data alongside other measures, provided value for usability (calmness, excitement, engagement, frustration). Cernea et al
evaluation. Results indicate that although there are a lot of found that EPOC’s predictions are accurate in 70% to 100%
objective measurements available from the brain scanner, the cases, with the exception of excitement but concluded that it may
largest value came from qualitatively identifying EEG patterns, be a hard to define excitement as an intrinsically mental activity.
and correlating them with think aloud data. Recommendations for Vi and Subramanian [5] were able to accurately detect confusion
using CUES and for future developments are both provided. created by user interface design using the EPOC.
Despite providing difficult to use, and often noisy data, these
Categories and Subject Descriptors overall positive experiences of using brain scanners as research
H5.2. Information interfaces and presentation (User Interfaces): tools, as well as the lack of functioning systems for running user
evaluation/methodology, screen design. studies, have motivated us to build CUES. The following sections
describe CUES and provide a formative study of the value
Keywords provided by the EEG data.
Information Seeking, Cognitive Load Theory, EEG, Usability
1. INTRODUCTION
Evaluation of user interfaces is typically restricted by what can be
observed in specifically designed experimental environments or
through fieldwork. Aside from objective measures like time to
complete a task, researchers use questionnaires, interviews, think-
aloud protocols, and subjective observations to determine how
satisfied or frustrated the users are with a particular interface. In
some way, mouse movements, eye-tracking patterns, or
differences in heart rate can indicate emotional state of the
subject, but arguably, looking at the brain activity directly would
be more effective and accurate. Different neuro-imaging devices
and electromagnetic brain scanners have been recently introduced Figure 1: A screenshot of the CUES Visualiser.
as tools that can assist interface evaluation [1, 5] and they were
found to be accurate [2, 3]. In 2011, Wilson argued that brain 2. CUES
scanning devices might be useful for evaluating search user CUES is a collection of applications that allows researchers to
interfaces and their impact on a user’s cognitive load [6]. manage, automate and visualise user studies, as described below.
In this paper, we introduce the Cognitive Usability Evaluation 2.1 Study Setup and Recording
System, or CUES, as a universal tool to integrate cognitive EEG
CUES provides an intuitive interface for managing study related
data with other standard usability measurements. CUES can be
data, including: participant details, study tasks, study conditions
used to run studies with multiple participants and capture various
and the data sources to capture during a study. CUES’ study setup
data that may assist researchers in performing the evaluation.
component provides a range of settings for managing participants
CUES is designed to capture brain activity, as returned by an off-
in different study conditions. Having configured the study, it can
the shelf EEG-device Emotiv EPOC1. In addition, CUES
then be “run” within CUES. CUES is designed to capture
visualizes the captured outputs as shown in Figure 1, such as
interactions between participants and web pages. To facilitate this,
mouse movements (callout #2), audio (#3), and EEG data (#5).
CUES provides a simple customised web browser, which to a
To our knowledge, CUES is the first usability evaluation system participant appears indifferent from their everyday browser. In the
that features a brain scanning device as an integral part. However, background, however, the browser is capturing: audio, brain data,
others have reported experimental results of using alternative screenshots, mouse trails, and JavaScript based web events.
Audio is captured from the machines input device (e.g.
Presented at EuroHCIR2012. Copyright © 2012 for the individual papers by the
papers' authors. Copying permitted only for private and academic purposes. This Microphone). Brain data is acquired from the Emotiv EPOC
volume is published and copyrighted by its editors.
1
http://www.emotiv.com/
device. As well as providing the raw Electroencephalography with alternative visualizations. The last two allow users to browse
(EEG) signal, CUES collects the EPOC’s pre-classified emotions carefully designed taxonomies aiming at more expert audience.
(Engagement, Excitement, Frustration, etc.) and facial features MeSH in particular, is highly specialized and is used mostly for
(Smile Extent, Frown Extent, etc.). Screenshots of the webpage, automatic indexing tasks. This variation ensured different
as seen by the user, are captured at specified intervals, with a full- reactions from the participants. We chose this particular scenario,
page capture occurring upon each page load. Finally, JavaScript as it aligned with our other interests. Our findings about these
web events are captured via a custom JavaScript library that is taxonomy interfaces will be presented in a separate future paper.
injected into each page by the browser. These events allow CUES Six digital economy graduate students with different backgrounds,
to capture user interactions with the web page such as button including graphical design, geography, and economics, were
clicks, highlighted text, data entry etc. All captured data is stored recruited to take part in the study. Gender was balanced, and age
in a suitable, open format (Audio: wav, Screenshots: JPEG, Other: ranged between 22 and 45. Participation involved: 1) consent
XML) allowing the data to be analysed using other software as form (approved by the institution’s ethics committee) and setup of
well as the Visualiser (described below). the EPOC Emotive EEG scanner, 2) Phase 1: non-interactive
2.2 Visualising the Study brain response to systems’ designs, 3) Phase 2: content-agnostic
The Visualiser, shown in Figure 1, provides a way of correlating exploration of the systems, 4) Phase 3: applied exploration of the
various types of data in a time series. CUES also offers options to systems, and 5) a final debriefing interview. The applied Phase-3
customize of the visualisations, such as choosing which brain data involved participants searching for content relating to their current
and/or web events to include on the timelines. Further, and research, whereas Phase 2 always began with the initial query:
perhaps most importantly, the evaluator can stack multiple records ‘Schools’. Participation took 1 hour, where participants were
on top of each other for comparison, allowing them to compare, allowed to take breaks from wearing the Emotiv if needed.
for example, one participant’s performance on multiple tasks, or Participants were given an Amazon voucher for their time.
several participants’ performance on a certain task. 3.2 Quantitative Analysis
A hierarchical tree (#1) is provided for browsing the available In analysing the system, we first found that certain outputs from
recordings, which is ordered by study tasks, conditions and the EPOC had more value than others. Frustration, Short Term
participants. Once selected, each recording is opened within its Excitement (STE), and Engagement were the three emotions that
own tab (#2). Every visualisation within a single recording is showed most variance during interaction. While Meditation
linked to the audio waveform display (#3). The waveform showed almost no variation at any point in the study, Long Term
visualises the audio captured during the study, and optionally has Excitement (LTE) showed some usable variation for recordings of
the ability to be played with sound or muted (useful when 10 minutes or longer. These were infrequent in the study, and so
comparing many recordings at once). our analysis focused on Frustration, STE and Engagement.
Brain data are plotted on a 2D graph (#4), and emotions can be Although apparently a form of objective measurement, analysing
selectively added to the graph via the emotion selector (#5). EEG data does not lend itself comfortably to summarisation or
Additionally, web events (such as page loads, mouse overs and statistical comparison. As can be seen in Figure 1, the data varies
mouse clicks) can be selectively added to the graph (#4) through dramatically throughout a task phase. One may hypothesise that
the event selector (#6). Each event is added at the bottom of the average emotive values would help find the “most exciting” or the
chart at the point in time that the event occurred. Hovering over “most frustrating” system. However, as can be seen in Figure 3,
the event’s box on the chart gives additional event details (#7). participation averages tend to approximate with each other as they
Finally, there are additional visualisations that utilise the captured go through a number of peaks and troughs. Further, from the very
screenshots and mouse data. A researcher may select their desired first interaction with the system, participation diverges. This
visualisation from the tab component (#8). In Figure 1, we see divergence in behaviour means that the data at t=20s for one user
that the participant’s view of the web page at time X is overlaid is based on a completely different interaction for another
with their recent mouse trail (#9). Other visualisations include a participant. Consequently, to make a standard comparison, we
heat map of the cursor position, trail location on the entire page, must take a common event and examine the corresponding data.
the entire page by itself, and the visible region view. In our study, this was most obviously represented by Phase 1, in
which all participants were shown every UI one at a time, creating
3. FORMATIVE EVALUATION OF CUES data that could be compared both within and between participants.
To study the utility of the brain data we adapted the RITE method To further investigate the types of analyses that the CUES
[4] to iteratively make changes to the methodology as we learned Visualiser could support in future developments, we performed
about CUES’ capabilities. This process involved reflecting on the some manual analyses of the example data, shown in Figures 2-4.
utility and value of the data captured after each participant, and
trialling alternative configurations, such as: capturing facial Statistical Analyses. Figure 2 summarizes the average responses
expressions with the camera, turning off features, separating or for frustration and STE for three of the participants (p3, p4 and
joining the recording of subtasks to find the right level for p6). The comparison shows that different systems create varying
analysis, and so on. This process allowed us to examine and initial and delayed emotions. MeSH and ESD, for example, create
contrast recommendations for using and improving CUES. initial peaks of frustration, but drop lower after 20s, while Yippy
creates a form of frustration that peaks later. It is possible to take
3.1 Procedure and Participants some statistics, with the peak of STE for MeSH being almost
To create a scenario within which to trial CUES, tasks were significantly highest at t=11s (F(2)=6.47, p=0.056).
designed to evaluate the design of 4 very different taxonomy Summarising Data. Figure 3 shows graphs from Phase 2 that
interfaces: Yippy, CarrotSearch, MeSH, and ESD. Taxonomies compare results from different participants for the same system.
like these are a common form of Search User Interface feature. We should note that there is some data capture issues in places.
The first two of these systems present automatically generated The engagement data for participants p4 and p5, for example, are
categorisations of web search results, yet Carrot provides users
almost identical and appear to represent missing flat-lined data. occasions were silences during peaks of STE and -engagement.
Notably, however, general averages across the 3-5 minute tasks Approach 2: Event Detection. This approach involved a more
were quite even, indicating that averaging the data will not be predictive style. After determining common patterns, described
especially valuable for analysis. Similarly, Figure 4 shows the below, we were focusing on these patterns as we analysed each
participants’ average emotions throughout Phases 2 and 3 while participant’s brain data. As content was playing back in real time,
evaluating the 4 systems. Although we were hoping we’d see we could ‘see ahead’ which parts of the system the user would
relationships between other forms of usability data, such as find confusing or when the user was about to figure something
subjective preferences captured in interviews, we were unable to out. Beyond giving us these specific insights, the patterns also
find any obvious relationships.. In the future, we will investigate allowed us to examine the times of high confusion; or to examine
other quantitative approaches that might be relatable to other the times of effective progress.
forms of usability measures, such counting the number of EEG
graphs’ peaks and troughs above and below given thresholds, as Common EEG patterns:
well as their scale, and allowing summarisations during certain - High frustration and low excitement
events or time-periods, rather than for entire tasks. o often indicating confusion
- A peak of frustration followed by a peak of excitement
3.3 Qualitative Analysis o often indicating comprehension
Of all the data comparisons above, it is very difficult to draw any
- Low excitement and frustration, with high engagement
conclusions about ‘average data’ having much value, because
o often indicating effective progress
average data across an entire task means very little. Even
- Low frustration and high excitement
averaging across participants at a given time is difficult, when o often indicating (good) discovery
interaction diverges. In our experience, however, the most
valuable insights gathered from the brain data were in watching 4. DISCUSSION
for patterns in the signal curves and investigating the correlated Overall, we experimented with both quantitative and qualitative
subjective data, such as the think-aloud data and the mouse trails, data captured by CUES, as well as approaches to analysing them
for additional insights. This combination was much more valuable using CUES. Overwhelmingly, we found that the best value
than the other combinations we tried, such as recording the facial provided by the brain scanner was in qualitative analysis, where
expressions with a camera. This valuable qualitative process the data allowed us to a) augment the verbal protocol, b) see
involved two approaches, described below. ahead of the verbal protocol, and c) explore and examine specific
Approach 1: Validating Think Aloud. This first approach parts of the verbal protocol. In each of these cases, we found it
involved playing back the brain, think-aloud, and mouse trail data extremely helpful to also see the user’s view, mouse trail, and
in real time, which allowed us to qualify utterances in the think- logged interactions.
aloud approach. For example, using think-aloud alone, there were Despite appearing as a quantitative source, the qualitative value
many occasions where participants would utter a comment gained from augmenting other more common usability metrics.
indicating that they did not understand something. Using levels of The specific added value came in two areas. First, the brain data
frustration and engagement, we could clearly see which of these provided additional insight and context into the content of the
occasions was creating a significant barrier to use, and which verbal protocol, which is otherwise often ambiguous and open to
were unimportant. Further, we could identify possible reasons for the interpretation of the investigator. Second, the brain data added
silence during the verbal-protocol, with some peaking in a visual dimension to the verbal protocol, which is what allowed
frustration when, for example, pages were not loading. Other us to both see ahead and specifically explore the data.
Figure 2: Initial response time-curves in the first 45s of seeing a UI
Figure 3: Consistency between users in Phase 2
Figure 4: Internal Consistency between users in Phases 2 and 3. Favourite and least favourite system and systems with
favourite features are also shown (based on the interview questions at the end of the study).
4.1 Limitations rate from the mouse data, as the current mouse trail was limited to
Despite finding a lot of value in analyzing the EPOC data the frame rate chosen for screen capture. In order to avoid data
qualitatively, there are still some well-known limitations to using loss, it was suggested that a warning appear during tasks when
EEG data. First and foremost, EEG data is easily confounded by any of the EPOC sensors lost its signal. In regards to the
body movement. The motor control of fingers, hands, and arms, Visualiser, greater control was desired to easily see all the
for example, can create noisy data and arbitrary peaks. In CUES, elements when stacking several records on top of each other. In
however, the cross-validation in the think-aloud and brain data this paper, we also explored alternative visualisation and analyses,
allows for some of this noise to be ignored. So far, however, we which we hope to integrate in the future. Feedback also indicated
have not specifically measured body movement. that global controls, rather than per record, were desirable, to
avoid constant reconfiguration from the default. Finally, the
Further, we frequently saw, especially during the interviews, motion and control over the viewport and playback is currently
frustration correlate with speaking. Although it seems like a tied; future versions will allow independent control.
verbal protocol may, therefore, completely mask the data, we
found it was times when participants had to think and explain 5. CONCLUSIONS
what was happening. In this case, the verbal protocol often made This paper described CUES, a prototype system designed to
frustration and lack of understanding more visible in the system. utilise cheap off the shelf EEG brain scanners to help run usability
There are many other limitations to the study, which was only a studies. A formative evaluation provided many insights into the
formative investigation into the utility of CUES , using a scenario value of different features. Despite being primarily objective in
focused on evaluating a single form Search User Interface feature. nature, we found that the EEG data was most effective when
We plan to run a much larger hypothesis-driven evaluation of analysed qualitatively in parallel with think-aloud data. The EEG
CUES in the future. data a) helped to validate or qualify ambiguous think aloud
comments, and b) added a visual dimension to the verbal protocol
4.2 Recommendations for using CUES allowing us to look ahead at their experience and explore the data
Good data. Despite concerns, we were able to get good data for certain events. Ultimately, we conclude that a lot of value can
regardless of hair length, etc. However, one must watch out for be gained from using CUES to investigate EEG brain
flat-lined data from one or 2 bad sensors, which leads to data loss. measurements in parallel with other usability measures such as
Waiting for data. We discovered that there is a 10s lead time as logs, screen captures, and think-aloud protocols.
certain pre-classified measures begin to show. Short tasks, such as 6. REFERENCES
visual exposure, need to be extended to include this lead time. [1] Cernea, D., Olech, P.-S., Ebert, A. and Kerren, A., EEG-
LTE required tasks must be 10+ minutes long to have value.
Based Measurement of Subjective Parameters in Evaluations.
Comfort. We learnt that participants could wear the device for In HCII'11 - Posters, 279-283. 2011
sustained periods of time. Some participants experienced mild [2] Kitamura, Y., Yamaguchi, Y., Hiroshi, I., Kishino, F. and
discomfort after wearing the device for more than 40 minutes. Kawato, M., Things happening in the brain while humans
Task Chunking. Correctly separating out tasks is important. If learn to use new tools. In CHI'03, 417-424. 2003
you want to compare a person’s response to System A versus [3] Liu, Y., Sourina, O. and Nguyen, M.K., Real-Time EEG-
System B, they must be in separate recordings to facilitate easy Based Human Emotion Recognition and Visualization. In
comparison and analysis. CW'10, 262-269. 2010
[4] Medlock, M., Wixon, D., Terrano, M., Romero, R. and
4.3 CUES Improvements Fulton, B., Using the RITE method to improve products; a
Conducting the study allowed us to identify several areas for definition and a case study. In Usability Professionals
improving CUES. One feature of the system captures an entire Association. 2002
website, rather than just the page view, but this created an unusual [5] Vi, C. and Subramanian, S., Detecting error-related negativity
page load event that, in turn, created artificial levels of frustration for interaction design. In CHI'12, 493-502. 2012
in our first participant. To be useful, this element needs an [6] Wilson, M.L., Evaluating the Cognitive Impact of Search
alternative implementation to have no visual effect on the user. User Interface Design Decisions. In EuroHCIR 2011, 27-30.
Further, we also wished to separate the viewport capture frame 2011