=Paper= {{Paper |id=None |storemode=property |title=CUES: Cognitive Usability Evaluation System |pdfUrl=https://ceur-ws.org/Vol-909/poster5.pdf |volume=Vol-909 |dblpUrl=https://dblp.org/rec/conf/eurohcir/PikeWDM12 }} ==CUES: Cognitive Usability Evaluation System== https://ceur-ws.org/Vol-909/poster5.pdf
                   CUES: Cognitive Usability Evaluation System
                Matthew Pike                                            Max L. Wilson                   Anna Divoli & Alyona Medelyan
   Department of Computer Science                                    Mixed Reality Lab                            Pingar Research
      Swansea University, UK                                   University of Nottingham, UK                    anna.divoli@pingar.com
        matpike@gmail.com                                      m.l.wilson@nottingham.ac.uk                  alyona.medelyan@pingar.com


ABSTRACT                                                                           devices for measuring brain activity as users perform specific
A Cognitive Usability Evaluation System, CUES, was constructed                     tasks. Kitamura et al used fMRI outputs to show that after
to allow the simple integration of cognitive data from a                           repeating the task of learning how to use chopsticks the neural
commercialized EEG brain scanner, with other common usability                      activity patterns indeed indicate learning [2]. Cernea et al [1] used
measures, such as interaction logs, screen capture, and think                      the same EEG-device as the one used in CUES, the EPOC1. Their
aloud. CUES was iteratively evaluated with a small number of                       goal was to evaluate EPOC’s accuracy as it predicts users’ facial
participants to understand whether and how the visualisation of                    expressions (smiling, blinking) and their emotional state
EEG data alongside other measures, provided value for usability                    (calmness, excitement, engagement, frustration). Cernea et al
evaluation. Results indicate that although there are a lot of                      found that EPOC’s predictions are accurate in 70% to 100%
objective measurements available from the brain scanner, the                       cases, with the exception of excitement but concluded that it may
largest value came from qualitatively identifying EEG patterns,                    be a hard to define excitement as an intrinsically mental activity.
and correlating them with think aloud data. Recommendations for                    Vi and Subramanian [5] were able to accurately detect confusion
using CUES and for future developments are both provided.                          created by user interface design using the EPOC.
                                                                                   Despite providing difficult to use, and often noisy data, these
Categories and Subject Descriptors                                                 overall positive experiences of using brain scanners as research
H5.2. Information interfaces and presentation (User Interfaces):                   tools, as well as the lack of functioning systems for running user
evaluation/methodology, screen design.                                             studies, have motivated us to build CUES. The following sections
                                                                                   describe CUES and provide a formative study of the value
Keywords                                                                           provided by the EEG data.
Information Seeking, Cognitive Load Theory, EEG, Usability

1. INTRODUCTION
Evaluation of user interfaces is typically restricted by what can be
observed in specifically designed experimental environments or
through fieldwork. Aside from objective measures like time to
complete a task, researchers use questionnaires, interviews, think-
aloud protocols, and subjective observations to determine how
satisfied or frustrated the users are with a particular interface. In
some way, mouse movements, eye-tracking patterns, or
differences in heart rate can indicate emotional state of the
subject, but arguably, looking at the brain activity directly would
be more effective and accurate. Different neuro-imaging devices
and electromagnetic brain scanners have been recently introduced                             Figure 1: A screenshot of the CUES Visualiser.
as tools that can assist interface evaluation [1, 5] and they were
found to be accurate [2, 3]. In 2011, Wilson argued that brain                     2. CUES
scanning devices might be useful for evaluating search user                        CUES is a collection of applications that allows researchers to
interfaces and their impact on a user’s cognitive load [6].                        manage, automate and visualise user studies, as described below.
In this paper, we introduce the Cognitive Usability Evaluation                     2.1 Study Setup and Recording
System, or CUES, as a universal tool to integrate cognitive EEG
                                                                                   CUES provides an intuitive interface for managing study related
data with other standard usability measurements. CUES can be
                                                                                   data, including: participant details, study tasks, study conditions
used to run studies with multiple participants and capture various
                                                                                   and the data sources to capture during a study. CUES’ study setup
data that may assist researchers in performing the evaluation.
                                                                                   component provides a range of settings for managing participants
CUES is designed to capture brain activity, as returned by an off-
                                                                                   in different study conditions. Having configured the study, it can
the shelf EEG-device Emotiv EPOC1. In addition, CUES
                                                                                   then be “run” within CUES. CUES is designed to capture
visualizes the captured outputs as shown in Figure 1, such as
                                                                                   interactions between participants and web pages. To facilitate this,
mouse movements (callout #2), audio (#3), and EEG data (#5).
                                                                                   CUES provides a simple customised web browser, which to a
To our knowledge, CUES is the first usability evaluation system                    participant appears indifferent from their everyday browser. In the
that features a brain scanning device as an integral part. However,                background, however, the browser is capturing: audio, brain data,
others have reported experimental results of using alternative                     screenshots, mouse trails, and JavaScript based web events.
                                                                                   Audio is captured from the machines input device (e.g.
 Presented at EuroHCIR2012. Copyright © 2012 for the individual papers by the
 papers' authors. Copying permitted only for private and academic purposes. This   Microphone). Brain data is acquired from the Emotiv EPOC
 volume is published and copyrighted by its editors.
                                                                                   1
                                                                                       http://www.emotiv.com/
device. As well as providing the raw Electroencephalography              with alternative visualizations. The last two allow users to browse
(EEG) signal, CUES collects the EPOC’s pre-classified emotions           carefully designed taxonomies aiming at more expert audience.
(Engagement, Excitement, Frustration, etc.) and facial features          MeSH in particular, is highly specialized and is used mostly for
(Smile Extent, Frown Extent, etc.). Screenshots of the webpage,          automatic indexing tasks. This variation ensured different
as seen by the user, are captured at specified intervals, with a full-   reactions from the participants. We chose this particular scenario,
page capture occurring upon each page load. Finally, JavaScript          as it aligned with our other interests. Our findings about these
web events are captured via a custom JavaScript library that is          taxonomy interfaces will be presented in a separate future paper.
injected into each page by the browser. These events allow CUES          Six digital economy graduate students with different backgrounds,
to capture user interactions with the web page such as button            including graphical design, geography, and economics, were
clicks, highlighted text, data entry etc. All captured data is stored    recruited to take part in the study. Gender was balanced, and age
in a suitable, open format (Audio: wav, Screenshots: JPEG, Other:        ranged between 22 and 45. Participation involved: 1) consent
XML) allowing the data to be analysed using other software as            form (approved by the institution’s ethics committee) and setup of
well as the Visualiser (described below).                                the EPOC Emotive EEG scanner, 2) Phase 1: non-interactive
2.2 Visualising the Study                                                brain response to systems’ designs, 3) Phase 2: content-agnostic
The Visualiser, shown in Figure 1, provides a way of correlating         exploration of the systems, 4) Phase 3: applied exploration of the
various types of data in a time series. CUES also offers options to      systems, and 5) a final debriefing interview. The applied Phase-3
customize of the visualisations, such as choosing which brain data       involved participants searching for content relating to their current
and/or web events to include on the timelines. Further, and              research, whereas Phase 2 always began with the initial query:
perhaps most importantly, the evaluator can stack multiple records       ‘Schools’. Participation took 1 hour, where participants were
on top of each other for comparison, allowing them to compare,           allowed to take breaks from wearing the Emotiv if needed.
for example, one participant’s performance on multiple tasks, or         Participants were given an Amazon voucher for their time.
several participants’ performance on a certain task.                     3.2 Quantitative Analysis
A hierarchical tree (#1) is provided for browsing the available          In analysing the system, we first found that certain outputs from
recordings, which is ordered by study tasks, conditions and              the EPOC had more value than others. Frustration, Short Term
participants. Once selected, each recording is opened within its         Excitement (STE), and Engagement were the three emotions that
own tab (#2). Every visualisation within a single recording is           showed most variance during interaction. While Meditation
linked to the audio waveform display (#3). The waveform                  showed almost no variation at any point in the study, Long Term
visualises the audio captured during the study, and optionally has       Excitement (LTE) showed some usable variation for recordings of
the ability to be played with sound or muted (useful when                10 minutes or longer. These were infrequent in the study, and so
comparing many recordings at once).                                      our analysis focused on Frustration, STE and Engagement.
Brain data are plotted on a 2D graph (#4), and emotions can be           Although apparently a form of objective measurement, analysing
selectively added to the graph via the emotion selector (#5).            EEG data does not lend itself comfortably to summarisation or
Additionally, web events (such as page loads, mouse overs and            statistical comparison. As can be seen in Figure 1, the data varies
mouse clicks) can be selectively added to the graph (#4) through         dramatically throughout a task phase. One may hypothesise that
the event selector (#6). Each event is added at the bottom of the        average emotive values would help find the “most exciting” or the
chart at the point in time that the event occurred. Hovering over        “most frustrating” system. However, as can be seen in Figure 3,
the event’s box on the chart gives additional event details (#7).        participation averages tend to approximate with each other as they
Finally, there are additional visualisations that utilise the captured   go through a number of peaks and troughs. Further, from the very
screenshots and mouse data. A researcher may select their desired        first interaction with the system, participation diverges. This
visualisation from the tab component (#8). In Figure 1, we see           divergence in behaviour means that the data at t=20s for one user
that the participant’s view of the web page at time X is overlaid        is based on a completely different interaction for another
with their recent mouse trail (#9). Other visualisations include a       participant. Consequently, to make a standard comparison, we
heat map of the cursor position, trail location on the entire page,      must take a common event and examine the corresponding data.
the entire page by itself, and the visible region view.                  In our study, this was most obviously represented by Phase 1, in
                                                                         which all participants were shown every UI one at a time, creating
3. FORMATIVE EVALUATION OF CUES                                          data that could be compared both within and between participants.
To study the utility of the brain data we adapted the RITE method        To further investigate the types of analyses that the CUES
[4] to iteratively make changes to the methodology as we learned         Visualiser could support in future developments, we performed
about CUES’ capabilities. This process involved reflecting on the        some manual analyses of the example data, shown in Figures 2-4.
utility and value of the data captured after each participant, and
trialling alternative configurations, such as: capturing facial          Statistical Analyses. Figure 2 summarizes the average responses
expressions with the camera, turning off features, separating or         for frustration and STE for three of the participants (p3, p4 and
joining the recording of subtasks to find the right level for            p6). The comparison shows that different systems create varying
analysis, and so on. This process allowed us to examine and              initial and delayed emotions. MeSH and ESD, for example, create
contrast recommendations for using and improving CUES.                   initial peaks of frustration, but drop lower after 20s, while Yippy
                                                                         creates a form of frustration that peaks later. It is possible to take
3.1 Procedure and Participants                                           some statistics, with the peak of STE for MeSH being almost
To create a scenario within which to trial CUES, tasks were              significantly highest at t=11s (F(2)=6.47, p=0.056).
designed to evaluate the design of 4 very different taxonomy             Summarising Data. Figure 3 shows graphs from Phase 2 that
interfaces: Yippy, CarrotSearch, MeSH, and ESD. Taxonomies               compare results from different participants for the same system.
like these are a common form of Search User Interface feature.           We should note that there is some data capture issues in places.
The first two of these systems present automatically generated           The engagement data for participants p4 and p5, for example, are
categorisations of web search results, yet Carrot provides users
almost identical and appear to represent missing flat-lined data.       occasions were silences during peaks of STE and -engagement.
Notably, however, general averages across the 3-5 minute tasks          Approach 2: Event Detection. This approach involved a more
were quite even, indicating that averaging the data will not be         predictive style. After determining common patterns, described
especially valuable for analysis. Similarly, Figure 4 shows the         below, we were focusing on these patterns as we analysed each
participants’ average emotions throughout Phases 2 and 3 while          participant’s brain data. As content was playing back in real time,
evaluating the 4 systems. Although we were hoping we’d see              we could ‘see ahead’ which parts of the system the user would
relationships between other forms of usability data, such as            find confusing or when the user was about to figure something
subjective preferences captured in interviews, we were unable to        out. Beyond giving us these specific insights, the patterns also
find any obvious relationships.. In the future, we will investigate     allowed us to examine the times of high confusion; or to examine
other quantitative approaches that might be relatable to other          the times of effective progress.
forms of usability measures, such counting the number of EEG
graphs’ peaks and troughs above and below given thresholds, as          Common EEG patterns:
well as their scale, and allowing summarisations during certain          -    High frustration and low excitement
events or time-periods, rather than for entire tasks.                           o often indicating confusion
                                                                         -    A peak of frustration followed by a peak of excitement
3.3 Qualitative Analysis                                                        o often indicating comprehension
Of all the data comparisons above, it is very difficult to draw any
                                                                         -    Low excitement and frustration, with high engagement
conclusions about ‘average data’ having much value, because
                                                                                o often indicating effective progress
average data across an entire task means very little. Even
                                                                         -    Low frustration and high excitement
averaging across participants at a given time is difficult, when                o often indicating (good) discovery
interaction diverges. In our experience, however, the most
valuable insights gathered from the brain data were in watching         4. DISCUSSION
for patterns in the signal curves and investigating the correlated      Overall, we experimented with both quantitative and qualitative
subjective data, such as the think-aloud data and the mouse trails,     data captured by CUES, as well as approaches to analysing them
for additional insights. This combination was much more valuable        using CUES. Overwhelmingly, we found that the best value
than the other combinations we tried, such as recording the facial      provided by the brain scanner was in qualitative analysis, where
expressions with a camera. This valuable qualitative process            the data allowed us to a) augment the verbal protocol, b) see
involved two approaches, described below.                               ahead of the verbal protocol, and c) explore and examine specific
Approach 1: Validating Think Aloud. This first approach                 parts of the verbal protocol. In each of these cases, we found it
involved playing back the brain, think-aloud, and mouse trail data      extremely helpful to also see the user’s view, mouse trail, and
in real time, which allowed us to qualify utterances in the think-      logged interactions.
aloud approach. For example, using think-aloud alone, there were        Despite appearing as a quantitative source, the qualitative value
many occasions where participants would utter a comment                 gained from augmenting other more common usability metrics.
indicating that they did not understand something. Using levels of      The specific added value came in two areas. First, the brain data
frustration and engagement, we could clearly see which of these         provided additional insight and context into the content of the
occasions was creating a significant barrier to use, and which          verbal protocol, which is otherwise often ambiguous and open to
were unimportant. Further, we could identify possible reasons for       the interpretation of the investigator. Second, the brain data added
silence during the verbal-protocol, with some peaking in                a visual dimension to the verbal protocol, which is what allowed
frustration when, for example, pages were not loading. Other            us to both see ahead and specifically explore the data.




                                 Figure 2: Initial response time-curves in the first 45s of seeing a UI




                                           Figure 3: Consistency between users in Phase 2
        Figure 4: Internal Consistency between users in Phases 2 and 3. Favourite and least favourite system and systems with
                      favourite features are also shown (based on the interview questions at the end of the study).

4.1 Limitations                                                           rate from the mouse data, as the current mouse trail was limited to
Despite finding a lot of value in analyzing the EPOC data                 the frame rate chosen for screen capture. In order to avoid data
qualitatively, there are still some well-known limitations to using       loss, it was suggested that a warning appear during tasks when
EEG data. First and foremost, EEG data is easily confounded by            any of the EPOC sensors lost its signal. In regards to the
body movement. The motor control of fingers, hands, and arms,             Visualiser, greater control was desired to easily see all the
for example, can create noisy data and arbitrary peaks. In CUES,          elements when stacking several records on top of each other. In
however, the cross-validation in the think-aloud and brain data           this paper, we also explored alternative visualisation and analyses,
allows for some of this noise to be ignored. So far, however, we          which we hope to integrate in the future. Feedback also indicated
have not specifically measured body movement.                             that global controls, rather than per record, were desirable, to
                                                                          avoid constant reconfiguration from the default. Finally, the
Further, we frequently saw, especially during the interviews,             motion and control over the viewport and playback is currently
frustration correlate with speaking. Although it seems like a             tied; future versions will allow independent control.
verbal protocol may, therefore, completely mask the data, we
found it was times when participants had to think and explain             5. CONCLUSIONS
what was happening. In this case, the verbal protocol often made          This paper described CUES, a prototype system designed to
frustration and lack of understanding more visible in the system.         utilise cheap off the shelf EEG brain scanners to help run usability
There are many other limitations to the study, which was only a           studies. A formative evaluation provided many insights into the
formative investigation into the utility of CUES , using a scenario       value of different features. Despite being primarily objective in
focused on evaluating a single form Search User Interface feature.        nature, we found that the EEG data was most effective when
We plan to run a much larger hypothesis-driven evaluation of              analysed qualitatively in parallel with think-aloud data. The EEG
CUES in the future.                                                       data a) helped to validate or qualify ambiguous think aloud
                                                                          comments, and b) added a visual dimension to the verbal protocol
4.2 Recommendations for using CUES                                        allowing us to look ahead at their experience and explore the data
Good data. Despite concerns, we were able to get good data                for certain events. Ultimately, we conclude that a lot of value can
regardless of hair length, etc. However, one must watch out for           be gained from using CUES to investigate EEG brain
flat-lined data from one or 2 bad sensors, which leads to data loss.      measurements in parallel with other usability measures such as
Waiting for data. We discovered that there is a 10s lead time as          logs, screen captures, and think-aloud protocols.
certain pre-classified measures begin to show. Short tasks, such as       6. REFERENCES
visual exposure, need to be extended to include this lead time.           [1] Cernea, D., Olech, P.-S., Ebert, A. and Kerren, A., EEG-
LTE required tasks must be 10+ minutes long to have value.
                                                                              Based Measurement of Subjective Parameters in Evaluations.
Comfort. We learnt that participants could wear the device for                In HCII'11 - Posters, 279-283. 2011
sustained periods of time. Some participants experienced mild             [2] Kitamura, Y., Yamaguchi, Y., Hiroshi, I., Kishino, F. and
discomfort after wearing the device for more than 40 minutes.                 Kawato, M., Things happening in the brain while humans
Task Chunking. Correctly separating out tasks is important. If                learn to use new tools. In CHI'03, 417-424. 2003
you want to compare a person’s response to System A versus                [3] Liu, Y., Sourina, O. and Nguyen, M.K., Real-Time EEG-
System B, they must be in separate recordings to facilitate easy              Based Human Emotion Recognition and Visualization. In
comparison and analysis.                                                      CW'10, 262-269. 2010
                                                                          [4] Medlock, M., Wixon, D., Terrano, M., Romero, R. and
4.3 CUES Improvements                                                         Fulton, B., Using the RITE method to improve products; a
Conducting the study allowed us to identify several areas for                 definition and a case study. In Usability Professionals
improving CUES. One feature of the system captures an entire                  Association. 2002
website, rather than just the page view, but this created an unusual      [5] Vi, C. and Subramanian, S., Detecting error-related negativity
page load event that, in turn, created artificial levels of frustration       for interaction design. In CHI'12, 493-502. 2012
in our first participant. To be useful, this element needs an             [6] Wilson, M.L., Evaluating the Cognitive Impact of Search
alternative implementation to have no visual effect on the user.              User Interface Design Decisions. In EuroHCIR 2011, 27-30.
Further, we also wished to separate the viewport capture frame                2011