<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparing the Usability of two Multi-Agents Systems DSLs: SEA_ML++ and DSML4MAS Study Design</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>João Silva</institution>
          ,
          <addr-line>Ankica Barišic ́ , Vasco Amaral , Miguel Goulão</addr-line>
          ,
          <country>Baris Tekin Tezel</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Context: The “Physics of Notations” (PoN) supports a systematic improvement of the cognitive effectiveness of visual modelling languages. Problem: PoN focuses on the concrete syntax of a language, building on a predefined abstract syntax. We should also consider the abstract syntax of a language when developing efforts to improve it by choosing the most adequate language constructs (concepts and their relationships). We instantiate this challenge by comparing two Multi-Agent Systems Domain Specific Languages: SEA_ML++ and DSML4MAS, and assessing the extent to which their respective constructs affect the developer experience. Method: We will perform a quasiexperiment for comparing how practitioners use both languages to solve similar modelling challenges. The experiment will have a cross-over within-subjects design and will focus on the extent to which the different language constructs impact on developer experience. These tasks will be monitored, so that we can assess their success and effort involved, including eye-tracking information. Results: This paper reports on the planned study design for this empirical comparison of two DSLs for MAS.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>In the last two decades, technologies like modelling
workbenches made it easier to design, prototype and deploy
diagrammatic languages used more often for capturing
abstractions in modelling. Extensive experience with the development
of domain-specific languages (DSLs) lead to a new discipline,
Software Languages Engineering (SLE), with the goal of
making systematic the process of developing a software language.</p>
      <p>
        SLE follows an iterative life-cycle [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] that starts with
domain analysis, followed by language design, implementation
and evaluation. Unfortunately, the first and the last steps are
still not at a mature phase. Besides taking into account the
evaluation of expressiveness of a given language, the language
design (coverage of the language goals) needs to make use of
empirical studies to assess the language usability.
      </p>
      <p>
        The “Physics of Notations” (PoN) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] created a valuable
framework to evaluate the language’s concrete syntax, and is
extensively used to support a systematic improvement of the
cognitive effectiveness of visual modelling languages with a
fixed abstract syntax in some language metamodel or grammar.
      </p>
      <p>Improving the concrete syntax is very important, but we
should also consider the abstract syntax of a language. We
should be able to choose and validate the adequate language
constructs (concepts and their relationships) and the models
(or language sentences) we can express with those. However,
there is a lack of guidelines reported in the literature for this
new level of assessment that could give a languages engineer
a “recipe” for doing this sort of evaluations. These would be
valuable when developing and improving a given language.</p>
      <p>
        We are currently improving a Multi-Agent Systems (MAS)
DSL: SEA_ML++ [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In this context, we are planning an
empirical comparison with another MAS DSL: DSML4MAS
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], to assess the extent to which these DSLs respective
constructs, and combinations, affect language usability.
      </p>
      <p>
        We will perform a quasi-experiment for comparing how
practitioners use both languages to solve similar modelling
challenges. We will have a cross-over within-subjects design
with a focus on how the different language constructs impact
on their usability by modellers. We will monitor these tasks to
assess their success and effort involved, including eye-tracking
information and a usability questionnaire (SUS [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]).
      </p>
      <p>This paper is organised as follows: Section II describes
DSMLs for MAS languages as the object of our case study.
Section III presents the planned quasi-experiment, followed by
a discussion in Section IV. Section V summarises this paper.</p>
    </sec>
    <sec id="sec-2">
      <title>II. BACKGROUND</title>
      <sec id="sec-2-1">
        <title>A. Multi-Agent Systems DSMLs</title>
        <p>
          Software agents are autonomous entities which contain
intelligence that serves for solving their selfish or common
problems and to achieve certain goals. The study of
MultiAgent Systems (MASs) focuses on those systems in which
many intelligent agents interact with each other. In
agentoriented software engineering (AOSE), the application of
model-driven development (MDD) and the use of
domainspecific modelling languages (DSMLs) for MAS development
are quite popular since the implementation of MAS is naturally
complex, error-prone and costly due to the autonomous and
proactive properties of the agents [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          In the last decade, several MAS modelling languages and
DSMLs (e.g. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]) were proposed to
support development of MASs. For example, DSML4MAS [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
introduces a general MAS metamodel with various viewpoints
that enable the development of MAS for many application
domains. A DSL is introduced in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] to provide a
language for the development of mobile agents. In addition, [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]
introduces a modelling language enabling the model-driven
development within the scope of Prometheus methodology for
agent development. In [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] a graphical DSML (called
SEA_ML) and textual DSL (called SEA_L) are proposed
for MAS working in semantic web environments including 8
viewpoints. MAS-ML 2.0 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] is a modelling language which
supports the MAS modelling with different agent architectures
such as: Simple Reflex Agents, Model-Based Reflex Agents,
Goal-Based Agents and Utility-Based Agents. DSML4BDI
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] is another modelling language specific for Jason agent
programming language. In [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], the authors propose a
toolsupported development method that applies MDD techniques
to design and implement agents based on the
belief-desireintention architecture with a sophisticated plan selection
process.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>B. The “Physics of Notations” (PoN)</title>
        <p>
          Moody proposed the “Physics of Notations” [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to support
the construction of more effective software languages. A major
concern is on how to evaluate the cognitive effectiveness
of visual languages (see, for example [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]). The framework
concentrates on the physical properties (concrete syntax) of
the symbols and not on their structure (abstract syntax) or
semantics (ignoring semantics of both the ontological and
language target semantic Domain). In figure 1, we present the
dimensions at the instance level (in grey) that are explored
by the current work. Here we study the composition of visual
elements and its structure to form sentence instances. This
figure is adapted from [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], where the authors refer to their
focus on the top left corner (Visual Notations).
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>C. Related studies</title>
        <p>Most of the available DS(M)Ls proposed for MASs have
been evaluated by just providing a case study demonstrating
how the related language can be used for design and
implementation of MAS. A quantitative analysis and/or qualitative
evaluation considering e.g. the development time performance,
generation performance, and/or the usability of the language
are not considered in these studies.</p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], we proposed an evaluation framework which
provides the systematic assessment of both the language
constructs and the use of agent DSMLs according to various
dimensions and criteria. The study also provides an assessment
of SEA_ML [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. However, it does not take into account the
effect of language constructs in the developer’s modelling
process while using the languages. This evaluation framework
is adopted in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] for the assessment of the
proposed MAS DSMLs. Another MAS DSML evaluation
feature exists in [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] for a textual DSL, called JADEL, providing
four abstractions, namely agents, behaviours, communication
ontologies, and interaction protocols to the well-known JADE
agent development framework. However, the study evaluates
solely JADEL’s code generation performance.
        </p>
        <p>
          In recent years, we have seen several studies to identify
language improvement opportunities, identifying problems with
their concrete syntax and how they impact developer
experience. These studies have covered a diversity of languages,
including UML [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], BPMN [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], KAOS [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ],
[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], i* [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ], OutSystems BPT [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ], and
SEA_ML++ [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Some of these languages were also analysed
from the perspective of the impact of diagram layout in the
understandability of models, namely UML [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ], [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ], [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] and
i* [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ]. Other studies have compared alternative DSLs for a
similar domain (e.g. Lego Mindstorms vs. Gyro [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ]).
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>III. EXPERIMENT PLANNING</title>
      <p>This section describes the experimental planning for this
evaluation. Further details, including documentation and
evaluation materials, can be found in this paper’s companion site1.</p>
      <sec id="sec-3-1">
        <title>A. Goals</title>
        <p>
          Broadly, we are interested in assessing the usability of two
MAS DSMLs, SEA_ML++ [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and DSML4MAS [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] in
the context of solving modelling challenges. We use the
GoalQuestion-Metric [
          <xref ref-type="bibr" rid="ref36">36</xref>
          ] template to describe our research goals:
        </p>
        <p>Our first goal (G1) is to analyse the effect of using
SEA_ML++ or DSML4MAS, for the purpose of evaluation,
with respect to the correctness with which a developer models
a MAS system, from the viewpoint of researchers, in the
context of an experiment conducted with graduate students
from Universidade Nova de Lisboa and Ege University. Our
second goal (G2) is to analyse the effect of using SEA_ML++
or DSML4MAS, for the purpose of evaluation, with respect to
the speed with which a developer models a MAS system, from
1https://sites.google.com/fct.unl.pt/hufamo2018masstudydesign/home
the viewpoint of researchers, in the context of an experiment
conducted with graduate students from Universidade Nova
de Lisboa and Ege University. Our third goal is to analyse
the effect of using SEA_ML++ or DSML4MAS, for the
purpose of evaluation, with respect to the rework involved in
modelling a MAS system, from the viewpoint of researchers, in
the context of an experiment conducted with graduate students
from Universidade Nova de Lisboa and Ege University.</p>
      </sec>
      <sec id="sec-3-2">
        <title>B. Experimental units</title>
        <p>The participants in this evaluation will be professional
software developers from Lisbon, and graduate students trained
in several universities, namely Universidade Nova de Lisboa,
Instituto Superior Técnico and Instituto Universitário de
Lisboa. We will have a close replica of these evaluations with
subjects from the Ege University, in Turkey. We will use
convenience sampling to recruit participants, in all these sites.
Each participant will be randomly assigned to one of four
groups, keeping a balanced sample on each of the four groups.</p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Tasks</title>
        <p>Each subject will be asked to perform two modelling tasks:
one using SEA_ML++, the other DSML4MAS. The two tasks
will have similar complexity and will consist in modelling a
MAS system from a natural language description of that
system. They will use an Eclipse-based editor, which is essentially
similar. The editor only varies in the language constructs and
composition rules offered to participants, depending on which
language is being used. The participant will make his best
to correctly model a system with each of these languages.
Regardless of the particular development task, the user will
see a split screen, with the majority of it being occupied by
the editor, on the left side, and a smaller portion with the
case study the user is to model, on the right side. Figure
2 presents the starting point for performing the task with
SEA_ML++. Both the textual description of the model to
build, on the right side, and the editor, on the left, are sized
so that the whole exercise can be performed without the
need to resizing or scrolling any window. Figure 3 presents
the starting point for performing the task with DSML4MAS.
Again, window sizes will be similar, and no need to resize or
scroll is expected. Indeed, participants will be instructed not
to change windows sizes, to increase comparability among
sessions. After performing both modelling tasks, participants
are asked to answer a System Usability Scale (SUS) test on
SEA_ML++ and DSML4MAS.</p>
        <p>The tasks involve three different viewpoints: the agent
viewpoint, the MasAndOrg viewpoint and the Interaction
viewpoint. For the sake of illustration, we provide here the
agent viewpoint, in both languages, Figure 4 (SEA_ML++)
and Figure 5 (DSML4MAS). Further materials, including
large-sized versions of these diagrams can be found in our
companion site.</p>
      </sec>
      <sec id="sec-3-4">
        <title>D. Hypotheses, parameters and variables</title>
        <p>For each of our high-level goals, we define the null (H0) and
alternative (H1) hypotheses. Similar hypotheses can be written
for contrasting SEA_ML++ with DSML4MAS in terms of their
effect on correctness, speed, amount of rework, visual effort,
and perceived usability of the languages.</p>
        <p>H0Correctness: Using SEA_ML++ rather than DSML4MAS
does not influence the produced models correctness.</p>
        <p>H1Correctness: Using SEA_ML++ rather than DSML4MAS
influences the produced models correctness.</p>
        <p>H0Speed: Using SEA_ML++ rather than DSML4MAS does
not influence the speed of model production.</p>
        <p>H1Speed: Using SEA_ML++ rather than DSML4MAS
influences the speed of model production.</p>
        <p>H0Rework: Using SEA_ML++ rather than DSML4MAS does
not influence the amount of rework during model production.</p>
        <p>H1Rework: Using SEA_ML++ rather than DSML4MAS
influences the amount of rework during model production.</p>
        <p>H0Effort: Using SEA_ML++ rather than DSML4MAS does
not influence the visual effort involved during model
production.</p>
        <p>H1Effort: Using SEA_ML++ rather than DSML4MAS
influences the visual effort involved during model production.</p>
        <p>H0Usability: Using SEA_ML++ rather than DSML4MAS does
not influence the perceived effort involved during model
production.</p>
        <p>H1Usability: Using SEA_ML++ rather than DSML4MAS
influences the perceived usability of model production.</p>
        <p>1) Assessing correctness: For each of the proposed
challenges, we have a “gold standard” model defined in both
languages, with which we can compare the models built by
our participants. The correctness of the proposed models is
measured in terms of their precision, recall, and F-measure,
defined here as follows:
precision – the percentage of model elements and
relationships in the model built by the participant that
correctly address the challenge (even if the participant
chose alternative ways of modelling the MAS when
compared to the “gold standard”, as long as they are
considered correct.
recall – the percentage of model elements and
relationships in the “gold standard” model that are correctly
addressed by the participant’s model.</p>
        <p>F-measure – a measure that combines precision and
recall, computed as 2 (P recision Recall) ; this measure
(P recision+Recall)
provides an harmonic mean of precision and recall.</p>
        <p>Higher values of precision, recall and the F-measure support
the claim for higher correctness, with 0 representing totally
incorrect and 1 totally correct models.</p>
        <p>2) Assessing speed: We assess speed by measuring the
amount of time (measured in seconds) taken by our participant
to build a MAS model. Lower values of this metric support
the claim of better language usage efficiency.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3) Assessing rework: We assess rework by identifying,</title>
        <p>through the analysis of the model building screencast, the
moments where the participant discarded parts of the solution
he was building (e.g. by removing a previously added element,
or relationship).</p>
        <p>4) Assessing visual effort: We assess visual effort using eye
tracking data collected through the screencast. In particular,
we will analyse heat-maps of the screencasts to compare,
for example, whether there are significant differences in the
amount of time spent exploring modelling options available in
the language toolbar and whether there is some relationship
between these exploring moments and patterns of rework.</p>
        <p>
          Table I outlines our cross-over within subjects design, with
two challenges from different domains (D1 and D2), but with
a similar complexity. Each participant will solve those two
challenges using a different language in each of them. To
cancel learning effects, we will balance the number of times
5) Assessing the perceived usability: We assess the per- the participants start with each of the languages and each of
ceived usability through an SUS questionnaire which provides the problems. In other words, we will balance the participants
a SUS score from 0 to 100, with an average value of 68 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. in groups A, B, C and D.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Higher values support the claim for a better usability.</title>
      <sec id="sec-4-1">
        <title>E. Design</title>
      </sec>
      <sec id="sec-4-2">
        <title>F. Procedure</title>
        <p>
          As depicted in Table I, before starting, each participant
will sign a letter of consent, adapted from [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ] and fill in
a demographic questionnaire, so that we record information
about our participants, including country, age, genre, academic
level, previous experience with MAS and, in particular, with
each of the two analysed languages. This is followed by
viewing a short tutorial on both languages. Then, the subject
will perform an eye tracking device calibration, so that the eye
tracking data of the session can be recorded with precision. To
maximise eye tracking recording precision, participants will
be comfortably seated at a distance of about 60cm from a
full HD 22 inch monitor and instructed not to move much
during the whole session. An EyeTribe eye tracker 2 will be
placed below the monitor. The participant will also have a
keyboard and a mouse, to be able to build a MAS model.
After these preparatory tasks, the experiment itself can start.
During the whole session, a screencast of the contents of the
screen will be recorded. Furthermore, eye tracking data will
also be collected, in sync with the screencast of the session.
The participant will have no time limit to finish his task, but
our pilot sessions point to a duration of about 20 minutes to
perform the given tasks. Finally, the subject answers a SUS
test [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], so that we may contrast his opinions on the usability
of SEA_ML++ and of DSML4MAS.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>G. Analysis procedure</title>
        <p>
          The data collected during the experiment sessions will be
analysed using a combination of automated data collection for
the questionnaires and eye tracking data, with manual data
collection, combining the visual inspection of the screencast with
the synchronised recorded audio of the think aloud protocol.
Concerning descriptive statistics, we will normally collect the
following ones, adjusting the actual set of descriptive statistics
to the scale type (nominal, ordinal, interval or ratio) of each
variable: number of cases, mean, median, mode, standard
deviation, skewness, kurtosis, the p-value of the Shapiro-Wilk
normality test. We will then use appropriate statistics tests.
For example, we plan to use the Welch t test, which is a more
robust alternative to the t-test [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ]) to compare the distributions
of correctness obtained with SEA_ML++ vs. DSML4MAS.
The statistics analysis will be run using SPSS 3.
        </p>
        <p>1) Correctness: The data concerning correctness will be
collected through visual inspection of the solutions created
by the participants in our study. This implies a qualitative
2http://www.theeyetribe.com/
3https://www.ibm.com/analytics/spss-statistics-software
assessment of those solutions in a process which is somewhat
similar to grading the result of a modelling exercise, in an
academic context, following the criteria detailed in section
III-D1. We will then compute descriptive statistics for the
collected metrics and test for significant differences between
the level of correctness achieved with each language.</p>
        <p>2) Speed: The data concerning speed will be collected
during the visual inspection of the screencast of the sessions,
by annotating the timestamps marking the begin and the end of
each task. We will then compute descriptive statistics for the
collected metrics and test for significant differences between
the duration of the tasks using each language.</p>
        <p>3) Rework: The data concerning rework will be collected
through visual inspection of the screencast. In particular, we
will collect and annotate with timestamps events of creation,
deletion, or update of model elements and associations among
those elements. This will provide us with a timeline of the
model construction process for further analysis. Concerning
rework, we will analyse activities that undo previous work (e.g.
a model element that was previously added to the model and
now is deleted). This will allow identifying when the
participant is convinced he made a mistake and decides to backtrack.
Ultimately, we will explore whether the different languages
lead to different levels of rework, both in general, and with
particular sub-groups of participants, divided according to their
background (e.g. by level of expertise with MAS).</p>
        <p>4) Visual effort: The eye tracking data is collected
automatically during the execution of the experiment. This produces
a time series of eye tracking events, namely fixations and
saccades, with their duration, location, etc. The screen area
will be annotated with relevant areas of interest, so that we
can use the eye tracking data to monitor how each participant
navigated through those areas, during the process. We will use
custom-made tools from the NOVA LINCS team to support
this analysis. In the end, we expect to use heat maps to analyse
where the most important focuses of visual attention were, and
scanpath analysis to better understand the model navigation
strategies of our participants.</p>
        <p>5) Perceived usability: We will assess usability through
a SUS test. The SUS instrument is available in the testing
environment as a web form. The collected data will be directly
fed into SPSS so that we may proceed with the comparative
analysis of the distributions of the usability scores.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>IV. DISCUSSION</title>
      <sec id="sec-5-1">
        <title>A. Expected results and implications</title>
        <p>
          We are interested in assessing how the usability is influenced
by the selection of one of these languages over the other.
Rather than using these results as a way of promoting the
usage of one of the languages, our goal is to identify language
improvement opportunities, on the one hand, and learning
from the “competition”, on the other. This process is, in that
sense, similar to the one the NOVA LINCS team has followed
for supporting the Gyro language evolution [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] through a
series of developer experience evaluations. We have advocated
elsewhere [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] that software language development should
be iterative and incremental, including (possibly lightweight)
evaluations after each iteration, so that improvement
opportunities are identified as soon as possible, and, when feasible
and adequate, followed on in the next version of the language.
        </p>
        <p>
          Apart from the more “traditional” analysis of effectiveness,
here regarded from the perspective of correctness, and
efficiency, viewed considering the speed, we expect our
exploratory study on the process of building the models, with
an analysis of the time annotated sequences of insertions,
deletions and changes while constructing models to provide
us insights on the main bottlenecks language users experience
during the model building process and, conversely, where they
seem to experience less difficulties. The eye tracking data is
expected to provide further context for better identifying
language improvement opportunities. In a longer run, the lessons
learned in this and similar studies have the potential for
helping us designing more usable software modelling languages.
This will also help us better understanding how people from
different backgrounds interact with each modelling language,
building on earlier works that explored how different personal
characteristics (e.g. gender) impacted on the learning, problem
solving and information processing style [
          <xref ref-type="bibr" rid="ref39">39</xref>
          ]. Finally, the SUS
usability questionnaire will help us better understanding how
the differences between both languages impact usability.
B. Threats to validity
        </p>
        <p>1) Conclusion validity: Although we plan to have a
reasonable amount of participants (over 30), considering the nature of
this study, sample size is a likely threat, due to the difficulty in
recruiting participants. Our mitigation strategy is to have two
teams performing the study in two different countries. The
exercise of preparing the experimental replication package so
that it can be run both in Portugal and Turkey will help us
fine tune it making the package more reusable to third-party
replications. This will directly mitigate the sample size risk,
as we will have participants in both countries, and indirectly,
by facilitating potential third-party replications.</p>
        <p>2) Internal validity: There is a potential learning effect
from solving one challenge to the next. We mitigate this risk
by having the crossover design so that half of the
participants start with a SEA_ML++ model while the other starts
with DSML4MAS. Another threat could be that a particular
problem would by accident favour one of the languages. To
mitigate it, both problems will be modelled in both languages,
by different participants. We chose two languages for which
the tool support is at a similar level, and with a close look
and feel, so that tooling does not play a role in differentiating
among the two DSLs. We also made efforts so that all
materials were easily readable in a 22 inch monitor and that
the models to be developed would fit nicely in a canvas on
this kind of monitor, without requiring the user to scroll or
zoom the image. Monitor size and the general layout for the
experiment, including the distance of the participant to the
monitor were constrained by the technical specifications of
our eye tracking device. In spite of these constraints, the tasks
are already challenging to our participants.</p>
        <p>3) External validity: Our participants will not have, in
general, much experience with MAS and with the two languages.
As such, our participants are better representatives of
developers who are learning these languages. Further research is
necessary to assess how these languages compare, when used
by modellers who are experienced with the two languages. The
conclusions of this study will be applicable to these two MAS
DSMLs. Replications with other languages, not necessarily
for MAS, are required before we can generalise this study’s
conclusions to other contexts.</p>
        <p>4) Construct validity: After watching a short tutorial about
both languages, participants will solve a couple of challenges,
one with each language. This may cause an evaluation
apprehension threat. We mitigate this by informing participants that
the languages are being evaluated, not the participants. The
experimental process is built so that we express no bias toward
any of the languages, to mitigate the risk of accidentally
favouring SEA_ML++. Our goal is to identify opportunities
to improve SEA_ML++ rather than the comparison with
DSML4MAS itself. Our measures to mitigate this risk include
choosing for the author of the recorded tutorials someone with
no vested interest in any of the languages and doing the same
for the researchers performing the data analysis. Further, in the
interest of transparency and replicability, the data used in these
evaluations and data analysis scripts for SPSS will be made
publicly available. Last, but not the least, this paper discussing
the experimental design to be used in this evaluation serves as
a manifest of interest in performing this particular experiment.
This creates an opportunity for a sanity check, where the initial
goals of this study will be directly comparable with what is
actually tested in the experiment, and reported later, mitigating
the potential for selective publishing, where only favourable
results would be published.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>V. SUMMARY</title>
      <p>We presented the experimental planning for the evaluation
of the case study of DSLs for Multi-agents Systems. Our goal
is to go beyond the evaluation of the language’s notations
(concrete syntax) and evaluate the constructs composition at
the level of the instance sentence level (abstract syntax).</p>
      <p>It is expected that the results of the evaluation planned
in this paper will help in identifying effective improvement
opportunities for the developer experience with SEA_ML++.</p>
      <p>The work triggers future research in that it departs from the
more commonly explored part of visual modelling languages
(their visual notation) to other relevant perspectives, namely
at the instance (sentence) level.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENT</title>
      <p>The authors would like to thank the following: i) the
Scientific and Technological Research Council of Turkey
(TUBITAK) under grant 115E591, and ii) Portuguese grants
NOVA LINCS Research Laboratory (Grant: FCT/MCTES
PEst UID/ CEC/04516/2013) and DSML4MAS Project
(Grant: FCT/MCTES TUBITAK/0008/2014).</p>
      <p>The authors would also like to thank the COST Action
networking mechanisms and support of IC1404 Multi-Paradigm
Modeling for Cyber-Physical Systems (MPM4CPS). COST is
supported by the EU Framework Programme Horizon 2020.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Barisic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Amaral</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          , “
          <article-title>Usability driven DSL development with USE-ME,”</article-title>
          <source>Computer Languages, Systems &amp; Structures</source>
          , vol.
          <volume>51</volume>
          , pp.
          <fpage>118</fpage>
          -
          <lpage>157</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mernik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heering</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Sloane</surname>
          </string-name>
          , “
          <article-title>When and how to develop domain-specific languages,” ACM Comput</article-title>
          . Surv., vol.
          <volume>37</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>316</fpage>
          -
          <lpage>344</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Moody</surname>
          </string-name>
          , “
          <article-title>The “physics” of notations: toward a scientific basis for constructing visual notations in software engineering,” IEEE T Software Eng</article-title>
          , vol.
          <volume>35</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>756</fpage>
          -
          <lpage>779</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Challenger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Demirkol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Getir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mernik</surname>
          </string-name>
          , G. Kardas, and T. Kosar, “
          <article-title>On the use of a domain-specific modeling language in the development of multiagent systems</article-title>
          ,
          <source>” Eng Appl Artif Intel</source>
          , vol.
          <volume>28</volume>
          , pp.
          <fpage>111</fpage>
          -
          <lpage>141</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Miranda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Challenger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. T.</given-names>
            <surname>Tezel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. F.</given-names>
            <surname>Alaca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Amaral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          , and G. Kardas, “
          <article-title>Improving the usability of a mas dsml</article-title>
          ,” in 6th International Workshop on Engineering Multi-Agent
          <source>Systems (EMAS</source>
          <year>2018</year>
          ). Stockholm, Sweden: Springer, July, 14
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hahn</surname>
          </string-name>
          , “
          <article-title>A domain specific modeling language for multiagent systems</article-title>
          ,”
          <source>in Proceedings of the 7th international joint conference on Autonomous agents and multiagent systems-Volume</source>
          <volume>1</volume>
          ,
          <year>2008</year>
          , pp.
          <fpage>233</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Brooke</surname>
          </string-name>
          et al., “
          <article-title>Sus-a quick and dirty usability scale,” Usability evaluation in industry</article-title>
          , vol.
          <volume>189</volume>
          , no.
          <issue>194</issue>
          , pp.
          <fpage>4</fpage>
          -
          <lpage>7</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kardas</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Gomez-Sanz</surname>
          </string-name>
          , “
          <article-title>Special issue on model-driven engineering of multi-agent systems in theory and practice</article-title>
          ,
          <source>” Comput Lang Syst Str</source>
          , vol.
          <volume>50</volume>
          , pp.
          <fpage>140</fpage>
          -
          <lpage>141</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Beydoun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Low</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Henderson-Sellers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mouratidis</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. J. GomezSanz</surname>
          </string-name>
          , J. Pavon, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Gonzalez-Perez</surname>
          </string-name>
          , “
          <article-title>Faml: a generic metamodel for mas development</article-title>
          ,
          <source>” IEEE T Software Eng</source>
          , vol.
          <volume>35</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>841</fpage>
          -
          <lpage>863</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Ciobanu</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Juravle</surname>
          </string-name>
          , “
          <article-title>Flexible software architecture and language for mobile agents</article-title>
          ,” Concurr
          <string-name>
            <surname>Comp-Pract</surname>
            <given-names>E</given-names>
          </string-name>
          , vol.
          <volume>24</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>559</fpage>
          -
          <lpage>571</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>J. M. Gascueña</surname>
          </string-name>
          , E. Navarro,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Fernández-Caballero</surname>
          </string-name>
          , “
          <article-title>Model-driven engineering techniques for the development of multi-agent systems</article-title>
          ,
          <source>” Eng Appl Artif Intel</source>
          , vol.
          <volume>25</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>159</fpage>
          -
          <lpage>173</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E. J. T.</given-names>
            <surname>Gonçalves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Cortés</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A. L.</given-names>
            <surname>Campos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. S.</given-names>
            <surname>Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Freire</surname>
          </string-name>
          , V. T. da
          <string-name>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. S. F. de Oliveira</surname>
          </string-name>
          , and M. A. de Oliveira, “
          <article-title>Masml 2.0: Supporting the modelling of multi-agent systems with different agent architectures,” J Syst Software</article-title>
          , vol.
          <volume>108</volume>
          , pp.
          <fpage>77</fpage>
          -
          <lpage>109</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Demirkol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Challenger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Getir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kosar</surname>
          </string-name>
          , G. Kardas, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mernik</surname>
          </string-name>
          , “
          <article-title>A dsl for the development of software agents working within a semantic web environment</article-title>
          ,
          <source>” Computer Science and Information Systems</source>
          , vol.
          <volume>10</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>1525</fpage>
          -
          <lpage>1556</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kardas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. T.</given-names>
            <surname>Tezel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Challenger</surname>
          </string-name>
          , “
          <article-title>Domain-specific modelling language for belief-desire-intention software agents</article-title>
          ,
          <source>” IET Softw</source>
          , vol.
          <volume>12</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>356</fpage>
          -
          <lpage>364</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Faccin</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Nunes</surname>
          </string-name>
          , “
          <article-title>A tool-supported development method for improved bdi plan selection</article-title>
          ,
          <source>” Engineering Applications of Artificial Intelligence</source>
          , vol.
          <volume>62</volume>
          , pp.
          <fpage>195</fpage>
          -
          <lpage>213</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Moody</surname>
          </string-name>
          and J. van Hillegersberg,
          <article-title>“Evaluating the visual syntax of uml: An analysis of the cognitive effectiveness of the uml family of diagrams,”</article-title>
          <source>in International Conference on Software Language Engineering</source>
          . Springer,
          <year>2008</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>D</surname>
          </string-name>
          . L. Moody, P. Heymans, and R. Matulevicˇius, “
          <article-title>Visual syntax does matter: improving the cognitive effectiveness of the i* visual notation,” Requir Eng</article-title>
          , vol.
          <volume>15</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>141</fpage>
          -
          <lpage>175</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Challenger</surname>
          </string-name>
          , G. Kardas, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Tekinerdogan</surname>
          </string-name>
          , “
          <article-title>A systematic approach to evaluating domain-specific modeling language environments for multi-agent systems</article-title>
          ,” Software Qual
          <string-name>
            <surname>J</surname>
          </string-name>
          , vol.
          <volume>24</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>755</fpage>
          -
          <lpage>795</lpage>
          , Sep.
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kardas</surname>
          </string-name>
          , E. Bircan, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Challenger</surname>
          </string-name>
          , “
          <article-title>Supporting the platform extensibility for the model-driven development of agent systems by the interoperability between domain-specific modeling languages of multiagent systems</article-title>
          ,
          <source>” Comput Sci Inf Syst</source>
          , vol.
          <volume>14</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>875</fpage>
          -
          <lpage>912</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>F.</given-names>
            <surname>Bergenti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Iotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Monica</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Poggi</surname>
          </string-name>
          , “
          <article-title>Agent-oriented modeldriven development for jade with the jadel programming language,” Comput Lang Syst Str</article-title>
          , vol.
          <volume>50</volume>
          , pp.
          <fpage>142</fpage>
          -
          <lpage>158</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>El Kouhen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gherbi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dumoulin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Khendek</surname>
          </string-name>
          , “
          <article-title>On the semantic transparency of visual notations: Experiments with uml</article-title>
          ,” in
          <source>International SDL Forum</source>
          . Springer,
          <year>2015</year>
          , pp.
          <fpage>122</fpage>
          -
          <lpage>137</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>N.</given-names>
            <surname>Genon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Heymans</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Amyot</surname>
          </string-name>
          , “
          <article-title>Analysing the cognitive effectiveness of the bpmn 2.0 visual notation</article-title>
          ,”
          <source>in Proceedings of the Third International Conference on Software Language Engineering</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>377</fpage>
          -
          <lpage>396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>D</surname>
          </string-name>
          . L. Moody, “
          <article-title>Why a diagram is only sometimes worth a thousand words: An analysis of the bpmn 2.0 visual notation</article-title>
          ,
          <source>” Hämtat</source>
          <year>2012</year>
          -
          <volume>06</volume>
          -19 från http://www. business. uq. edu. au/sites/default/files/event/supportingD ocs/
          <source>Analysis% 20of% 20BPMN% 202.0% 20Visual% 20Syntax. pdf, Tech. Rep.</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>R.</given-names>
            <surname>Matulevicˇius</surname>
          </string-name>
          and P. Heymans, “
          <article-title>Visually effective goal models using kaos</article-title>
          ,” in International Conference on Conceptual Modeling. Springer,
          <year>2007</year>
          , pp.
          <fpage>265</fpage>
          -
          <lpage>275</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R.</given-names>
            <surname>Matulevicius</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Heymans</surname>
          </string-name>
          , “
          <article-title>Comparing goal modelling languages: An experiment</article-title>
          ,” in International Working Conference on Requirements Engineering: Foundation for Software Quality,
          <year>2007</year>
          , pp.
          <fpage>18</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>M.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gralha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          , and J. a. Araujo, “
          <article-title>Increasing the semantic transparency of the kaos goal model concrete syntax</article-title>
          ,
          <source>” in 37th International Conference on Conceptual Modeling (ER</source>
          <year>2018</year>
          ).
          <article-title>Xi'an</article-title>
          , China: Springer, October,
          <fpage>22</fpage>
          -
          <lpage>25</lpage>
          2018.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>P.</given-names>
            <surname>Caire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Genon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Heymans</surname>
          </string-name>
          , and D. L. Moody, “
          <article-title>Visual notation design 2.0: Towards user comprehensible requirements engineering notations</article-title>
          ,” in RE'13. IEEE,
          <year>2013</year>
          , pp.
          <fpage>115</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>N.</given-names>
            <surname>Genon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Caire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Toussaint</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Heymans</surname>
          </string-name>
          , and D. Moody, “
          <article-title>Towards a more semantically transparent i* visual syntax</article-title>
          ,” in International Working Conference on Requirements Engineering: Foundation for Software Quality,
          <year>2012</year>
          , pp.
          <fpage>140</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>M.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gralha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>a. Araujo, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreira</surname>
          </string-name>
          , “
          <article-title>On the impact of semantic transparency on understanding and reviewing social goal models</article-title>
          ,
          <source>” in 26th IEEE International Conference on Requirements Engineering (RE</source>
          <year>2018</year>
          ). Banff, Canada: IEEE,
          <year>August</year>
          ,
          <fpage>20</fpage>
          -
          <lpage>24</lpage>
          2018.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>H.</given-names>
            <surname>Henriques</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lourenço</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Amaral</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          , “
          <article-title>Improving the developer experience with a low-code process modelling language,”</article-title>
          <source>in ACM/IEEE 21st International Conference on Model Driven Engineering Languages and Systems (MODELS)</source>
          . Copenhagen, Denmark: ACM,
          <year>October 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>H.</given-names>
            <surname>Störrle</surname>
          </string-name>
          , “
          <article-title>On the impact of layout quality to understanding uml diagrams,” in Visual Languages and Human-Centric Computing (VL/HCC</article-title>
          ),
          <source>2011 IEEE Symposium on. IEEE</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>135</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>H.</given-names>
            <surname>Storrle</surname>
          </string-name>
          , “
          <article-title>On the impact of layout quality to understanding uml diagrams: Diagram type and expertise,” in Visual Languages and HumanCentric Computing (VL/HCC</article-title>
          ),
          <source>2012 IEEE Symposium on. IEEE</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>49</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>H.</given-names>
            <surname>Störrle</surname>
          </string-name>
          , “
          <article-title>On the impact of layout quality to understanding uml diagrams: size matters</article-title>
          ,
          <source>” in International Conference on Model Driven Engineering Languages and Systems</source>
          . Springer,
          <year>2014</year>
          , pp.
          <fpage>518</fpage>
          -
          <lpage>534</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>M.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gralha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Araújo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreira</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Cambeiro</surname>
          </string-name>
          , “
          <article-title>What is the impact of bad layout in the understandability of social goal models?” in 24th IEEE International Requirements Engineering Conference</article-title>
          (RE'16). Beijing, China: IEEE, September,
          <fpage>12</fpage>
          -
          <lpage>16</lpage>
          2016.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Barišic´</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cambeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Amaral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          , and T. Mota, “
          <article-title>Leveraging teenagers feedback in the development of a domain-specific language: the case of programming low-cost robots</article-title>
          ,”
          <source>in Proceedings of the 33rd Annual ACM Symposium on Applied Computing. ACM</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1221</fpage>
          -
          <lpage>1229</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basili</surname>
          </string-name>
          , G. Caldiera, and
          <string-name>
            <given-names>H.</given-names>
            <surname>Rombach</surname>
          </string-name>
          , “Goal Question Metric Paradigm,” Encyclopedia of Software Eng., vol.
          <volume>1</volume>
          , pp.
          <fpage>528</fpage>
          -
          <lpage>532</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>P.</given-names>
            <surname>Runeson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Host</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rainer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Regnell</surname>
          </string-name>
          ,
          <article-title>Case study research in software engineering: Guidelines and examples</article-title>
          . Wiley,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>B. L.</given-names>
            <surname>Welch</surname>
          </string-name>
          , “
          <article-title>The generalization of 'student's' problem when several different population variances are involved</article-title>
          ,
          <source>” Biometrika</source>
          , vol.
          <volume>34</volume>
          , no.
          <issue>1-2</issue>
          , pp.
          <fpage>28</fpage>
          -
          <lpage>35</lpage>
          ,
          <year>1947</year>
          . [Online]. Available: http: //dx.doi.org/10.1093/biomet/34.1-
          <fpage>2</fpage>
          .
          <fpage>28</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>L.</given-names>
            <surname>Beckwith</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Burnett</surname>
          </string-name>
          , “
          <article-title>Gender: An important factor in end-user programming environments?” in Visual Languages and Human Centric Computing, 2004 IEEE Symposium on</article-title>
          . IEEE,
          <year>2004</year>
          , pp.
          <fpage>107</fpage>
          -
          <lpage>114</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>