<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Analysis of the Performance of Italian Schools in Bebras and in the National Student Assessment INVALSI</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Violetta Lonati</string-name>
          <email>violetta.lonati@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universita degli Studi di Milano</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Anna Morpurgo</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Computer Science</institution>
          ,
          <addr-line>Via Celoria 18, 20133 | Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <abstract>
        <p>This paper analyzes the results of the Bebras Challenge on Informatics and Computational Thinking held in Italy in the last three years and it compares them to the overall performance of Italian schools in the national INVALSI assessment of the standardized levels reached by students in Italian, Mathematics, and English. The main research question is if the mean regional performance at INVALSI tests can predict the performance of schools of the same region in the Bebras challenge. The answer is positive at the grossest level: macro regional areas with INVALSI results below the national average tend to perform worse also in the Bebras challenge. At regional level, a high correlation between Bebras and INVALSI was found among the regions whose results di er signi cantly.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The Bebras International Challenge on Informatics
and Computational Thinking (http://bebras.org) is
a yearly contest organized since 2004 [Dag10, HCD11].
In 2018 almost three million participants from 54
countries took part to one of the locally organized events.
The contest, open to pupils of all school levels (from
primary up to upper secondary), is based on tasks
rooted on core informatics concepts and computational
thinking, yet independent of speci c previous
knowledge such as for instance that acquired during
curricular activities. In fact Bebras tasks avoid the use of
jargon and are especially aimed at a non-vocational
audience, focusing on that part of informatics that
should become familiar to everyone, not just
computing professionals. The tasks are supposed to provide
an entertaining learning experience, and they are
designed by the Bebras community to be moderately
challenging and solvable in a relatively short time.
The setting of the contest is slightly di erent in each
country, but in general participants have to solve a
set of about 10-15 tasks in an average time of three
minutes for each. In Italy, the Bebras is open to
teams of 3 or 4 pupils, divided in ve age groups:
I (grades 4{5, ages 9{10), II (grades 6{7, ages 11{
12), III (grade 8, age 13), IV (grades 9{10, ages 14{
15), V (grades 11{13, ages 16{18). In the last three
editions we had 36,018 teams, from schools located in
all the 20 administrative regions Italy is subdivided
into (see Table 3). Besides being used during the
contests, Bebras tasks are an opportunity for
educational activities [DS16, LMM+17, CAC+18].
Moreover, Bebras was used to measure improvements of
students' attitude to computational thinking [SBS17].
The study examined 21 schools (children aged 9{11)
which participated in \Code Clubs". The primary
outcome measure was a set of Bebras tasks, which 317
pupils completed at baseline and endpoint. We
wonder, instead, if the performances in Bebras follow the
general level of competencies of the schools
participating to the contest. In order to answer this question,
one should have a measure of the curricular
achievements of the schools (or even the classes) involved,
but unfortunately these data are not publicly available.
In fact, one of the Bebras' goals is to spread the
acquaintance with informatics and computational
thinking among every school population, even (or maybe
especially) those not naturally attracted by computing.
To this end, we avoid any participation fee and we try
to keep the competition at a level such that nobody
should feel ashamed to participate: Bebras should be
perceived as an opportunity to have fun and learn
something, not to show o the performances of the
schools. For example in Italy, although every teacher
receives ranking data about their teams, only the very
top of the ranking is published (the best eight teams
in each age group, with at most one team per school).
Thus, we do not want to ask teachers about the marks
of their pupils in the curricular activities or other
proxies of their academic success. Instead, we tried to
understand if the results in the Italian Bebras contest
were somewhat correlated with the general school
performances in the same territory. For this, we resort to
INVALSI data, the national student assessment
program, similar to OECD's Programme for International
Student Assessment (PISA) or IEA's Trends in
International Mathematics and Science Study (TIMSS)
and Progress in International Reading Literacy Study
(PIRLS).</p>
      <p>Since school year 2005/6, all the pupils of the
Italian school system at the end of grade 2, 5, 8, and 10
are evaluated by an INVALSI standardized test, aimed
at measuring their pro ciency in Italian,
Mathematics, English listening and English reading.
According to the 2018 INVALSI report, the performances of
the twenty Italian regions di er in a signi cant way,
at least from grade 8 and up. Thus, we set up a
study aimed at understanding if these di erences are
re ected in the results we see in the Bebras contest.
The number of Bebras teams is much smaller than the
number of students involved in the INVALSI
assessment (even by considering that their public data are
based on a sample, see below), moreover Bebras
participation depends on teachers' interest, while INVALSI
is mandatory. Nevertheless, we wanted to understand
if Bebras data re ect the general geographic pattern
of the wider population of Italian schools.</p>
      <p>The paper is organized as follows: in Section 2 we
formalize our research questions, in Section 3 we
describe our approach, in Section 4 we report our
analyses, and nally in Section 5 we draw some conclusions.</p>
    </sec>
    <sec id="sec-2">
      <title>The research questions</title>
      <p>The 2018 INVALSI assessment [INV18] tested 29,520
grade 5 classes (562,635 pupils), 29,032 grade 8 classes
(574,506 pupils), and 26,361 grade 10 classes (543,296
pupils). In order to guarantee data quality, a sample
of students was observed directly during the test: the
data reported publicly is based on this direct analysis
of 29,371 grade 5 students, 31,300 grade 8 students,
and 48,664 grade 10 students. In grade 5 the test is
paper based and manually marked, while in the other
grades the test is computer based and automatically
marked. Results are separately assessed for four
areas of competence: `Italian', `Mathematics', `English
listening', and `English reading' (in 2018, grade 10
was not tested for English). Public data cover all the
twenty Italian regions (Trentino-Alto Adige is actually
divided into two autonomous provinces, since in the
region live communities with di erent mother-tongues,
no aggregated regional data are provided). Results are
provided at two levels of aggregations:
1. ve geographic macro-areas: North-West,
North</p>
      <p>East, Center, South, South-Islands;
2. 21 administrative regions (19 regions and 2
autonomous provinces).</p>
      <p>Table 1 shows the mean performance by area; the
average is set at 200 (with a standard deviation of 40).</p>
      <p>According to the INVALSI report [INV18], the
differences among the areas at grade 5 are small1.
Instead, the di erences are considered increasingly
signi cant in the higher grades. Overall they are claimed
to match similar results in PISA assessment (surveyed
internationally every three years) with the North part
of the country performing better than the national
average, and the South part worse than the national
average; the Center instead re ects the national average.
The report also mentions that the Northern part of
the country has better than average results in recent
TIMMS assessments.</p>
      <p>The regional data are more detailed, since they
report also the standard deviation of the distributions,
not only the means. The data are shown in Table 2.</p>
      <p>In this study, our goal is to understand if this
variability is re ected in the results of the Italian Bebras.
We have homogeneous data for the last three editions
(2016, 2017, 2018). The total number of partecipating
teams is reported in Table 3.</p>
      <p>Bebras data involve a smaller number of schools
with respect to INVALSI (which aims at being
\universal" in the Italian school system: the participation
1Grade 2 has even smaller di erences; it was not considered
here, since the Italian Bebras involves pupils from grade 4 up to
grade 13
area
Center
North-East
North-West
South
South-Islands
Center
North-East
North-West
South
South-Islands
Center
North-East
North-West
South
South-Islands
grade
is mandated by law. In the past it was also used to
mark students at grade 8, but the 2018 edition was
not used for this purpose). Nevertheless we would like
to use them to try to answer the following research
questions.</p>
      <p>Is there any correlation between the average ability
of Bebras teams in a speci c region and the regional
performance in INVALSI tests?
Is there any correlation between the average ability of
Bebras teams in a geographic macro area and the area
performance in INVALSI tests?
Is the overall performance trend at INVALSI tests,
with Northern schools performing better than the
national average and Southern schools performing worse,
re ected also in Bebras results?
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>We estimated the ability of the Bebras teams by tting
an Item Response Theory (IRT) [HS85] model with
two parameters. IRT is routinely used to evaluate
massive educational assessment studies like OECD's PISA,
and it has already been applied to Bebras and other
informatics competitions [KVC06, HM14, BLM+15].
Moreover, a similar IRT model is behind the INVALSI
data as described in [Des18].</p>
      <p>IRT models each solver with an ability ( )
parameter and links it to the probability of a correct solution
via a logistic function. Such a function is a
characteristic of each task (item) and it de nes its response
to the solver ability. Response functions are described
by a number of parameters: we used a model with two
parameters, the di culty ( ) of a task and its
discrimination ( ). Di culty locates the response function:
if the ability of the solver is greater than the di culty
of a task, the probability of solving it is greater than
0:5. Discrimination de nes the slope of the response
curve: a high discrimination means that a small
increase in the ability of the solver has a great impact
on the probability of solving the task; a discrimination
= 0 de nes a task in which the ability of the solver does
not matter at all. Figure 1 shows some examples of
logistic response functions. It is worth noting that all
that counts in the model are the relative values of the
parameters (there is no absolute measure of ability):
thus to t it to data it is necessary to identify ability
with conventional values. In order to be comparable
with INVALSI data, we deviated from the common
practice [GH06] of assuming that, overall, ability has
mean = 0 with respect to an arbitrary reference point
and standard deviation = 1. Instead, we assumed a
mean ability = 200 and a standard deviation = 40.</p>
      <p>In order to estimate the di culty and
discrimination of each task, we implemented the probabilistic
model with Stan [Sta16]. Stan is a software tool which,
given a statistical model, uses Hamiltonian Monte
Carlo sampling (a very e cient form of Markov chain
Monte Carlo sampling) to approximate the posterior
probability of the parameters of interest.</p>
      <p>P ( ijY ) i 2 teams
(1)
where i is the ability of team i. The statistical model
sampled is a hierarchical one, with the following prior
distributions:
ABRUZZO
BASILICATA
CALABRIA
CAMPANIA
EMILIA-ROMAGNA
FRIULI-VENEZIA GIULIA
LAZIO
LIGURIA
LOMBARDIA
MARCHE
MOLISE
PIEMONTE
PUGLIA
SARDEGNA
SICILIA
TOSCANA
TRENTINO-ALTO ADIGEa
UMBRIA
VALLE D'AOSTA
VENETO
ABRUZZO
BASILICATA
CALABRIA
CAMPANIA
EMILIA-ROMAGNA
FRIULI-VENEZIA GIULIA
LAZIO
LIGURIA
LOMBARDIA
MARCHE
MOLISE
PIEMONTE
PUGLIA
SARDEGNA
SICILIA
TOSCANA
TRENTINO-ALTO ADIGEa
UMBRIA
VALLE D'AOSTA
VENETO
area
area
Center
Center
Center
Center
North-East
North-East
North-East
North-East
North-West
North-West
North-West
North-West
South
South
South
South
South-Islands
South-Islands
South-Islands
South-Islands
LAZIO
MARCHE
TOSCANA
UMBRIA
Total
EMILIA-ROMAGNA
FRIULI-VENEZIA GIULIA
TRENTINO-ALTO ADIGE
VENETO
Total
LIGURIA
LOMBARDIA
PIEMONTE
VALLE D'AOSTA
Total
ABRUZZO
CAMPANIA
MOLISE
PUGLIA
Total
BASILICATA
CALABRIA
SARDEGNA
SICILIA
Total</p>
      <p>Grade 5</p>
      <p>Grade 8</p>
      <p>Grade 10</p>
      <p>In this model we assumed a Cauchy weakly
informative prior distribution on hyper-parameters | the
mean di culty used as a reference point in the
logistic |, , and | the standard deviation
respectively of di culty and discrimination |. The ability
is then supposed to be normally distributed with mean
= 200 and standard deviation = 40, the di culty
normally distributed with mean = 0 and standard
deviation = , and the logarithm of discrimination is
normally distributed with mean = 200 and standard
deviation = . The correctness y of each item is
nally sampled according to a Bernoulli process where
the probability of success is computed with the
logistic model described above. These are quite standard
choices for Bayesian IRT (see [GH06, Sta16]). We
sampled the Stan Monte Carlo model for 2,000 iterations,
throwing away the rst 1000 results (50% warm-up
iterations). The results have all the typical properties
of converging models, in particular the R^ statistics is
close to 1 for every parameter of interest (a necessary,
but unfortunately not su cient, condition for
convergence). Results are indeed sensible, with descriptive
statistics consistent with score data, therefore we are
rather con dent that our model is plausible and useful
to infer latent parameters.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Data analysis</title>
      <p>
        In order to answer the research questions posed in
Section 2, we start by identifying which variations among
Bebras data are indeed signi cant. Ideally, we would
like to lter out the di erences due to statistical
uctuations. In fact, even the
        <xref ref-type="bibr" rid="ref4">INVALSI 2018</xref>
        report warns
the readers that the di erences in grade 5 results are
too small to be considered a true assessment of the
local competencies [INV18]. Unfortunately the report
does not give enough details to replicate the signi
cance test they used. We used a t -test between each
pair of areas and regions, and we considered as signi
cant those in which the t -test has a p-value &lt; 1 10 4
(i.e., the \null" hypothesis that the two generating
distributions have the same mean is less probable than
100100 ). Table 4 collects the signi cance of the
results grouped by macro-area: only a few di erences
are signi cant at grade 5, but the overall signi cance
increases with grades 8 and 10.
      </p>
      <p>A similar pattern is also found when the results are
grouped by regions, as reported in Table 5.
4.1</p>
      <sec id="sec-4-1">
        <title>Analysis at the regional level</title>
        <p>When one considers Bebras and INVALSI results
grouped by region, the correlation among the
rankings of the means is rather low. Tables 6,7, and 8 give
the Kendall rank correlation coe cients respectively
for grade 5, grade 8, and grade 10. The correlation
increases with grades, but several inversions among the
rankings remain.</p>
        <p>In order to also appreciate the impact of the
standard deviation of the results, we give the pictures of
the distributions too (see Figures 2, 3, and 4,
respectively grade 5, 8, and 10), approximated with a
Gaussian with the same mean and standard deviation.</p>
        <p>We also investigated if, whenever the di erence in
Bebras results between two regions is considered
signi cant (see Table 5), the di erence is in the same
\direction" of the di erence in INVALSI (please note,
however, that we do not have detailed enough data to
test if the di erence in INVALSI results is also signi
cant). For example, VENETO and CAMPANIA have
a signi cant di erence in Bebras results: VENETO
performed better than CAMPANIA, and the same is
true with respect to INVALSI tests.</p>
        <p>For grade 5, we found 10 signi cant di erences
between regions, the di erences have the same direction
for 4 pairs. In the other 6 pairs, the directions di er:
Bebras di erence has the same direction of `English
reading' in 5 cases, of `English listening' in 4 cases, of
`Italian' in 4 cases, of `Mathematics' in 4 cases; thus,
17 cases out 24 are in the same direction.</p>
        <p>For grade 8, we found 31 signi cant di erences
between regions and the di erences have the same
direction for all.</p>
        <p>For grade 10, we found 78 signi cant di erences
between regions, the di erences have the same direction
for 63 pairs. In the other 15 pairs, the directions
differ: Bebras di erence has the same direction of
`Italian' in 1 case, in all other 29 cases the direction of
Bebras di erence is opposite of the di erence in
Italian and Mathematics, which instead are consistent
between them.</p>
        <p>All in all, we believe we have preliminary evidence
that the answer to RQ1 is somewhat positive: at least
when the di erence is signi cant, the di erence in
Bebras mostly matches INVALSI di erences.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Analysis at the level of macro-areas</title>
        <p>With the exception of grade 5 (see Table 9, but at
this grade, as noted above, the di erences are mostly
not signi cant), the correlation among the rankings
of the means grouped by macro-areas is rather high.
Tables 10 and 11 give the Kendall rank correlation
coe cients respectively for grade 8 and grade 10.</p>
        <p>Thus, also for RQ2 we believe we have evidence
to answer positively, at least for the grades 8 and 10,
where the di erences between the results of the
macroareas are considered signi cant.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Analysis at the grossest level</title>
        <p>
          The
          <xref ref-type="bibr" rid="ref4">INVALSI 2018</xref>
          report claims that the overall
INVALSI results generally match PISA results: the
Northern part of Italy performs better than the
national average, while the Southern part performs
worse. This pattern, with the best mean results in
the two Northern macro-areas and the worst mean
results in the two Southern macro-areas, is found also in
Bebras. According to Bebras data, the Center
macroarea performs slightly below the national average.
        </p>
        <p>Thus, RQ3 seems also positively supported by our
data.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Threats to validity</title>
        <p>The 2018 INVALSI report does not give the details
about the signi cance tests used to mark the di
erences at grade 5 as not signi cant, while at grades 8
and 10 they were considered so. Also, no pairwise (at
both regional and macro-area levels) signi cance was
reported. Since the Bebras sample is much smaller,
we used a rather tight criterion: a t -test with a
pvalue threshold &lt; 1 10 4. The underlying
statistical model is the same in INVALSI and Bebras
(2parameter IRT), but we do not know the tting
apArea
Center
North-East
North-West
South
South-Islands</p>
        <p>Center</p>
        <p>North-East</p>
        <p>North-West
|
10
5
8
grades in which the p-value is less than 1
distributions have the same mean.
10 3, the threshold we used to reject the hypothesis that the two
10 4, the threshold we used to reject the hypothesis that the two distributions
Italian</p>
        <p>Mathematics</p>
        <p>for grade 10 INVALSI and Bebras
proach used in INVALSI: to get numerically
comparable results we used Normal distributions located in
200, with scale of 40. We adopted sensible prior
parameter choices, common in the IRT literature, but we
do not know if a di erence considered signi cant in our
model would be marked as such also by the INVALSI
approach.</p>
        <p>The main threat to validity, however, is the bias
intrinsic in the Bebras sample. While INVALSI data
cover every school in Italy and the sample surveyed in
[INV18] was supposedly chosen with statistical goals
in mind, we just used all the data of the teams who
participated to the last three editions of the Italian
Bebras and were able to ship a result with our
online platform [BCL+18]. Bebras pupils are thus drawn
from the classes and schools with teachers interested
in computational thinking and informatics (although
this special interest is not necessarily shared by their
pupils) and had the equipment and the logistic context
suitable to participate. Also, while INVALSI tests
individuals, Bebras is played in teams of 3{4 students.</p>
        <p>For INVALSI we used the data as reported in
[INV18], since we have no access to raw data. The
data source is incomplete, for example no pieces of
information are given about the numbers of sampled
students by region or even macro-area. This makes
it impossible to aggregate data in di erent ways with
respect to the ones given or to put together INVALSI
data related to di erent school years.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We can conclude that yes, the data of the last three
editions of the Italian Bebras support the hypothesis
that the general INVALSI national assessment of
Italian schools can be used to predict the performance
of students in the Italian edition of the Bebras
International Challenge on Informatics and Computational
Thinking. This result is not completely obvious, since
Bebras avoids tasks based on curricular subjects and
technical jargon and INVALSI assesses competencies
in linguistic and mathematical areas, not directly
addressed by Bebras. In fact, Italian schools do not
have curricular informatics in grades 5 and 8. The
national guidelines for primary and lower secondary
schools somewhat mention computational thinking,
but the adoption in school and its perception by
teachers is rather discontinuous [CLN17a, CLN17b]. Even
in grade 10, informatics appears only in vocational
curricula and science oriented programs. A more
coherent proposal is under discussion (see [FLL+18]),
but currently we can safely assume that informatics
and computational thinking are not routinely faced by
the general population of Italian schools.
Nevertheless, the Bebras snapshot seems to re ect the general
geographic trend of Italian schools, even if the
participants come from schools with a special interest in
computational thinking and informatics. This could
be an important result, because Bebras data can be
used to assess the computational skills of the students
and, according to our study, they have the potential
to be generalized to a wider population.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The authors wish to thank Federico Pedersini and
Massimo Santini for discussing early drafts of this
paper.
Italian
Mathematics
Eng. listening
Eng. reading
Bebras
Italian
Mathematics
Eng. listening
Eng. reading
Bebras</p>
      <p>for grade 5 INVALSI and Bebras results (macro-areas)</p>
      <p>for grade 8 INVALSI and Bebras results (macro-areas)</p>
      <p>Valentina Dagiene_. Sustaining informatics
education by contests. In Proceedings of
ISSEP 2010, volume 5941 of Lecture Notes
in Computer Science, pages 1{12, Zurich,
Switzerland, 2010. Springer.
[DS16]
[FLL+18]
[GH06]
[HCD11]
[HM14]</p>
      <p>
        Marta Desimoni. Prove
        <xref ref-type="bibr" rid="ref4">INVALSI
2018</xref>
        , chapter Le prove carta e matita
per la rilevazione nazionale degli
apprendimenti
        <xref ref-type="bibr" rid="ref4">INVALSI 2018</xref>
        :
aspetti metodologici.
        <xref ref-type="bibr" rid="ref4">INVALSI, 2018</xref>
        .
https://invalsi-areaprove.cineca.
it/docs/2019/Parte_II_capitolo_2_
aspetti_metodologici_P&amp;P_2018.pdf.
      </p>
      <p>Valentina Dagiene_ and Sue Sentance. It's
computational thinking! bebras tasks in
the curriculum. In Proceedings of ISSEP
2016, volume 9973 of Lecture Notes in
Computer Science, pages 28{39, Cham,
2016. Springer.</p>
      <p>Luca Forlizzi, Michael Lodi, Violetta
Lonati, Claudio Mirolo, Mattia Monga,
Alberto Montresor, Anna Morpurgo, and
Enrico Nardelli. A core informatics
curriculum for Italian compulsory schools. In
Pozdniakov S. and Dagiene_ V., editors,
Informatics in schools. fundamentals of
computer science and software
engineering. ISSEP 2018., volume 11169 of LNCS,
pages 141{153. Springer, Cham, 2018.</p>
      <p>Andrew Gelman and Jennifer Hill. Data
analysis using regression and
multilevel/hierarchical models. Cambridge
university press, Cambridge, UK, 2006.</p>
      <p>Bruria Haberman, Avi Cohen, and
Valentina Dagiene_. The beaver contest:
Attracting youngsters to study computing.</p>
      <p>In Proceedings of ITiCSE 2011, pages 378{
378, Darmstadt, Germany, 2011. ACM.</p>
      <p>Peter Hubwieser and Andreas Muhling.</p>
      <p>Playing PISA with Bebras. In Proceedings
[HS85]
[INV18]
[KVC06]</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [BCL+18]
          <string-name>
            <surname>Carlo</surname>
            <given-names>Bellettini</given-names>
          </string-name>
          , Fabrizio Carimati, Violetta Lonati, Riccardo Macoratti, Dario Malchiodi, Mattia Monga, and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Morpurgo</surname>
          </string-name>
          .
          <article-title>A platform for the Italian Bebras</article-title>
          .
          <source>In Proceedings of the 10th international conference on computer supported education (CSEDU</source>
          <year>2018</year>
          ) | Volume
          <volume>1</volume>
          , pages
          <fpage>350</fpage>
          {
          <fpage>357</fpage>
          .
          <string-name>
            <surname>SCITEPRESS</surname>
          </string-name>
          ,
          <year>2018</year>
          .
          <article-title>Best poster award winner</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [BLM+15]
          <string-name>
            <surname>Carlo</surname>
            <given-names>Bellettini</given-names>
          </string-name>
          , Violetta Lonati, Dario Malchiodi, Mattia Monga, Anna Morpurgo, and
          <string-name>
            <given-names>Mauro</given-names>
            <surname>Torelli</surname>
          </string-name>
          .
          <article-title>How challenging are Bebras tasks? an IRT analysis based on the performance of Italian students</article-title>
          .
          <source>In Proceedings of ITiCSE 2015</source>
          , pages
          <fpage>27</fpage>
          {
          <fpage>32</fpage>
          , Vilnius, Lithuania,
          <year>July 2015</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [CAC+18]
          <string-name>
            <surname>Giuseppe</surname>
            <given-names>Chiazzese</given-names>
          </string-name>
          , Marco Arrigo, Antonella Chifari, Violetta Lonati, and
          <string-name>
            <given-names>Crispino</given-names>
            <surname>Tosto</surname>
          </string-name>
          .
          <article-title>Exploring the e ect of a robotics laboratory on computational Ronald K. Hambleton</article-title>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Swaminathan</surname>
          </string-name>
          .
          <source>Item Response Theory: Principles and Applications</source>
          . Springer-Verlag, Berlin,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          INVALSI.
          <article-title>Rapporto prove INVALSI 2018</article-title>
          .
          <article-title>Technical report</article-title>
          , INVALSI,
          <year>2018</year>
          . Only in Italian, available at https://www.invalsi.it/invalsi/ doc_evidenza/2018/Rapporto_prove_ INVALSI_
          <year>2018</year>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Graeme</given-names>
            <surname>Kemkes</surname>
          </string-name>
          , Troy Vasiga, and
          <string-name>
            <surname>Gordon</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Cormack</surname>
          </string-name>
          .
          <article-title>Objective scoring for computing competition tasks</article-title>
          .
          <source>In Proceedings of 2nd ISSEP</source>
          , volume
          <volume>4226</volume>
          of Lecture Notes in Computer Science, pages
          <volume>230</volume>
          {
          <fpage>241</fpage>
          , Berlin, Germany,
          <year>2006</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [LMM+17]
          <string-name>
            <surname>Violetta</surname>
            <given-names>Lonati</given-names>
          </string-name>
          , Mattia Monga, Anna Morpurgo, Dario Malchiodi, and
          <string-name>
            <given-names>Annalisa</given-names>
            <surname>Calcagni</surname>
          </string-name>
          .
          <article-title>Promoting computational thinking skills: would you use this Bebras task? In Proceedings of the international conference on informatics in schools: situation, evolution and perspectives (ISSEP2017</article-title>
          ), Lecture Notes in Computer Science, Cham,
          <string-name>
            <surname>CH</surname>
          </string-name>
          ,
          <year>2017</year>
          . Springer International Publishing AG. To appear.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [SBS17] [Sta16]
          <string-name>
            <given-names>Suzanne</given-names>
            <surname>Straw</surname>
          </string-name>
          , Susie Bamford, and
          <string-name>
            <given-names>Ben</given-names>
            <surname>Styles</surname>
          </string-name>
          .
          <article-title>Randomised controlled trial and process evaluation of code clubs</article-title>
          .
          <source>Technical Report CODE01</source>
          , National Foundation for Educational Research, May
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>Available at: https://www.nfer.ac.uk/ publications/CODE01.</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Stan</given-names>
            <surname>Development Team</surname>
          </string-name>
          .
          <article-title>Stan modeling language users guide and reference manual version 2</article-title>
          .19.0. http://mc-stan.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>