<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building structured synthetic datasets: The case of Blackbird Language Matrices (BLMs)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Paola Merlo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Samo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vivi Nastase</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chunyang Jiang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Linguistics, University of Geneva</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Our goal is to investigate, ultimately to enhance, to what degree existing LLM learn disentangled rule-based, compositional linguistic representations. We take the approach of developing curated synthetic data on a large scale, with specific properties, and using them to study sentence representations built using pretrained language models. Inspired by IQ tests, we develop a new multiple-choice task. Finding a solution to this task requires a system detecting complex linguistic patterns and paradigms in text representations. We present formal specifications of this task, illustrate it with two problems and present their benchmarking results. Il nostro obiettivo è indagare, allo scopo di migliorare, quanto gli LLM esistenti apprendano rappresentazioni linguistiche composte, basate su regole districate. Il nostro approccio consiste nello sviluppare dati sintetici curati su larga scala, con proprietà specifiche, e nell'utilizzarli per studiare le rappresentazioni di frasi costruite con modelli linguistici pre-addestrati. Ispirandoci ai test del QI, abbiamo sviluppato un nuovo task a scelta multipla. Trovare la soluzione di questo task richiede che il sistema individui schemi e paradigmi linguistici complessi nelle rappresentazioni testuali. Presentiamo le specifiche formali di questo task, lo illustriamo con due problemi e presentiamo i risultati del benchmarking.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;synthetic structured data</kwd>
        <kwd>formal definitions of grammatical phenomena</kwd>
        <kwd>diagnostic studies of deep learning models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1. Introduction
spread, fill, stuf and load, that describe covering surfaces
or filling volumes [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. They occur in two
subcategorisation frames, related to each other in a regular way: the
object of the preposition with is the subject of the onto
frame, while the object of the onto prepositional phrase
is the subject of the with frame.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Current consensus about LLM, and NNs in general, is</title>
      <p>that to reach better, possibly human-like, abilities, we
need to develop tasks and data that help us understand
their current generalisation abilities and help us train
or tune them towards more complex and compositional
skills. (1) John loaded the truck with hay.</p>
      <p>
        Humans are good generalizers. A large body of lit- Agent Locative Theme
erature has demonstrated that the human mind is pre- John loaded hay onto the truck.
disposed to generate rules from data and combine these Agent Theme Locative
rules, in ways that have been argued to be distinct from
the patterns of activation of neural networks [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. To learn the structure of such a complex alternation
One possible approach to develop more robust methods, automatically, a neural network must be able to identify
then, is to drive the network to learn disentangled decom- the elements manipulated by the alternation, and their
positions of complex observations and learn underlying relevant attributes, and recognize the operations that
regularities [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. manipulate these objects, across more than one sentence.
      </p>
      <p>
        Let’s look at an illustrative example of what complex To study what factors lead to learning more
disendecomposition of covert rules would be necessary. Con- tangled linguistic representations —representations that
sider complex argument structure relations in the lexicon: reflect the underlying linguistic rules of grammar— we
for example, the Spray/load alternation in English, shown take the approach of developing curated synthetic data on
in (1). a large scale, building diagnostic models from pretrained
This alternation applies to verbs such as spray, paint, representations of these data and investigating the
models’ behaviour. To this end, we develop a new linguistic
CLiC-it 2023: 9th Italian Conference on Computational Linguistics, task, inspired by the IQ test RPM (Raven 1938), which we
Nov 30 — Dec 02, 2023, Venice, Italy call Blackbird Language Matrices (BLMs). BLMs define a
* Corresponding author. prediction task to learn complex linguistic patterns and
$ Paola.Merlo@unige.ch (P. Merlo); Giuseppe.Samo@unige.ch paradigms [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
(cGh.unSaymanog);.jviainvig.a4.2n@asgtmasaei@l.cgommai(lC.c.oJmian(Vg.) Nastase); In this paper, we present precise formal specifications
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License of the BLM task, illustrate it with the instantiations of
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org) two BLM problems and their benchmarking results. This
shows that the general formalism can be used to generate
datasets with the same format, and similar specification.
      </p>
      <p>
        Expanding the covered phenomena can thus be done
systematically, allowing for studies that combine or work
across multiple phenomena and languages [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. We
believe this task takes us closer to investigations of human
linguistic intelligence.
      </p>
      <sec id="sec-2-1">
        <title>2. RPMs and BLMs</title>
        <p>
          Raven’s progressive matrices are IQ tests consisting of
a sequence of images, called the context, connected in a
logical sequence by underlying generative rules [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The
task is to determine the missing element in this visual
sequence, the answer. An instance is shown in Figure 1:
given a matrix (left), choose the last element of the matrix
from given options. The matrices are built according to
generative rules that span the whole sequence of stimuli
and the answers are constructed to be similar enough
that the solution can be found only if the rules are
identified correctly. For example in Figure 1, the matrix is
constructed according to two rules: Rule 1: row-wise,
from left to right, the red dot moves one place clockwise
each time. Rule 2: column-wise, from top to bottom, the
blue square moves one place anticlockwise each time.
Identifying these rules leads to the correct answer, the
only cell that continues the generative rules correctly.
        </p>
        <p>
          A similar task has been developed called Blackbird
Language Matrices (BLMs) [
          <xref ref-type="bibr" rid="ref11 ref7 ref8">7, 8, 11</xref>
          ] for linguistic problems,
as given in Figure 2, which illustrates the template of a
BLM agreement matrix. As can be seen, the agreement
rules are implicitly expressed by patterns in the sentence,
with or without intervening attractor elements, and
alternate in a combinatorial pattern across the sentences
(shown in colour), so that only one answer concludes the
sequence.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>3. Formal Specifications of BLMs</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>We define here the new Blackbird’s Language Matrices (BLMs) task and data format.</title>
      <p>Context
1 NP-sing PP1-sing VP-sing
2 NP-plur PP1-sing VP-plur
3 NP-sing PP1-plur VP-sing
4 NP-plur PP1-plur VP-plur
5 NP-sing PP1-sing PP2-sing VP-sing
6 NP-plur PP1-sing PP2-sing VP-plur
7 NP-sing PP1-plur PP2-sing VP-sing
8 ???</p>
      <p>Answer
1 NP-sing PP1-sing et NP2 VP-sing Coord
2 NP-plur PP1-plur PP2-sing VP-plur correct
3 NP-sing PP-sing VP-sing WNA
4 NP-sing PP1-sing PP2-sing VP-plur AE
5 NP-plur PP1-sing PP1-sing VP-plur WN1
6 NP-plur PP1-plur PP2-plur VP-plur WN2</p>
      <sec id="sec-3-1">
        <title>3.1. Defining the linguistic phenomenon</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>The first step in the definition of the problem consists in formally defining the linguistic grammatical phenomenon as a paradigm.</title>
    </sec>
    <sec id="sec-5">
      <title>Definition Let a linguistic phenomenon LP be given.</title>
      <p>LP is exhaustively defined by a grammar  =
(, , , , ) s.t.
 is the set of objects
 is the set of attributes of the objects in 
 is the set of external observed rules
 is the set of unobserved internal rules
 is the lexicon of objects in , attributes in ,
and operators in  ∪ .</p>
      <p>For example, as shown in the example Figure 3, in
the subject-verb agreement phenomenon, the agreement
rule is the primary production in , while the fact that
agreement can occur independently of the distance of
the elements expresses the fact that agreement applies
to structural representations, a rule in . Sometimes, but
not always, I acts as a confusing factor.</p>
      <p>Rules are triples of objects (shown in red), attributes
(in green) and operations (in blue). Objects are usually
phrases, attributes are usually morpho-syntactic
properties of the phrases and operations are typical grammatical
operations: feature match, movement (becomes), lexical
substitution (changes).</p>
      <sec id="sec-5-1">
        <title>3.2. Defining the matrices</title>
        <p>Definition A  matrix is a tuple=(, ,  ) s.t.
 is the shape of the matrix
 are the relational operators that connect the
items of the matrix
 is the set of items of the matrix.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Shape (, ) is the shape of the matrix, which consists</title>
      <p>of  items and each item can be at most of length
.</p>
    </sec>
    <sec id="sec-7">
      <title>The length of the items can vary. The items</title>
      <p>can be sentences or elements in a morphological
paradigm. The choice of  depends on how many
items need to be shown to illustrate the paradigm
and on whether the illustration is exhaustive or
sampled.</p>
    </sec>
    <sec id="sec-8">
      <title>For example, a matrix of size eight is exhaus</title>
      <p>tive for an agreement problem with three noun
phrases and a two-way number diferentiation
(singular, plural), but can only present a sample
of the information for the spray/load alternation.
Subject-verb number agreement
Violation of E: wrong subject-verb agreement
Violation of I: wrong agreement on N2 or N3
Violation of R: wrong number of attractors
Spray/load alternation
Violation of E: Wrong lexical choice of preposition
Violation of I: Subject of active voice is not Agent
Violation of R: Wrong number of arguments</p>
    </sec>
    <sec id="sec-9">
      <title>The items  are defined by</title>
      <p>= (, , , , ) and they are drawn
from the set  .</p>
    </sec>
    <sec id="sec-10">
      <title>The matrix is created by sampling (, , ) triples. The</title>
      <p>ways in which  ∈  can apply to a given (, ) pair has
to be predefined, as it is not entirely context-free.</p>
      <sec id="sec-10-1">
        <title>3.3. Defining the answer set</title>
        <p>The answer set  consists of a set of items like those
in . One item in W, , is the correct answer to
complete the sequence defined by . The other items are
the contrastive set. They are items that violate  , the
rules of construction of the context matrix C, either in
the primary rules , in the auxiliary rules , or in the
matrix operators .</p>
        <p>Sometimes they are built almost automatically,
sometimes by hand. The cardinality of the answer set is
determined by how many facets of the linguistic phenomenon
need to be shown to have been learned.</p>
      </sec>
      <sec id="sec-10-2">
        <title>3.4. Augmenting the matrices</title>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>Diferent levels of lexical and structural complexity can be obtained by changing the lexical items (completely or partially), in a given matrix.</title>
      <p>Definition An augmented BLM is a quadruple
(, ,  , ).
 is the shape of the matrix,  are the relational
operations that connect the  items of the matrix.
 is a set of operations defined to augment the
cardinality of  , while keeping S and R constant. 
is defined by controlled manipulations of Os and
As in  to collect similar elements.
Experimental</p>
      <p>Datasets</p>
      <p>Grammar repositories
Devised
exnovo structures</p>
      <p>Natural y
occurring
examples
Masked augmentation</p>
      <p>Lexical set seed</p>
      <p>Contexts and answers sets</p>
      <sec id="sec-11-1">
        <title>4.1. BLM-AgrF – subject-verb agreement in French</title>
        <p>We augment the sentence set  by modifying the noun
phrases of the items in  . We generate alternatives with
a language model choosing among the top , within an
acceptability margin from the original sentence. The
margin is set with a variable-size window and collects
the top 10 alternative noun phrases. The acceptability of
the resulting sentences is validated manually. In the next
sections, we illustrate the data for two BLM problems,
and baseline benchmarking results.</p>
        <p>
          In BLM-AgrF [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], a BLM problem for subject-verb
agreement consists of a context set of seven sentences that
share the subject-verb agreement phenomenon, but
differ in other aspects – e.g. number of intervening noun
phrases between the subject and the verb, called
attractors because they can interfere with the agreement,
different grammatical numbers for these attractors, and
4. Example of two BLM problems diferent clause structures. Each context is paired with a
set of candidate answers. The answer sets contain
miniThe creation of structured datasets can be a challenging mally contrastive examples built by corrupting some of
task, depending on the type of linguistic problem being the generating rules. This helps investigate the kind of
investigated, the available linguistic resources, and the information and structure learned, by error analysis. An
size of the lexical factors involved in the problem. Figure example template is illustrated in Figure 2, and an actual
5 summarizes the pipeline that needs to be followed for example in Figure 6.
the whole process: identifying the data for the linguistic The dataset comprises three subsets, of increasing
lexphenomenon under investigation, developing a lexical ical complexity. Type I data is generated based on
manset seed of lexical items for creating context and answer ually provided seeds (Franck et al. 2002), illustrated in
sets for the BLMS, which are then combined to construct Figure 9, and a template that captures the rules mentioned
desired context templates and answer sets. From the lin- above. Type II data is generated based on Type I data, by
guistic phenomenon to the creation of the lexical set seed, introducing lexical variation with the aid of a transformer,
various approaches can be pursued based on the type of by generating alternatives for masked nouns. Type III
linguistic phenomenon being investigated. This choice data is generated by combining sentences from diferent
might depend on whether the phenomenon has already instances from the Type II data, while maintaining the
been extensively studied in experimental linguistics, the structure of the sequence. The structural variations alter
scale of the lexical components involved in the linguistic the distance and relative depth of the subject and verb
phenomenon, and the available resources in the target and produce a variety of conditions. The diferent levels
language. We then employ a fill-mask task with trans- of lexical variation will allow us to investigate the impact
formers to automatically generate additional, plausible of lexical variation on the ability of a system to detect
constituents for the desired structures. grammatical patterns. We include complete instances –
1 Il vaso
2 I vasi
3 Il vaso
4 I vasi
5 Il vaso
6 I vasi
7 Il vaso
8 ???
        </p>
        <p>Context
con il fiore
con il fiore
con i fiori
con i fiori
con il fiore
con il fiore
con i fiori
del giardino
del giardino
del giardino
si è rotto.
si sono rotti.
si è rotto.
si sono rotti.
si è rotto.
si sono rotti.</p>
        <p>si è rotto.</p>
        <p>Answer set
1 Il vaso con i fiori e il giardino si è rotto. coord
2 I vasi con i fiori del giardino si sono rotti. correct
3 Il vaso con il fiore si è rotto. WNA
4 Il vaso con il fiore del giardino si sono rotti. AE
5 I vasi con i fiori del giardino si sono rotti. WN1
6 I vasi con i fiori dei giardini si sono rotti. WN2</p>
        <p>Answers
1 NP-Agent Verb PP-Loc correct
2 NP-Agent Verb *PP-Loc WPrep
3 NP-Agent Verb *[NP-Theme PP-Loc] PP-Loc WPP
4 NP-Agent Verb NP-Loc because PP-Theme Adv
5 NP-Agent Verb NP-Loc *PPLoc-Theme WT
6 *NP-Loc Verb PP-Loc WS
meAgent, Repeat in Template 2), and others that govern
in French – in Appendix A. the proper lexical selection (Wrong Preposition,
Adverbial).</p>
        <p>Diferent answer sets, that focus on diferent rule
sub4.2. BLM-s/lE.v0 – spray/load verb patterns, will allow for detailed investigations of the type
alternations in English of information that is more easily or more dificult to
deIn the BLM-s/lE problem developed to exhibit the tect, and to determine principles of designing the answer
spray/load alternation (discussed in the introduc- set for the most informative phenomenon investigation.
tion), each sentence can be described in terms of one Like the BLM-AgrF dataset, the BLM-s/lEv.0 dataset
distribution-of-three-values rule, governing the semantic presents lexically varied versions, Type I, II, and III, of
roles (Agent, Theme, Locative), and two distribution- increasing variability. The structure of the context and
anof-two-values rules governing syntactic types (nominal swer set of one alternant is presented in Figure 7. Figure
phrase NP vs. prepositional phrases PP) and the mood of 11 for the other template and relevant lexical examples
the verb, whether active (Verb) or passive (VerbPass). of both templates are given in appendix B.
We created two templates, targeting the syntax-semantic
mapping of the arguments. 5. Benchmarking systems</p>
        <p>In the contrastive answer set, the target sentence is to
be chosen from a set of candidates that exhibit minimal Our goal is to investigate – and ultimately use this
knowldiferences. The semantic-syntactic mapping of the alter- edge to enhance – textual representations built using
nation can be decomposed into a set of smaller patterns pretrained large language models. To determine whether
that describe the sentences in the alternation and that such representations encode linguistic rules, and to what
can be violated to construct incorrect answers. Difer- degree they are compositional, we use BLM tasks that
ent subsets of patterns can be used to develop diefrent provide data generated using specific rules, and baseline
answer sets. systems that should be capable of detecting the patterns</p>
        <p>
          A variation of this dataset, presented in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] uses an that encode the relevant rules for the targeted phenomena
answer set that change the position of the agent in an in the distributed continuous sentence representations.
active sentence, the type of phrases following the verb, We choose a FFNN and a CNN as our baselines. The FFNN
embedding the PP in an NP, and changes in prepositions should be able to discover patterns distributed
throughthat introduce diferent types of arguments. out a sentence, and throughout a sequence of sentences,
        </p>
        <p>
          The dataset presented here, changes subpatterns that while the CNN could discover localized patterns – both
govern the correct learning of the syntactic form of the in the sentence and the sequence.
sentence (Wrong Theme, Wrong Subject, Wrong PP As presented above, a BLM problem instance consists
in Template 1; SwapLocAgent, NoAgent, SwapThe- of a context and an answer set. The context is a sequence
of 7 sentences, and the answer set is a set of 6 sentences,
one of which is a correct continuation of the input se- CNN FFNN
quence. All sentences are encoded using BERT [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] – as All training data
the embedding of the [CLS] token on the last layer of the
model. We used the pretrained
"BERT-base-multilingualcased" model1. The sentence representations are
combined in diferent ways, depending on the baseline system
– a FFNN or a CNN – used.
        </p>
        <p>The input to the FFNN is the concatenation of sentence
embeddings in the BLM instance context, as a vector of
size 7 * 768. This input is processed through 3 fully
connected layers, which progressively compress the input Same training data
size (7 * 76→−8−− − 1 3.5 * 76→−8−− − 2 3.5 * 76→−8−− − 3
768) to obtain the size of a sentence representation. The
FFNN’s interconnected layers enable it to capture
patterns that are distributed throughout the entire input
vector.</p>
        <p>The input to the CNN is the stacked sentence
embeddings in the BLM instance context, as a (7 x 768)
array. This input undergoes three consecutive layers of
2-dimensional convolutions, where each convolutional Figure 8: F1 results (averages over 5 runs) for alternant one.
layer uses a kernel size of (3x3) and a stride of 1, without
dilation. The resulting output from the convolutional
process is then passed through a fully connected layer,
which compresses it to the size of the sentence represen- II data whether testing on type II or III. It appears, then,
tation (768). By using a kernel size of (3x3), stride=1, and that on smaller datasets, the template patterns is perhaps
no dilation, this configuration emphasizes the detection better learnt in type II, while still retaining the notion of
of localized patterns within the sentence sequence array. lexical variation that is lacking in type I.</p>
        <p>The output of both systems is a vector representing a Inspecting the increase in performance with diferent
sentence embedding. This is compared to the sentence training data sizes, shown in Figure 12 in the appendix,
representations in the answer set, and the one with the it is confirmed that learning is very fast and plateaus
highest score is considered the correct answer. Details already with a few thousand examples for all train-test
are included in appendix C. combinations, with the exception of training on Type I
and testing on type III, which is clearly too dificult.</p>
        <p>
          Both baseline systems lead to good results despite the
6. Results variations in the input – structural and lexical – that
superficially obfuscate these phenomena, and the
nearPrevious published work from our group and current miss incorrect answers, confirming that the phenomena
ongoing work has benchmarked the problems generated we target are encoded in the sentence representations.
by these datasets and analysed the errors, with interesting Because of their diferent architectures and the type of
results [
          <xref ref-type="bibr" rid="ref11 ref14">11, 14</xref>
          ]. patterns they discover, the high performance of both
sys
        </p>
        <p>
          We report here the novel results on the BLM-s/lE tems indicates that relevant patterns for our two targeted
dataset, which are qualitatively similar to those reported phenomena are localized in BERT sentence embeddings.
for BLM-AgrF [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], thus confirming some general trends. Further steps can take advantage of the structured way
Figure 8 shows the results. The top panels shows results the data was constructed to attempt to disentangle the
using all the data, the bottom panel shows results using various generative rules and additional factors in the
only data sizes that match type I data size, hence smaller inputs.
data sizes for type II and type III.
        </p>
        <p>
          Globally, the results are very good. But interesting
diferences emerge if we vary data sizes. If we train on 7. Related Work
all data, more lexically-varied data (types II and III) give
better results, but if we train on equally sized datasets,
we see an improvement of the results in training on type
Previous work has focussed on understanding the
automatic learning of verb alternations in terms of syntactic
and semantic properties of the verbs and their argument
structures [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. These properties have been explored in
        </p>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>1https://huggingface.co/bert-base-multilingual-cased</title>
      <p>
        relation to their representation in LLMs, across various
dimensions of performance for diferent models [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ].
In particular, [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] suggest that LLMs with contextual
embeddings encode linguistic information on verb
alternation classes, at both the word and sentence levels. In
their work, [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] build upon [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] observations and
highlight the superior performance of one transformer Electra
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] compared to other large language models.
      </p>
      <p>
        The automatic generation of RPM-like matrices,
whether in vision or in language, is technically
challenging. In computer vision, several formalisms have been
proposed ([
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] formulate RPMs with first-order logic;
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] propose Procedurally Generated Matrices (PGM)
datasets through relation-object-attribute triple
instantiations; [21] use the Attributed Stochastic Image Grammar
(A-SIG [22]). Structured synthetic datasets have been
mostly developed to study issues of generalisation and
disentaglement, in vision [23], with full-fledged
experimentation and for language in a preliminary, nonRPM
-like dataset, consisting of simple examples containing a
few morphological markings [24]. The simplicity of the
sentences does not provide a suficiently realistic
challenge from a linguistic point of view. Very recent work
has started exploring the picture-naming potential of
language to solve problems in vision [25].
      </p>
      <sec id="sec-12-1">
        <title>8. Conclusions</title>
        <p>
          In this paper, we have presented the new BLM task,
provided its formal specifications and illustrated the first
instances of BLM problems and benchmarking results
with baseline architectures. Current work is developing
new dedicated architectures based on Variation
Autoencoders [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and developing new BLM problems. Future
work lies in further automating the data development
pipeline, to make the creation on BLM data sets also
accessible to less computationally-oriented linguists and
investigating the structure and nature of the information
encoded in the learned inner representations.
        </p>
      </sec>
      <sec id="sec-12-2">
        <title>Acknowledgments</title>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>We gratefully acknowledge the partial support of this work by the Swiss National Science Foundation, through grants #51NF40_180888 (NCCR Evolving Language) and SNF Advanced grant TMAG-1_209426 to PM.</title>
      <sec id="sec-13-1">
        <title>A. BLM-AgrF problem</title>
        <p>Example subject NPs from [26]
L’ordinateur avec le programme de l’experience
The computer with the program of the experiments
Manually expanded and completed sentences
L’ordinateur avec le programme de l’experience est en panne.</p>
        <p>The computer with the program of the experiments is down.</p>
        <p>Jean suppose que l’ordinateur avec le programme de l’experience est en panne.</p>
        <p>Jean thinks that the computer with the program of the experiments is down.</p>
        <p>L’ordinateur avec le programme dont Jean se servait est en panne.</p>
        <p>The computer with the program that John was using is down.</p>
        <p>A seed for language matrix generation
Jean suppose que l’ordinateur avec le programme
Jean thinks that the computer with the program
de l’experience est en panne
of the experiment is down
les ordinateurs avec les programmes
the computers with the programs
sont en panne
are down</p>
        <p>Contexts
Example Translation
1 La conférence sur l’histoire a commencé plus tard que prévu. The talk on history has started later than expected.
2 Les responsables du droit vont démissionner. Those responsible for the right will resign.
3 L’ exposition avec les peintures a rencontré un grand succès. The show with the paintings has met with great success.
4 Les menaces de les réformes inquiètent les médecins. The threats of reforms worry the doctors.
5 Le trousseau avec la clé de la cellule repose sur l’étagère. The bunch of keys of the cell sits on the shelf.
6 Les études sur l’efet de la drogue apparaîtront bientôt. The studies on the efect of the drug will appear soon.
7 La menace des réformes dans l’ école inquiète les médecins. The threat of reforms in the school worries the doctors.
Answers
Example Translation
1 Les nappes sur les tables et le banquet brillent au soleil. The tablecloths on the table and the console shine in the sun.
2 Les copines des propriétaires de la villa dormaient sur la The friends of the owners of the villa were sleeping on the beach.
plage.
3 Les avocats des assassins vont revenir. The laywers of the murderers will come back.
4 Les avocats des assassins du village va revenir. The lawyers of the murderers of the village will come back.
5 La visite aux palais de l’ artisanat approchent. The visit of the palace of the crafts is approaching.
6 Les ordinateurs avec le programme des expériences sont en panne. The computers with the program of the experiments are broken.</p>
        <p>B. BLM-s/lE</p>
        <p>Context
1 NP-Agent Verb NP-Theme
2 NP-Agent Verb PP-Theme
3 NP-Agent Verb NP-Loc
4 NP-Agent Verb PP-Loc
5 NP-Agent Verb NP-Theme PP-Loc
6 NP-Theme VerbPass PP-Loc
7 NP-Agent Verb NP-Loc
8 ???
All systems used a learning rate of 0.001 and Adam optimizer, and batch size 100. The training was done for 120
epochs. The experiments were run on an HP PAIR Workstation Z4 G4 MT, with435 an Intel Xeon W-2255 processor,
64G RAM, and a MSI GeForce RTX 3090 VENTUS 3X OC 24G GDDR6X GPU.</p>
        <p>
          We tested BERT sentence embeddings with baseline CNN and FFNN baseline architectures. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The sentence
embeddings are the encoding of the [CLS] token on the last layer of the model.
        </p>
        <p>The FFNN receives the input as a concatenation of sentence embeddings in a sequence, with a size of 7 * 768.
This input is then processed through 3 fully connected layers, which progressively compress the input size (7 * 768
→ 1 3.5 * 768 → 2 3.5 * 768 → 3 768) to obtain the size of a sentence representation. The FFNN’s
interconnected layers enable it to capture patterns that are distributed throughout the entire input vector.</p>
        <p>The CNN takes as input an array of embeddings with a size of (7 x 768). This input undergoes three consecutive
layers of 2-dimensional convolutions, where each convolutional layer uses a kernel size of (3x3) and a stride of 1,
without dilation. The resulting output from the convolutional process is then passed through a fully connected layer,
which compresses it to the size of the sentence representation (768). By using a kernel size of (3x3), stride=1, and no
dilation, this configuration emphasizes the detection of localized patterns within the sentence sequence array.</p>
        <p>Both networks produce the same output, which is a vector representing the sentence embedding of the correct
answer. The objective of learning is to maximize the probability of selecting the correct answer from a set of
candidate answers. To achieve this, we employ the max-margin loss function, considering that the incorrect answers
in the answer set are intentionally designed to have minimal diferences from the correct answer. This loss function
combines the distances between the predicted answer and both the correct and incorrect answers. Initially, we
calculate a score for each candidate answer’s embedding  in the answer set  with respect to the predicted sentence
embedding . This score is determined by the cosine of the angle between the respective vectors:
(, ) = (, )</p>
        <p>The loss function incorporates the max-margin concept, taking into account the diference between the score of
the correct answer  and each of the incorrect answers :
 = ∑︁[1 − (, ) + (, )]+</p>
        <p />
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lakretz</surname>
          </string-name>
          , G. Kruszewski,
          <string-name>
            <given-names>T.</given-names>
            <surname>Desbordes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hupkes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dehaene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <article-title>The emergence of number and syntax units in LSTM language models</article-title>
          , arXiv preprint arXiv:
          <year>1903</year>
          .
          <volume>07435</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lakretz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hupkes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vergallito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dehaene</surname>
          </string-name>
          ,
          <article-title>Mechanisms for handling nested dependencies in neural-network language models and humans</article-title>
          .,
          <string-name>
            <surname>Cognition</surname>
          </string-name>
          (
          <year>2021</year>
          ). doi:
          <year>2021</year>
          10.1016/j.cognition.
          <year>2021</year>
          .
          <volume>104699</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sablé-Meyer</surname>
          </string-name>
          , J. Fagot,
          <string-name>
            <given-names>S.</given-names>
            <surname>Caparos</surname>
          </string-name>
          , T. van Kerkoerle,
          <string-name>
            <given-names>M.</given-names>
            <surname>Amalric</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dehaene</surname>
          </string-name>
          ,
          <article-title>Sensitivity to geometric shape regularity in humans and baboons: A putative signature of human singularity</article-title>
          ,
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>118</volume>
          (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1073/pnas.2023123118.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vincent</surname>
          </string-name>
          ,
          <article-title>Representation learning: A review and new perspectives</article-title>
          ,
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>35</volume>
          (
          <year>2013</year>
          )
          <fpage>1798</fpage>
          -
          <lpage>1828</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Levin</surname>
          </string-name>
          ,
          <article-title>English verb classes and alternations: A preliminary investigation</article-title>
          , University of Chicago Press,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Beavers</surname>
          </string-name>
          , The spray/load alternation, The Wiley Blackwell Companion to Syntax, Second
          <string-name>
            <surname>Edition</surname>
          </string-name>
          (
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Merlo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Rodriguez</surname>
          </string-name>
          ,
          <article-title>Blackbird's language matrices (BLMs): a new benchmark to investigate disentangled generalisation in neural networks</article-title>
          ,
          <source>ArXiv: cs.CL 2205.10866</source>
          (
          <year>2022</year>
          ). URL: https://arxiv.org/abs/2205.10866. doi:
          <volume>10</volume>
          .48550/A RXIV.
          <volume>2205</volume>
          .10866.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Merlo</surname>
          </string-name>
          ,
          <article-title>Blackbird language matrices (BLM), a new task for rule-like generalization in neural networks: Motivations and formal specifications</article-title>
          ,
          <source>ArXiv cs.CL 2306.11444</source>
          (
          <year>2023</year>
          ). URL: https://doi.org/10.48550/a rXiv.
          <volume>2306</volume>
          .11444. doi:
          <volume>10</volume>
          .48550/arXiv.2306.11 444.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Merlo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Samo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nastase</surname>
          </string-name>
          ,
          <article-title>Blackbird Language Matrices Tasks for Generalization, in: GenBench: The first workshop on (benchmarking) generalisation in NLP</article-title>
          , Singapore,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Raven</surname>
          </string-name>
          , Standardization of progressive matrices,
          <source>British Journal of Medical Psychology</source>
          <volume>19</volume>
          (
          <year>1938</year>
          )
          <fpage>137</fpage>
          -
          <lpage>150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Rodriguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nastase</surname>
          </string-name>
          , P. Merlo,
          <string-name>
            <surname>BLM-AgrF</surname>
          </string-name>
          :
          <article-title>A new French benchmark to investigate generalization of agreement in neural networks</article-title>
          ,
          <source>in: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics</source>
          , Association for Computational Linguistics, Dubrovnik, Croatia,
          <year>2023</year>
          , pp.
          <fpage>1363</fpage>
          -
          <lpage>1374</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .eacl -main.99.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Samo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nastase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jiang</surname>
          </string-name>
          , P. Merlo,
          <article-title>BLM-s/lE: A structured dataset of English spray-load verb alternations for testing generalization in LLMs</article-title>
          ,
          <source>in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing</source>
          , Singapore,
          <year>2023</year>
          . [21]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , S. Zhu, RAVEN: A
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: dataset for relational and analogical visual reasonPre-training of deep bidirectional transformers for ing</article-title>
          ,
          <source>in: IEEE Conference on Computer Vision</source>
          and language understanding,
          <source>in: Proceedings of the Pattern Recognition, CVPR</source>
          <year>2019</year>
          , Long Beach, CA, 2019 Conference of the North American Chapter of USA, June 16-20,
          <year>2019</year>
          , Computer Vision Foundathe Association for Computational Linguistics: Hu- tion / IEEE,
          <year>2019</year>
          , pp.
          <fpage>5317</fpage>
          -
          <lpage>5327</lpage>
          . URL: http:/
          <source>/openac man Language Technologies</source>
          , Volume
          <volume>1</volume>
          (
          <article-title>Long and cess</article-title>
          .thecvf.com/content_CVPR_2019/html/Zhang Short Papers),
          <article-title>Association for Computational Lin-</article-title>
          _
          <string-name>
            <surname>RAVEN</surname>
          </string-name>
          _
          <article-title>A_Dataset_f or_Relational_and_Analog guistics</article-title>
          , Minneapolis, Minnesota,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>ical</lpage>
          _Visual_REasoNing_CVPR_
          <year>2019</year>
          <article-title>_paper</article-title>
          .
          <source>html. 4186</source>
          . URL: https://aclanthology.org/N19- 1423. doi:
          <volume>10</volume>
          .1109/CVPR.
          <year>2019</year>
          .
          <volume>00546</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          -1423. [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mumford</surname>
          </string-name>
          ,
          <article-title>A stochastic grammar of im-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>V.</given-names>
            <surname>Nastase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Merlo</surname>
          </string-name>
          , Grammatical information in ages,
          <source>Found. Trends Comput. Graph. Vis</source>
          .
          <volume>2</volume>
          (
          <year>2006</year>
          )
          <article-title>BERT sentence embeddings as two-dimensional 259-362</article-title>
          . URL: https://doi.org/10.1561/0600000018. arrays,
          <source>in: Proceedings of the 8th Workshop on doi:10.1561/0600000018. Representation Learning for NLP (RepL4NLP</source>
          <year>2023</year>
          ), [23]
          <string-name>
            <surname>S. van Steenkiste</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Locatello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          , Toronto, Canada,
          <year>2023</year>
          .
          <string-name>
            <given-names>O.</given-names>
            <surname>Bachem</surname>
          </string-name>
          , Are disentangled representations help-
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>O.</given-names>
            <surname>Majewska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Korhonen</surname>
          </string-name>
          ,
          <article-title>Verb classification ful for abstract visual reasoning?</article-title>
          ,
          <source>in: NeurIPS</source>
          <year>2019</year>
          ,
          <article-title>across languages</article-title>
          ,
          <source>Annual Review of Linguistics</source>
          <year>2020</year>
          .
          <volume>9</volume>
          (
          <year>2023</year>
          )
          <fpage>313</fpage>
          -
          <lpage>333</lpage>
          . doi:
          <volume>10</volume>
          .1146/annurev-lingu [24]
          <string-name>
            <surname>A. M'Charrak</surname>
          </string-name>
          ,
          <source>Deep Learning for Natural Language istics-030521-043632. Processing (NLP) using Variational Autoencoders</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Kann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Warstadt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Bowman</surname>
          </string-name>
          , (VAE),
          <source>Master's thesis</source>
          ,
          <source>ETH Switzerland</source>
          ,
          <year>2018</year>
          . URL:
          <article-title>Verb argument structure alternations in word</article-title>
          and https://pub.tik.ee.ethz.ch/students/2018-FS/MA-2
          <article-title>sentence embeddings</article-title>
          ,
          <source>in: Proceedings of the So- 018-22.pdf . ciety for Computation in Linguistics (SCiL)</source>
          <year>2019</year>
          , [25]
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Storks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chai</surname>
          </string-name>
          , In-context ana2019, pp.
          <fpage>287</fpage>
          -
          <lpage>297</lpage>
          . URL: https://aclanthology.org
          <article-title>logical reasoning with pre-trained language models</article-title>
          , /W19-0129. doi:
          <volume>10</volume>
          .7275/q5js-
          <fpage>4y86</fpage>
          .
          <source>in: Proceedings of the 61st Annual Meeting of the</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bruno</surname>
          </string-name>
          , J. Han,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zukerman</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Steinert- Association for Computational Linguistics (Volume Threlkeld, Probing for understanding of English 1: Long Papers), Association for Computational verb classes and alternations in large pre-trained Linguistics</article-title>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>1953</fpage>
          -
          <lpage>1969</lpage>
          .
          <article-title>language models</article-title>
          ,
          <source>in: Proceedings of the Fifth</source>
          Black- URL: https://aclanthology.org/
          <year>2023</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>109</volume>
          . boxNLP Workshop on Analyzing and Interpreting [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Franck</surname>
          </string-name>
          , G. Vigliocco,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nicol</surname>
          </string-name>
          ,
          <article-title>Subject-verb agreeNeural Networks for NLP, Association for Com- ment errors in french and english: The role of synputational Linguistics, Abu Dhabi, United Arab tactic hierarchy, Language and cognitive processes Emirates (Hybrid</article-title>
          ),
          <year>2022</year>
          , pp.
          <fpage>142</fpage>
          -
          <lpage>152</lpage>
          . URL: https:
          <volume>17</volume>
          (
          <year>2002</year>
          )
          <fpage>371</fpage>
          -
          <lpage>404</lpage>
          . //aclanthology.org/
          <year>2022</year>
          .blackboxnlp-
          <volume>1</volume>
          .
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>K.</given-names>
            <surname>Clark</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>T.</given-names>
            <surname>Luong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>Electra: Pre- training text encoders as discriminators rather than generators</article-title>
          ,
          <source>in: ICLR</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <article-title>Automatic generation of raven's progressive matrices</article-title>
          , in: Q.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          <string-name>
            <surname>Wooldridge</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence, IJCAI</source>
          <year>2015</year>
          ,
          <string-name>
            <given-names>Buenos</given-names>
            <surname>Aires</surname>
          </string-name>
          , Argentina,
          <source>July 25-31</source>
          ,
          <year>2015</year>
          , AAAI Press,
          <year>2015</year>
          , pp.
          <fpage>903</fpage>
          -
          <lpage>909</lpage>
          . URL: http: //ijcai.org/Abstract/15/132.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>D.</given-names>
            <surname>Barrett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Santoro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Morcos</surname>
          </string-name>
          , T. Lillicrap,
          <article-title>Measuring abstract reasoning in neural networks</article-title>
          , in: J. G. Dy,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Krause (Eds.),
          <source>Proceedings of the 35th International Conference on Machine Learning</source>
          ,
          <string-name>
            <surname>ICML</surname>
          </string-name>
          <year>2018</year>
          , Stockholmsmässan, Stockholm, Sweden,
          <source>July 10-15</source>
          ,
          <year>2018</year>
          , volume
          <volume>80</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>4477</fpage>
          -
          <lpage>4486</lpage>
          . URL: http://proceedings.mlr.press/ v80/santoro18a.html.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>