<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Transparent Visual Reasoning via Object-Centric Agent Collaboration</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benjamin Teoh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ben Glocker</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca Toni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Avinash Kori</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Imperial College London</institution>
          ,
          <addr-line>United Kingdon</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>A central challenge in explainable AI, particularly in the visual domain, is producing explanations grounded in human-understandable concepts. To tackle this, we introduce OCEAN (Object-Centric Explananda via Agent Negotiation), a novel, inherently interpretable framework built on object-centric representations and a transparent multi-agent reasoning process. The game-theoretic reasoning process drives agents to agree on coherent and discriminative evidence, resulting in a faithful and interpretable decision-making process. We train OCEAN end-to-end and benchmark it against standard visual classifiers and popular post-hoc explanation tools like Grad-CAM and LIME across two diagnostic multi-object datasets. Our results demonstrate competitive performance with respect to state-of-the-art black-box models with a faithful reasoning process, which was reflected by our user study, where participants consistently rated OCEAN's explanations as more intuitive and trustworthy.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Object-Centric Learning</kwd>
        <kwd>Slot Attention</kwd>
        <kwd>Visual Debates</kwd>
        <kwd>Agent Collaboration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Recent developments in visual models have led to highly
accurate image classifiers, but their “black-box” nature present
challenges in trust, accountability, and human-aligned
reasoning. At the crux of explainable AI (XAI) lies the
fundamental goal of clearing the opacity of machine
learning models, making them more intelligible to human users.
This need is particularly acute in high-stakes applications,
such as medical imaging [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and self-driving cars [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Explanations often arrive from post-hoc methods like
GradCAM [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and LIME [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which utilises visual elements like
heatmaps to highlight regions most influential to a model’s
prediction, ofering an intuitive way to visualise the decision.
However post-hoc explanations are disconnected form the
model’s reasoning mechanism, often lacking faithfulness to
the decision-making process.
      </p>
      <p>
        To navigate this hurdle, we draw inspiration from recent
work on ante-hoc explainability, where the model
architecture is carefully designed to expose internal reasoning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Notably, the Consensus Game [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] demonstrates how
multiagent interactions can be used to align generation with
discrimination in language models. In parallel, Visual
Debates [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and Free Argumentative Exchanges [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] explores
extracting transparent reasoning process as a means for
generating for post-hoc explanations. While these
methods difer in modalities and serve diferent purposes, both
emphasise learning and reasoning through reinforcement
learning, feature extraction, and dialogue-based
justification.
      </p>
      <p>In this work, we present OCEAN, an inherently
transparent and human understandable framework for visual
classification. OCEAN casts image classification as a
neurosymbolic collaborative game between agents who must
iteratively select visual symbols towards a shared classification
in the Consensus Game.</p>
      <p>
        OCEAN is informed by recent trends in neuro-symbolic
reasoning [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], concept learning models [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and
disentangled representation learning [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In Figure 1, we
demonstrate an intuitive model of our explainable classification
framework, which distils an explanation in the same
process it makes its prediction. Explanandum is a term we
use to describe the prediction-explanation pair, where the
prediction is solely determined based on the information
and reasoning of the explanation. Here, the explanandum
showed that the classification mechanism considered only
three objects when making the final prediction. Our key
contributions are threefold as detailed below:
• Collaborative Multi-Agent Classification Game :
We designed agent interactions, where agents
present arguments and iteratively converge to a
shared prediction. Through this, we investigate
whether collaborative agent-based reasoning can
achieve transparent classification and reasonable
prediction performance against baselines. Our
results suggest that this formulation not only exposes
the reasoning chain very well but also serves as a
strong foundation for future research in multi-agent
explainability.
• An End-to-End Learning Framework: We
developed a unified framework, OCEAN, that integrates
Slot Attention [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] as a means to extract
objectcentric representations, with our novel Consensus
Game module to perform object decomposition,
explanation generation and classification jointly, all of
which were trained end-to-end.
• Evaluation Against State-of-the-Art Methods:
We assess our method on synthetic diagnostic
datasets and compare it against common
classification methods such as ResNet [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In terms of
explainability, we compare our generated
explanations against the state-of-the-art mechanisms such
as Grad-CAM [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and LIME [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], along with user
studies.
      </p>
      <p>Together, these contributions demonstrate the impact of
integrating object-centric learning with agent-based
collaboration in building more interpretable, faithful, and
humanaligned visual classifiers.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Debate Games Recent research has explored using
structured agent interactions to enhance model explainability
through debate or dialogue, simulating a conversation
between agents. Visual Debates [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and FAX [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is one
approach that leverages debate. Kori et al. model explanations
as a sequential zero-sum debate game between two fictional
players. The player interactions aim to highlight the
classiifer’s reasoning paths, including uncertainties, thereby
offering a more human-aligned explanation structure. While
this approach yielded interpretable explanations, it relied
on surrogate models reasoning over quantised latent
features, limiting its faithfulness and the interpretability of
features presented. In spite of that, the limitations of Visual
Debates strongly inform and influence the design choices
of our OCEAN framework.
      </p>
      <p>
        In parallel, the Consensus Game [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] applied similar
conversational elements in the language domain to align
generative and discriminative reasoning in large language models
(LLM). This is achieved via a game-theoretic framework
which models the interaction between a Generator agent
and a Discriminator agent as an iterative sequential game
with the goal of reaching agreement between the two agents.
The authors provide a compelling foundation for aligning
agent decisions in spite of their diferent roles.
      </p>
      <p>
        Explainability methods Without any need to modify
the underlying architecture of pre-trained models, post-hoc
methods are able to generate explanations by diagnosing
the model’s decision. These methods are widely used due
to their flexibility, as they allow explainability to be applied
to highly performant models without any alterations.
Examples include Grad-CAM [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which highlights regions
in an image that contribute most to a model’s prediction,
and LIME [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which approximates feature importance by
perturbing input data. They often use techniques, such as
feature importance attribution, saliency maps and heatmaps,
to analyse the model’s decision-making process. This makes
post-hoc methods particularly flexible, as they can be
applied to a wide variety of models while preserving
performance. However, post-hoc methods often lack solutions for
producing faithful explanations that accurately represent
the model’s internal reasoning, which may lead to
misleading interpretations.
      </p>
      <p>
        Conversely, ante-hoc methods aim to embed
explainability directly into a model’s decision-making process, making
the model inherently interpretable. These methods seek
to avoid the pitfalls of post-hoc explanations by producing
faithful explanations during the model’s inference time and
aligning with the model’s reasoning. Thus, the
explanations generated from the model stem from fully transparent
predictions. This makes these models very suitable for
highstakes applications where trust, accountability, and
transparency are essential. Designing such models requires a
rethink of the structure of making predictions, often achieved
by introducing architectural constraints, interpretable
feature transforms, or modularising parts of the model.
Examples include ProtoPNet [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], and Self-Explaining Neural
Networks [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In spite of the benefits, the performance
tradeof is significant, and building such architectures
requires more design efort and domain knowledge.
Object-centric learning Object-centric learning (OCL)
is a paradigm in machine learning that focuses on
decomposing a scene into discrete objects, each represented as an
independent entity [
        <xref ref-type="bibr" rid="ref12 ref16 ref17 ref18 ref19">12, 16, 17, 18, 19</xref>
        ]. This method aligns
with human perception, as understanding individual objects
and their relationships is essential for interpreting complex
scenes. By isolating objects, OCL enhances the
interpretability and adaptability of object discovery and set prediction
tasks. For example, in autonomous vehicles, separating
pedestrians, vehicles, and trafic lights provides actionable
insights, while in explainable AI, it highlights specific
objects of a scene that influence the model’s decision. OCL is
introduced to the project to justify the need for structured,
object-centric scene decomposition and understanding. The
object vectors can serve as input for a wide range of
downstream tasks, including classification and generative
modelling. In tandem with the downstream tasks, training is
typically done end-to-end in a unified pipeline. This allows
the learning process to optimise both the object encoding
and the task-specific model simultaneously, ensuring that
the object representations are not only good representations
of objects but are also task-relevant.
      </p>
      <sec id="sec-2-1">
        <title>Neuro-Symbolic learning and Visual reasoning</title>
        <p>
          Neuro-Symbolic Learning is a vast area of research based
on combining neural elements with symbolic reasoning to
improve interpretability and generalisation in deep learning.
In vision, frameworks like NeSy-XIL [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and Concept
Bottleneck Models [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] aim to disentangle high-level visual
concepts into symbols or concepts and utilise these for
decision-making. This exposes the intermediate decision
factors, allowing users to verify or correct reasoning steps.
While we appreciate the general framework allows for
all types of concept learning and symbolic reasoning
mechanisms, they often require supervision on concept
labels. A contrasting but similar approach can be seen in
ProtoPNet [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], which integrates concept-level matching
directly into classification by a matching algorithm of
image regions to prototypical examples learned during
training. This ofers attention-based interpretability
without concept-level supervision. These works highlight
the importance of exposing concept-level reasoning as part
of the model’s prediction.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>Notations We begin by introducing the notation used
throughout the paper. Each input images  ∈ R×  × 
is processed by the Slot Attention module, which outputs
a fixed number  of slot representations. Each slot is a
vector of dimension , resulting in a set of slot encodings
 = {1, . . . ,  } for each image.</p>
      <p>
        While the number of players can be changed, we
present our Consensus Game module as consisting of
two players,  1 and  2, who take turns presenting a
sequence of arguments, 1 = {11, 21, ..., 1} and 2 =
{1, 22, ..., 2}, before ultimately making their
respec2
tive final claims, 1 and 2. Players collaboratively select
the most salient slots as arguments to support their claims,
which are then aggregated to produce the final classification.
Finally, both players are equipped with a shared utility  ,
representing the efectiveness of their choices of arguments.
Similar to [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], we design our framework Γ as:
Γ =
      </p>
      <p>⟨, { 1,  2}, {1, 2}, {1, 2},  ⟩.</p>
      <p>The sequence of arguments put forward in the game
supports the final claims made by both players. An argument
 for player   at stage  ∈ {1, .., } composes of two
elements: a slot selection  and a claim , such as in the
following:</p>
      <p>= (, ).</p>
      <sec id="sec-3-1">
        <title>Symbolic learning via Slot Attention We employ Slot</title>
        <p>
          Attention (SA) [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] primarily due to its lightweight design,
efectiveness in unsupervised object discovery, and
architectural compatibility with the downstream task. The output
object representations (slots) serve as symbols for the
Consensus Game, the reasoning mechanism of our framework.
SA is trained as part of an encoder-decoder architecture that
encourages object-centricity through a reconstruction loss.
The encoder is a shallow CNN that maps an input image of
shape (, , ,  ) to a dense feature map. Slot Attention
then compresses these features into  slot vectors of
dimension , yielding an output of shape (, , ). The decoder
reconstructs the image from these slots, producing both the
full reconstruction and individual slot reconstructions with
attention masks.
        </p>
        <p>Agent modelling Player  ’s policy   : ℋ →  maps
the current game history ℋ, comprising of all arguments
put forward, at time step  to the next action , which is
an argument  at turn  of the game. Here, we emphasise
the distinction between time step  and turn . Multiple
players can share the same turn number , but the time step
 always increases uniquely with each action.</p>
        <p>Each player in the Consensus Game operates as a neural
agent composed of several cooperating components, as seen
in Figure 2. The architecture is designed to process partial
observations (slots) sequentially, maintain memory of prior
selections, and make decisions that improve classification
accuracy over time. The player is implemented as a
composition of six main components: a recurrent state encoder,
an index embedder, a modulator, a policy network, and a
classifier. These components are trained jointly to optimise
the performance objective of succinct informative selections
and accurate classifications.</p>
        <p>
          To reason over sequences of slot selections, each player
maintains an internal memory of the selection history ℋ
using a Recurrent Neural Network (RNN), capturing the
evolving state of the game from their perspective. At turn
, player   observes a selected slot with embedding 
and a index embedding  (adapted from transformers [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]),
where  denotes the player   who originally selected the
slot, which could be either player. The index embedder is a
learned embedding layer that maps slot indices to fixed-size
vectors, enabling the model to distinguish between diferent
slot indices.
        </p>
        <p>These embeddings are summed and passed through the
player-specific modulator network ℳ, a lightweight
multilayer perceptron (MLP) that reparameterises the input into
a more informative representation:
 = ℳ( + )
ℎ = ℛ(, ℎ− 1)</p>
        <p>This representation is then fed into the RNN ℛ of player
 , which updates its hidden state ℎ at time step :</p>
        <p>To enable the policy network to make informed
decisions from the very beginning of the game, we initialise the
RNN hidden state ℎ0 with knowledge of the entire slot pool.
This is done by sequentially processing all available slot
embeddings through the RNN prior to the start of the game.
However, this initialisation incorporates full information,
which is incompatible with our intention with the classifier,
where predictions should be based solely on cumulative
evidence. To resolve this, we maintain two separate RNN
hidden states per player   throughout the game. The first,
denoted ℎΠ,, is used by the policy network and is initialised
with all slots, providing a strong prior for action selection.
The second, ℎ,, is used by the classifier and is initialised to
zero, updating only as new slots are selected during
gameplay. This separation ensures that the classifiers operate
under partial information constraints, while the policy
network benefits from full-slot context for strategic selections.
The two RNNs may share architecture and parameters, but
they operate over disjoint information flows:
ℎΠ, = ℛΠ,(, ℎΠ− ,1)
ℎ, = ℛ,(, ℎ−,1)
(Policy path)
(Classifier path)</p>
        <p>The player’s policy network is an MLP that produces
a distribution over available actions at each time step. It
receives the hidden state ℎ as input and produces a logit
vector  , which is passed through a softmax to define a
stochastic policy:</p>
        <p>= Π (ℎΠ,)</p>
        <p>Central to each player is a classifier that integrates the
selected slot features to predict the true class of the input
instance. Each player maintains its own classifier instance,
which receives the cumulative set of slot embeddings
chosen up to the current time step. The classifier produces a
distribution over class labels via a softmax layer, and the
confidence assigned to the true class label is used to
compute both sparse and dense reward signals, depending on
the training regime. The classifier thus serves both as a
target for individual inference and as a feedback channel to
guide each player’s learning.</p>
        <p>Reasoning via consensus game The players as
described above participate in the Consensus Game, where
they aim to learn optimal strategies for correct, efective and
concise reasoning for classification. Our Consensus Game
module enables any number of agents to iteratively select
slot information with the goal of jointly identifying the
correct class label. This interaction is framed as a self-play
reinforcement learning game, where each agent must
communicate selectively to reach a consensus. The module can
be seen in two complementary ways: as an ante-hoc visual
explanation mechanism that reveals how decisions emerge
through inter-agent reasoning, or as a novel explainable
classification approach leveraging multi-agent reasoning.
Following that, we developed an environment (Figure 3) for
facilitating transparent player interactions. During training,
it provides reward feedback to the players.</p>
        <p>For reward feedback, we have a shared dense reward
function Γ based on player confidence toward the true
label. This reward structure encourages both agreement
and correctness, rewarding agents only when consensus
aligns with the ground truth, and penalising both blind
agreement and disagreement.</p>
        <p>At every turn  of the game of length , each player
outputs a probability distribution over class labels. Let 
denote the probability assigned to the ground-truth class 
by player  at turn . The reward for that turn is computed
as the gain in confidence relative to a baseline value 0,
which can be initialised as zero or the uniform probability
||− 1. The per-turn reward is thus:
 =  − 0</p>
        <p />
        <p>This formulation incentivises agents to select informative
slots that gradually build up confidence in the correct class,
even before reaching the final prediction step. We also
experimented with an alternative reward scheme based on
relative changes in confidence between players, defined as:
 =  −

−−  1</p>
        <p>This approach also improved learning eficiency but
introduced a risk of reward hacking, where agents could suppress
early confidence to artificially inflate perceived
improvement later. This undermined the integrity of the learning
signal. As a result, we adopted the more stable and
interpretable absolute-confidence-based dense reward
formulation. The shared utility function Γ is then defined as the
expected reward over the joint policy of both agents:
Γ( 1,  2) = E1,2∼ ( 1, 2) [︀ Γ(1, 2,  )]︀ ,
where  1 and  2 denote the policies of players 1 and
2, respectively. During training, the agents are optimised
jointly to maximise Γ, promoting cooperative behaviour
that balances consensus with predictive accuracy.</p>
        <p>In the Consensus Game, each player seeks to maximise
their expected utility by learning a policy that governs their
slot selection and final claim. Since both agents share a
reward function Γ and are penalised for disagreement or
incorrect consensus, their optimal strategy is cooperative.</p>
        <p>Let   denote the policy of player . Each policy maps
a game state history to a distribution over available
actions. The goal of each player is to learn a policy  * that
maximises their expected reward under the joint policy
 = ( 1,  2):
 * = arg max E [Γ]</p>
        <p />
        <p>The model is trained end-to-end through a combination
of reinforcement learning and supervised learning
objectives. Each epoch consists of multiple consensus games
played over batches of input instances. For every batch,
slot embeddings and class labels are extracted and passed
to the game environment, which orchestrates the game
dynamics such as the starting player, turn order, and game
termination. To mitigate ordering bias, the starting player
is alternated across batches. During training, the number of
turns per game is fixed, simplifying reward attribution and
ensuring consistent episode lengths. During inference, the
Consensus Game operates very similarly to training mode
depending on the end condition. Instead of running for a
ifxed number of turns, the game terminates once a certain
condition is satisfied or until a maximum number of turns.
This setup enables the agents to reach a dynamic consensus
based on selected factors, rather than being constrained to
a fixed interaction length. These end conditions are based
on one of or some of the following: repetition, confidence
and consensus, as elaborated in Section 4.</p>
      </sec>
      <sec id="sec-3-2">
        <title>End-to-end training The Expectation</title>
        <p>
          Maximisation [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] algorithm inspires our training
algorithm (Figure 4), where we have two maximisation
steps aimed at maximising the reward outcome. In the first
step, the Slot Attention module acts as the latent variable
model, with the Consensus Game frozen. Slot Attention is
updated to provide slot representations that better support
downstream decision-making. The update comes in the
form of an augmented reconstruction loss, where we
subtract a scaled mean reward, thereby encouraging slot
encodings to also correlate with high-reward outcomes. In
the second step, the Consensus Game module is updated
        </p>
        <sec id="sec-3-2-1">
          <title>CLEVR-Hans7 (CF)</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>CLEVR-Hans7 (NCF)</title>
        </sec>
        <sec id="sec-3-2-3">
          <title>CLEVR-Hans3 (CF)</title>
        </sec>
        <sec id="sec-3-2-4">
          <title>CLEVR-Hans3 (NCF)</title>
        </sec>
        <sec id="sec-3-2-5">
          <title>Multi-dSprites A B C</title>
          <p>A
B
C
A
B
C
A
B
C
A
B
C
based on the current slots, with upstream Slot Attention
module frozen, which is equivalent to the training protocol
of the pipelined version of our architecture. The alternating
updates decouple the learning dynamics of the two
components and reduce interference.</p>
          <p>A practical setup would also warm up the Slot Attention
module by training it without the downstream Consensus
Game. This allows for efective learning in the Consensus
Game when training does begin for the agents following the
warm up. Alternatively, in the interest of time, we typically
load in pre-trained weights.</p>
          <p>Following EM training, the slot reconstructions look
identical to the ones obtained from the original pipeline setup,
suggesting that the added reward signal did not compromise
the perceptual quality of the extracted slots. By framing the
optimisation steps as an EM loop, we decouple the pressure
from the combined loss from Consensus Game, while
allowing some reward signal to flow back to the slot autoencoder.
We also observed incremental improvements in accuracy
and quality of explanations as a result of E2E learning.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>Datasets Suitable datasets must feature multiple objects
with compositional variability, where the images exhibit
diverse combinations of object attributes such as colour,
shape and position. Importantly, the images should be
labelled based on logical rule-based conditions such as “a red
square left of a green circle”. These requirements ensures
that the task has a well-defined reasoning component. Our
experimental setup utilises two datasets that fit the criteria:
CLEVR-Hans and a relabelled version of Multi-dSprites. We
focus our experimental eforts towards the CLEVR-Hans
dataset, particularly due to the non-confoundedness of its
test split, although we also present accuracy metrics for
the confounded validation split. CLEVR-Hans also provides
versions with 3 labels (CLEVR-Hans3) and with 7 labels
(CLEVR-Hans7).</p>
      <p>Baselines As the focus of this research is explainability,
we will be evaluating both the predictive performance and
the explanation quality of our framework. We used the
standard feature extractor + multi-layer perceptron (MLP)
framework for our two prediction baselines. Typical models
use a Convolutional Neural Network (CNN) for feature
extraction, which informs our first baseline. The other swaps
out the CNN for SA with slot size 128, where classification
is based on slots. Baseline results were collected using a
standard ResNet18 implementation trained on images of
resolution 224 × 224, and a Slot-based classifier evaluated
at resolutions of 64 × 64 and 128 × 128 across all datasets
(except for Multi-dSprites on the 128 × 128 configuration).
The baselines serve as performance anchors, particularly in
scenarios where full image observability and unconstrained
inference are available. The prediction accuracies are
detailed in Table 3.</p>
      <p>Conversely, to test our framework’s transparency and
reasoning quality, we compare with standard post-hoc
mechanisms. Namely, we apply Grad-CAM, LIME and Gradient</p>
      <sec id="sec-4-1">
        <title>Method</title>
        <sec id="sec-4-1-1">
          <title>LIME</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Grad-CAM (Slot)</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>OCEAN</title>
        </sec>
        <sec id="sec-4-1-4">
          <title>Gradient SHAP</title>
        </sec>
        <sec id="sec-4-1-5">
          <title>Grad-CAM MdS CH7 MdS</title>
          <p>CH7
MdS
CH7
MdS
CH7
MdS
CH7
92
100
85
81
100
100
65
100
73
96
SHAP to both prediction baselines. Explanations
generated from these setups would be used to compare with our
framework, presented in a survey.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Configs/Input Size →</title>
      </sec>
      <sec id="sec-4-3">
        <title>Dataset ↓</title>
        <p>CLEVR-Hans7 (CF)
CLEVR-Hans7 (NCF)
CLEVR-Hans3 (CF)
CLEVR-Hans3 (NCF)</p>
        <p>Multi-dSprites</p>
        <p>ResNet</p>
        <p>224
90.32%
83.54%
93.11%
60.93%
83.20%</p>
        <p>64
86.70%
86.66%
86.62%
85.60%
92.72%</p>
        <p>Slot MLP</p>
        <p>128
85.34%
86.39%
87.90%
88.76%
—</p>
        <p>Config A</p>
        <p>64
77.64%
70.15%
78.50%
52.58%
79.76%
Configurations To evaluate the impact of diferent
gameplay dynamics and learning settings, we trained three
representative configurations of the Consensus Game module.
These configurations are denoted as Config A, B and C. They
are summarised as follows:
• Config A: We constrain to 64 × 64 images for lower
resolution slots in the Consensus Game, while fixing
to a consensus-based end condition. Here, games
conclude when players arrive to the same claim
while having a confidence that exceeds the threshold
of 70% confidence.
• Config B: We augment the end condition of the
64 × 64 variant to trigger on repeated selections by
any player instead (repetition-based). To encourage
this behaviour and end games early, we reward
repeated good selections, which reinforce ideas and
the shared prediction.
• Config C: We repeat Config A for the 128 × 128
variant of the Slot Attention module, with some
modifications in the network widths to cater the
increase in data.</p>
        <p>Evaluation Metrics For evaluating the performance of
our method, we measure and compare various properties of
the framework, including slot reconstruction loss, consensus
rate, game length, slot uniqueness, and empty slot usage.
Formally: (i) slot reconstruction loss is the mean squared
error of the reconstructed image and the original image,
(ii) consensus rate measures the frequency of the agents
aligning on the same final claim, (iii) game length is the
number of arguments presented by all players during the
game, (iv) slot uniqueness represents the rate at which the
players select a unique slot, and (v) empty slot usage is the
percentage of empty slots being presented; “empty” meaning
slots that do not map to objects in the image. We identify
empty slots using a heuristic based on the attention values
of the corresponding slots.</p>
        <p>Results We utilise three representative configurations of
our model, denoted as Config A, B, and C. We report results
(Table 1), including accuracy and F1-Score, on three
synthetic multi-object datasets, CLEVR-Hans3, CLEVR-Hans7,
and Multi-dSprites. This evaluation will help quantify how
well our framework performs classification while
maintaining the structural constraints needed for interpretability
and partial observability of the entire input. In contrast,
our framework operates under stricter constraints, where
agents make predictions on partial observations through
sequential slot selection, and are jointly optimised for both
prediction and interpretability. We observed that the
configurations with the “Consensus” end condition outperformed
Config B with its repetition-based end condition. The drop
in performance can be attributed to the agents adopting
repetition-based strategies that are less aligned with the goal
of achieving consensus. Interestingly, Config B achieves
comparable performance on CLEVR-Hans3 with the other
configurations, likely due to the less complex combinations
of objects that are needed for a conclusive classification. We
also consider the slot reconstructions, ensuring that
interpretability and object-centricity were not compromised as a
result of E2E training.</p>
        <p>While these results are competitive wrt baselines in terms
of classification accuracy, it is important to contextualise the
numbers. Our model is not solely optimised for accuracy,
but it is explicitly designed to support transparent reasoning
chains aligned with human interpretability, as demonstrated
in Figure 5. However, our results for the non-confounded
split for CLEVR-Hans7 show promise when compared with
the metrics of the confounded split, displaying only a small
drop in performance. We leave further investigation on the
robustness and generalizability aspects to future work.</p>
        <p>Generally, we attribute the performance hit to partial
observability of the input image, where our framework makes
classifications on the slots that the players select. However,
the poor predictive performance for the non-confounded
cases likely stems from overfitting and an underlying
confusion between objects of the same shape, regardless of the
other properties. A contributing factor is the entangled slot
representations, which group correlated features like size,
colour and shape during training, making it hard for our
framework to decouple those attributes when these
correlations no longer hold. This entanglement makes it dificult
for the players to reason over fine-grained details that are
needed to perform well for the non-confounded test set.
This behaviour reflects the model’s tendency to overfit to
prominent but irrelevant features rather than selecting
holistically over the entire set of objects in the scene, which is
seen in (a), (b), and (d) of Figure 5.</p>
        <p>Survey As the main investigation was on the
explainability and interpretability of our framework, we conducted a
survey asking participants to qualitatively assess five
explanation methods: LIME, Grad-CAM, Grad-CAM on Slot
Attention, Gradient SHAP and our OCEAN framework.
GradCAM was the only explanation baseline built on top the
slot-based classifier chosen for this survey. This is due to
the explanations for LIME and Gradient SHAP being
indistinguishable from their CNN-based counterpart.</p>
        <p>Participants were tasked with guessing the predicted class
of the model based on the explanations alone. Explanations
generated from most explanation mechanisms rely on the
input image for highlighting regions of the image. Thus,
the survey includes six statements that the participants rate
using a Likert scale. They score each statement from 1
(Strongly Disagree) to 5 (Strongly Agree):
• It is clear which parts of the image is the explanation.
• The explanation was easy to interpret without prior
technical knowledge.
• The explanation would help me identify errors or
biases in the model’s prediction.
• The explanation only focused on relevant
parts/objects of the image.
• The explanation helped me understand why the
model predicted the class.</p>
        <p>• I am satisfied with the quality of the explanation.</p>
        <p>Our survey results (Table 2) reveal distinct diferences in
the interpretability and user satisfaction among the five
explanation methods evaluated. For each question, our
framework leads in the average score, indicating that our
framework provided good explanations for the input
images. While the results are favourable, we should cautiously
interpret their implications. In the context of synthetic,
multi-object datasets, our framework performs well due to
our object-centric encoder and sequential reasoning
module, making our generations easier for users to interpret. As
such, the generalisability of our framework to real-world
image datasets should be investigated to determine whether
the interpretability advantages of our framework persist for
more complex or noisy images.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <p>
        Through this work, we gained insight into the delicate
tradeofs between performance, transparency, and
interpretability. Faithful explanations require an architecture where
they can be derived naturally from the decision-making
process. Our findings show that multi-agent interaction,
when structured carefully, provides a promising scafold
for such reasoning. An immediate improvement of this
work would be upgrading the encoder component for a
more scalable Slot Attention variant that uses vision
transformer architectures, such as in [22] and [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Enabling us
to apply the OCEAN framework to real-world datasets such
as COCO [23] or AFHQ [24] could test its interpretability
and reasoning in more complex settings. Another direction
for future work would be to explore intentional concept
learning strategies on top of object-centric learning, where
object representations are further factorised into object
attributes. With finer-grained information, the agents would
be able to reason on object characteristics as well and be
able to learn more complex attribute-based patterns as seen
in CLEVR-Hans.
      </p>
      <p>Collectively, these challenges point toward the
importance of future research into more scalable object-centric
models, principled design for interpretable systems, and
robust evaluation frameworks for explainability. We learned
that explanations are not merely outputs to be generated
after the fact, but processes that should be embedded into
opaque systems to ensure transparency and trust. This shift
in perspective is shared across the industry with broader
implications for AI trustworthiness, particularly in high-stakes
domains like medical diagnosis or autonomous
decisionmaking. As a result, we are heading straight on to a world
where understanding why a model makes a decision is just
as critical as the decision itself.</p>
      <p>Acknowledgments
This research was partially supported by the ERC under
the EU’s Horizon 2020 research and innovation programme
(grant no. 101020934), and by EPSRC postdoctoral prize
fellowship.</p>
      <p>Generative-AI Declaration
Authors have employed the use of Generative AI tools, such
as ChatGPT, to refine/rephrase certain parts of the paper,
while maintaining the conceptual originality of the work.
IEEE Signal Processing Magazine 13 (1996) 47–60.
doi:10.1109/79.543975.
[22] M. Seitzer, M. Horn, A. Zadaianchuk, D.
Zietlow, T. Xiao, C.-J. Simon-Gabriel, T. He, Z. Zhang,
B. Schölkopf, T. Brox, F. Locatello, Bridging the gap to
real-world object-centric learning, in: The Eleventh
International Conference on Learning Representations,
2023. URL: https://openreview.net/forum?id=b9tUk-f_
aG.
[23] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona,
D. Ramanan, P. Dollár, C. L. Zitnick, Microsoft coco:
Common objects in context, in: European conference
on computer vision, Springer, 2014, pp. 740–755.
[24] Y. Choi, Y. Uh, J. Yoo, J.-W. Ha, Stargan v2: Diverse
image synthesis for multiple domains, in: 2020 IEEE/CVF
Conference on Computer Vision and Pattern
Recognition (CVPR), 2020, pp. 8185–8194. doi:10.1109/
CVPR42600.2020.00821.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nasser</surname>
          </string-name>
          , U. K. Yusof,
          <article-title>Deep learning based methods for breast cancer diagnosis: a systematic review and future direction</article-title>
          ,
          <source>Diagnostics</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>161</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Badue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidolini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Carneiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Azevedo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. B.</given-names>
            <surname>Cardoso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Forechi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jesus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Berriel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Paixao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mutz</surname>
          </string-name>
          , et al.,
          <article-title>Self-driving cars: A survey</article-title>
          ,
          <source>Expert systems with applications 165</source>
          (
          <year>2021</year>
          )
          <fpage>113816</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Selvaraju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cogswell</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Vedantam</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Parikh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Batra</surname>
          </string-name>
          , Grad-cam:
          <article-title>Visual explanations from deep networks via gradient-based localization</article-title>
          ,
          <source>in: Proceedings of the IEEE international conference on computer vision</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>618</fpage>
          -
          <lpage>626</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          ,
          <article-title>"why should i trust you?": Explaining the predictions of any classiifer</article-title>
          ,
          <source>in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , KDD '16,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2016</year>
          , p.
          <fpage>1135</fpage>
          -
          <lpage>1144</lpage>
          . URL: https://doi.org/10.1145/2939672.2939778. doi:
          <volume>10</volume>
          . 1145/2939672.2939778.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sarkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vijaykeerthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sarkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. N.</given-names>
            <surname>Balasubramanian</surname>
          </string-name>
          ,
          <article-title>A framework for learning ante-hoc explainable models via concepts</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>10286</fpage>
          -
          <lpage>10295</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Jacob</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Farina</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Andreas,</surname>
          </string-name>
          <article-title>The consensus game: Language model generation via equilibrium search</article-title>
          ,
          <source>arXiv preprint arXiv:2310.09139</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Glocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>Explaining image classification with visual debates</article-title>
          ,
          <source>arXiv preprint arXiv:2210.09015</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rago</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>Free argumentative exchanges for explaining image classifiers</article-title>
          ,
          <source>arXiv preprint arXiv:2502.12995</source>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>W.</given-names>
            <surname>Stammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schramowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kersting</surname>
          </string-name>
          ,
          <article-title>Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations</article-title>
          ,
          <source>in: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>3618</fpage>
          -
          <lpage>3628</lpage>
          . doi:
          <volume>10</volume>
          . 1109/CVPR46437.
          <year>2021</year>
          .
          <volume>00362</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P. W.</given-names>
            <surname>Koh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. S.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mussmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pierson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Concept bottleneck models</article-title>
          , in: H.
          <string-name>
            <surname>D. III</surname>
          </string-name>
          , A. Singh (Eds.),
          <source>Proceedings of the 37th International Conference on Machine Learning</source>
          , volume
          <volume>119</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>5338</fpage>
          -
          <lpage>5348</lpage>
          . URL: https://proceedings. mlr.press/v119/koh20a.html.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          , W. Zhu,
          <article-title>Disentangled representation learning</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>46</volume>
          (
          <year>2024</year>
          )
          <fpage>9677</fpage>
          -
          <lpage>9696</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Locatello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weissenborn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mahendran</surname>
          </string-name>
          , G. Heigold,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          , T. Kipf,
          <article-title>Object-centric learning with slot attention</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>11525</fpage>
          -
          <lpage>11538</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barnett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rudin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <article-title>This looks like that: Deep learning for interpretable image recognition</article-title>
          , in: H.
          <string-name>
            <surname>Wallach</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
          </string-name>
          , F.
          <string-name>
            <surname>d'Alché- Buc</surname>
          </string-name>
          , E. Fox, R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>32</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2019</year>
          . URL: https: //proceedings.neurips.cc/paper_files/paper/2019/file/ adf7ee2dcf142b0e11888e72b43fcb75-Paper.pdf .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Alvarez Melis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jaakkola</surname>
          </string-name>
          ,
          <article-title>Towards robust interpretability with self-explaining neural networks</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>31</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Gref</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. L</surname>
          </string-name>
          . Kaufman, R. Kabra,
          <string-name>
            <given-names>N.</given-names>
            <surname>Watters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Burgess</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zoran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Matthey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Botvinick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lerchner</surname>
          </string-name>
          <article-title>, Multi-object representation learning with iterative variational inference</article-title>
          ,
          <source>in: International conference on machine learning, PMLR</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2424</fpage>
          -
          <lpage>2433</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Locatello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. D. S.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Glocker</surname>
          </string-name>
          ,
          <article-title>Grounded object centric learning</article-title>
          ,
          <source>arXiv preprint arXiv:2307.09437</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Locatello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Santhirasekaram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Glocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Sousa Ribeiro</surname>
          </string-name>
          ,
          <article-title>Identifiable objectcentric representation learning via probabilistic slot attention</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>37</volume>
          (
          <year>2024</year>
          )
          <fpage>93300</fpage>
          -
          <lpage>93335</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>N.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Eastwood</surname>
          </string-name>
          , R. Fisher,
          <article-title>Learning object-centric representations of multi-object scenes from multiple views</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>5656</fpage>
          -
          <lpage>5666</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>V.</given-names>
            <surname>Ashish</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ) I.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>T.</given-names>
            <surname>Moon</surname>
          </string-name>
          ,
          <article-title>The expectation-maximization algorithm,</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>