<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>WOA</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards Intelligent Pulverized Systems: a Modern Approach for Edge-Cloud Services⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Davide Domini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicolas Farabegoli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianluca Aguzzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mirko Viroli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Università di Bologna - ALMA MATER STUDIOURM</institution>
          ,
          <addr-line>Via Dell'Università 50, 47521 Cesena</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>25</volume>
      <fpage>8</fpage>
      <lpage>10</lpage>
      <abstract>
        <p>Emerging trends are leveraging the potential of the edge-cloud continuum to foster the creation of smart services capable of adapting to the dynamic nature of modern computing landscapes. This adaptation is achievable through two primary methods: by leveraging the underlying architecture to refine machine learning algorithms, and by implementing machine learning algorithms to optimize the distribution of resources and services intelligently. This paper explores the latter approach, focusing on recent advancements in pulverized architecture, collective intelligence, and many-agent reinforcement learning systems. This novel trend, which we refer to as intelligent pulverized system (IPS), aims to create a new generation of services that can adapt to the complex and dynamic nature of the edge-cloud continuum. Our proposed learning framework integrates many-agent reinforcement learning, graph neural networks, and aggregate computing to create intelligent services tailored for this environment. We discuss the application of this framework across diferent levels of the pulverization model, illustrating its potential to enhance the adaptability and eficiency of services within the edge-cloud continuum.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Edge Cloud Continuum</kwd>
        <kwd>Many-Agent Reinforcement Learning</kwd>
        <kwd>Pulverization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Recent technological developments are fostering a computational landscape that is increasingly
articulated and complex across various levels. The historical distinction between cloud, fog,
and edge computing is becoming progressively blurred, giving rise to what is known as the
edge-cloud continuum (ECC) [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. This shift is driven by the necessity for services developed
on these platforms to be highly opportunistic, capable of dynamically moving up and down the
continuum based on local needs and computational requirements.
      </p>
      <p>
        The ECC is a valuable resource for creating intelligent services that leverage recent
advancements in machine learning [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. These services must be able to utilize the continuum’s full
potential by adapting to the changing conditions and demands of the environment. Furthermore,
making the continuum itself more intelligent is essential for developing complex applications,
particularly those that are collective in nature. Modern applications, such as those in
pervasive [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], collective [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and ubiquitous [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] computing, often require collective computations
rather than mere local computations. Examples include applications for crowd control, trafic
management, and energy monitoring [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. These scenarios highlight the need for a holistic
view of the system rather than just local optimizations.
      </p>
      <p>
        This paper proposes a new vision of intelligent pulverized system (ISP) for creating a smarter
ECC by leveraging recent developments in many-agent reinforcement learning [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
pulverization [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ], and macroprogramming [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Pulverization is a modern approach that describes
collective computations capable of being “broken down” and distributed across multiple hosts,
which is crucial for fully exploiting the continuum [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Additionally, we explore how
macroprogramming approaches can efectively capture the collective aspect of these systems, providing
a comprehensive framework for intelligent service development within the ECC [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This
structured approach aims to enhance the intelligence of the ECC, enabling the development of
sophisticated, adaptive, and collective applications that can eficiently utilize the continuum’s
capabilities.
      </p>
      <p>The rest of the paper is structured as follows: Section 2 provides an overview of the
edgecloud continuum, its characteristics, and the concept of pulverization; Section 3 discusses the
integration of many-agent reinforcement learning, GNNs, and aggregate computing to develop
intelligent services; Section 4 details the application of intelligent collective services to various
levels of pulverization, including intelligent reconfiguration, adaptive communication, and
scheduling; and Section 5 summarizes our findings and outlines potential future directions for
research in this area.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <sec id="sec-2-1">
        <title>2.1. Edge-Cloud Continuum</title>
        <p>
          The rapid adoption of cloud technologies was driven by the benefit of having high availability,
and scalability of computational resources. The rise of cloud computing has led to the
proliferation of new applications involving smart devices; these applications are typically reified in
the Internet of Things (IoT) [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] and Cyber-Physical System (CPS) [14] domains. To meet the
increasing demand of low-latency and high-throughput of such applications, the edge
computing [15] paradigm has been introduced, to bring computational resources closer to end users,
thus reducing latency and network trafic.
        </p>
        <p>
          In this context, the highly distributed nature of such applications and the vast amount of
generated data have driven the shift from a centralized cloud computing paradigm to a more
distributed model, where the cloud is integrated with the edge. Such integration gave rise to a
new, hybrid paradigm, called edge-cloud continuum [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ]. Depending on the specific scenario
and scope, the Edge-Cloud Continuum (ECC) can have slightly diferent meanings [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Generally,
the ECC refers to a distributed computing environment that extends from the cloud to the edge,
where this continuum can be efectively leveraged to optimize metrics and quality of service.
        </p>
        <p>This new paradigm is characterized by a high level of heterogeneity in terms of devices,
networks, and services. Consequently, various types of hardware devices are part of the
continuum, ranging from high-end servers to wearable or embedded devices equipped with
microcontrollers. Ideally, any devices equipped with a computational unit and a network
connection can be part of the continuum.</p>
        <p>This paradigm includes diverse hardware and software stacks (e.g., Linux, embedded firmware,
Android Wear, Docker), and various network technologies (e.g., Ethernet, WiFi, ZigBee, LoRa,
5G/6G), making the ECC a highly heterogeneous and dynamic environment, challenging to
manage in terms of resource allocation and service deployment. While the continuum ofers new
possibilities to optimize the performance of applications, the complexity of the environment
makes it dificult to fully exploit the full potential of the ECC, fostering the need for novel
intelligent solutions.</p>
        <p>
          The rapid increase in data volumes from various applications is driving the evolution of
distributed digital infrastructures for data analytics and Machine Learning (ML). Traditionally
reliant on cloud infrastructures, the need for low-latency and secure processing has shifted
some processing to IoT edge devices, finding in the ECC a prominent solution for the distributed
intelligence [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Macroprogramming</title>
        <p>
          When dealing with large-scale systems, such as ECC, it is helpful to shift the focus from
individual devices to the collective system, as the behavior of the system as a whole is more
important than the behavior of individual devices. Consider, for example, a crowd congestion
alarm system. In this case, each smartphone could be part of an ECC system aimed at identifying
crowded areas and intelligently guiding the crowd to disperse, thereby avoiding emergencies.
Programming such behavior with a device-centric view is complicated because the collective
paradigm is not incorporated at that level. In this direction, a modern area of research aims to
change the programmer’s focus from the device to the aggregate. This area of research is called
macroprogramming [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], which has its roots in Wireless Sensor Networks (WSN) [16] and has
evolved to be used in opportunistic edge computing contexts today [17].
        </p>
        <p>The core of macroprogramming lies in identifying a collective abstraction that then becomes
a first-class citizen of the programming language. Various languages have been defined in this
scenario, but most of them are designed for specific applications (e.g., PyoT for IoT [ 18] and
Buzz [19] for robotics). A modern approach to macroprogramming is aggregate computing [20],
which is a functional top-down global-to-local paradigm that allows defining collective behaviors
through the manipulation of a distributed data structure called computational fields . This macro
view is then broken down into local executions of individual devices, which, by computing
iteratively and continuously, achieve the specified collective behavior.</p>
        <p>
          Aggregate computing has been used in the context of ECC to create additive behaviors based
on devices’ location in space [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. In this paper, however, we will focus on using this collective
abstraction to better guide machine learning in the ECC context. Indeed, aggregate computing
was shown to be a powerful tool to guide machine learning in the context of large-scale
distributed systems [21, 22, 23].
Pulverised
Component
        </p>
        <p>Logical
Device
Physical
Device
Virtual
Link
Physical
Link
e
goL ayL
l
h L
P</p>
        <p>Neighbor Devices
Logical Device





 
 
 
 


abstraction levels: the logical layer, the physical layer, and the infrastructure layer. On the right side,
the logical device pulverized into the five subcomponents.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Pulverization</title>
        <p>
          Pulverization [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ] is an approach to simplify the design and deployment of (collective)
distributed applications, by devising a peculiar partitioning into independently deployable
components. The main goal is to provide the developer with a way to specify the application
logic in a deployment-agnostic way, and let the platform/middleware take care of the deployment
and the communication between the components.
        </p>
        <p>In this context, this approach mainly targets the deployment of collective applications,
where through macroprogramming, the application is designed as a composition of collective
components (collective layer depicted in Figure 1). To achieve this, the pulverization model
defines two main abstraction levels: the logical layer and the physical layer, as depicted in Figure 1.
The former is the level at which the developer reasons about the application logic, while the
latter is the level at which the specified application is deployed and executed.
System Model at the logical level, the developer specifies the application as an ensemble
of application-specific devices forming an arbitrary topology. An application-specific device
can be “pulverized” if it can be decomposed into pulverized components representing: a set of
sensors ( ), a set of actuators ( ), a state holding the device’s knowledge ( ), a communication
component to interact with other devices ( ), and a behavior component defining the device’s
logic ( ). The Figure 1 shows a logical device pulverized into the five subcomponents (right side
of the figure).</p>
        <p>Typically, the logical layer is composed of several (hundreds or thousands) devices forming
a dynamic graph. Each logical device has a neighborhood with which it can interact. Such
neighbourhoods can vary over and the way the neighborhood is defined can be
applicationspecific (like proximity-based of based on physical connections). In pulverized system, the 
component is in charge of managing the communication between the devices, implementing
the neighborhood definition and the message exchange.</p>
        <p>Execution Model the components interact with each other to perform a MAPE-K like loop:
the sensors ( ) collect data from the environment and the neighbour’s messages are received
by the communication component ( ), the behavior component ( ) applies the logic taking
as input the state ( ) and the data from the sensors ( ), and the neighbor’s messages. The
produced output contains the new state ( ), the messages to be sent to the neighbors via the
communication component ( ), and the prescriptive actions to be executed by the actuators ( ).</p>
        <p>Notably, the flexibility of this model allows the implementation of such an interaction loop
in many ways. The most common way to implement this interaction is via a round-based
execution model, where each component at fixed intervals sends and/or receives messages
from the other components. However, more complex implementation can be devised, like a
pure reactive model where the components react to the messages received without a fixed
schedule; in this case, the  component leads (intelligently) the interaction loop. Another way
to implement such a loop is to adopt a “best-efort” approach, where all components requiring
input from the other components compute their logic against the most recent available data.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Flexible pulverized Deployments for the ECC The pulverization approach is particularly</title>
        <p>suitable for collective applications, enabling non-trivial deployments. When macroprogramming
like aggregate computing [20] are used, the pulverization model naturally integrates with them,
simplifying development, especially in ECC infrastructures. Otherwise, the pulverization can
be exploited to design custom systems with the pulverization in mind.</p>
        <p>Whenever an application is pulverized, a deployment mapping must be provided between
the logical and the infrastructure levels. Typically, in this mapping, a logical device is typically
deployed on either multiple physical devices or a single one. A logical device is often associated
with an application-level physical device that needs to be managed or controlled (e.g., a sensor,
a drone, or a person equipped with wearables). However, the capabilities of that physical
device might not be enough to host all the components, and the pulverization might require
the deployment of the components on diferent physical devices, namely infrastructure-level
devices. The main advantage of pulverization is enabling many deployments without afecting
or changing the application logic. Thanks to this feature, a system can be deployed on arbitrary
infrastructural devices as long as they provide the required capabilities to host the components
(for example, a sensor component can be executed only on devices hosting the appropriate
sensors, and the behavior component can be allocated on that host having suficient computational
power). In this sense, the ECC can be exploited to opportunistically allocate the components
making possible deployments otherwise impossible.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.4. Machine Learning for ECC</title>
        <p>The emergence of the edge-cloud continuum ofers significant potential for the development of
new intelligent systems, particularly for applications that rely on machine learning techniques.
In the following, we provide the current state of the art in the field of machine learning for the
edge-cloud continuum.</p>
        <p>
          ECC for ML Applications ECC has attracted significant attention from the ML community:
highly distributed networks composed of devices capable of generating a large amount of data
can provide new resources to enhance the quality of ML models. However, these devices often
have limited computing resources and latency constraints that prevent both local learning
and ofloading learning tasks to cloud servers. Conversely, the ECC provides a great platform
with edge servers that add computational resources close to where the data is generated. In
recent years, several studies in the literature have proposed leveraging these new resources,
for instance: i) Real-time video streaming analysis for applications such as trafic control,
surveillance, security, and real-time object detection [
          <xref ref-type="bibr" rid="ref14 ref15 ref16">24, 25, 26</xref>
          ]; ii) Smart vehicular systems for
pedestrian recognition and crash detection [
          <xref ref-type="bibr" rid="ref17 ref8">8, 27</xref>
          ]; and iii) Smart cities for energy management
and fault detection [
          <xref ref-type="bibr" rid="ref18 ref7">7, 28</xref>
          ].
        </p>
        <p>
          Machine Learning for ECC machine learning can be exploited to optimize the deployment of
services in the ECC, since leveraging its potential in real-life scenarios is a non-trivial task. The
devices are heterogeneous and have multiple constraints (e.g., energy, latency, computational
power). For this reason, it is essential to adapt the deployment dynamically and
opportunistically. In the literature, several approaches based on heuristics and meta-heuristics have been
proposed [
          <xref ref-type="bibr" rid="ref19 ref20 ref21">29, 30, 31</xref>
          ]; nonetheless, these solutions are crafted for specific use cases and do not
adapt over time. In recent years, solutions based on supervised learning techniques [
          <xref ref-type="bibr" rid="ref22 ref23 ref24">32, 33, 34</xref>
          ]
have been proposed to predict the characteristics of microservices in order to find solutions that
meet various constraints such as quality of service and energy consumption. This approach
requires a vast amount of data that represents the various possible states of the system and
the possible actions to be taken, which is challenging given the complexity of these systems.
For these reasons, reinforcement learning techniques are emerging as a natural approach. This
paradigm allows the optimization of the decision-making process by observing the evolution
of the system over time, without the initial need to collect a large amount of data. In [
          <xref ref-type="bibr" rid="ref25">35</xref>
          ] the
authors introduce a solution based on deep reinforcement learning (specifically using a Dueling
DQN model [
          <xref ref-type="bibr" rid="ref26">36</xref>
          ]) to optimize the deployment of microservices. Initially, the global state of the
system is observed (i.e., CPU usage, memory, and network bandwidth). These observations are
then passed to the neural network, which predicts end-to-end latency and peak throughput for
each possible action, thereby defining a new resource partitioning. The work in [
          <xref ref-type="bibr" rid="ref27">37</xref>
          ], instead,
proposes a solution based on distributed deep reinforcement learning (using an actor-critic
model) to address the task ofloading problem, i.e., deciding which available device will execute
a certain task. To achieve this, tasks are divided into three categories, namely: i) delay-sensitive
tasks; ii) energy-sensitive tasks; and iii) insensitive tasks. At this point, each device runs its
actor network to decide whether a task can be executed locally, on an edge server or in the cloud.
Finally, a centralized critic network evaluates the actions taken by the various devices, providing
feedback on their efectiveness. In addition to the aforementioned studies, other works leverage
reinforcement learning for various optimization tasks in the ECC. For instance, [
          <xref ref-type="bibr" rid="ref28">38</xref>
          ] employs
ofline RL for task scheduling, [
          <xref ref-type="bibr" rid="ref29">39</xref>
          ] uses RL to optimize data caching in edge servers, and [
          <xref ref-type="bibr" rid="ref30">40</xref>
          ]
applies RL to enhance network resource allocation.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Intelligent Collective Services for ECC</title>
      <p>The pulverization model proposes logical systems in which a series of logical entities form a
dynamic graph that exposes the following properties:
• large scale: an ECC system can be extensively scaled to encompass thousands of devices;
• locality: edge devices are spatially located, and their behavior may depend on their
location;
• partial observability: edge devices do not have a perception of the entire collective but
can perceive a certain neighborhood or aggregated information;
• heterogeneity: devices belonging to ECC are highly heterogeneous due to their diverse
nature.</p>
      <p>
        Applying standard supervised learning techniques to ECC services is challenging because
it is dificult to determine the correct behavior a priori, given the highly dynamic nature of
these systems. Therefore, considering these properties, we argue that a combination of
manyagent reinforcement learning [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] (to encode large-scale dynamics), graph neural networks [
        <xref ref-type="bibr" rid="ref31">41</xref>
        ]
(to encode spatial relationships), and aggregate computing (to encode collective feedback and
observations) can be a suitable approach for developing intelligent services for ECC. In particular,
many-agent reinforcement learning plays a crucial role in ECC, as it enables scalable policy
learning and influences behavior through delayed collective feedback.
      </p>
      <p>In the following, we introduce both many-agent reinforcement learning and graph neural
networks, and then discuss the proposed algorithm for developing intelligent services for ECC.</p>
      <sec id="sec-3-1">
        <title>3.1. Formalization</title>
        <p>
          Many-Agent Reinforcement Learning is an extension of reinforcement learning in which,
instead of having a single learning agent, there is an entire family of intelligent agents that learn
concurrently. They are referred to as “many-agent” to distinguish them from “multi-agent”,
as the number of devices can be extremely large, potentially involving thousands of devices.
Formally, a many-agent system can be modeled through a family of agent A = (, , , ℛ,  )
that lives and interact in an environment ℰ = (, A,  ,  ) [
          <xref ref-type="bibr" rid="ref32">42</xref>
          ]. Particularly, the agent A is
defined by:
• , , and  represent the sets of local states, observations, and actions, respectively.
• ℛ :  → R is the reward function, influenced by the environment.
•  :  →  is the policy mapping observations to actions, which can be deterministic or
stochastic.
        </p>
        <p>Where the environment ℰ is defined by:
•  is the fixed population of agents.
• A is the agent prototype governing each agent in .
•  :  ×   ×  →</p>
        <p>yielding a collective reward.
•  :  →  is the global observation model.</p>
        <p>R is the global transition function influenced by agent actions,</p>
        <sec id="sec-3-1-1">
          <title>The system’s evolution over time can be captured as follows:</title>
          <p>=  ( ()),</p>
          <p>+1 =  (, )</p>
          <p>
            At any given time , the system can be represented as a graph  = ( , ), where  is
constructed from a neighborhood relationship. Each node has an associated local observation
serving as input for a GNN, as demonstrated in previous works [
            <xref ref-type="bibr" rid="ref33 ref34 ref35">43, 44, 45, 21</xref>
            ].
 ∈ . This graph representation is crucial for both computing computational fields and
() through a process known as message passing [
            <xref ref-type="bibr" rid="ref36">46</xref>
            ].
          </p>
          <p>
            The process involves three main steps in each layer  of the GNN:
Graph Neural Network is a novel neural network model designed to process
graphstructured data [
            <xref ref-type="bibr" rid="ref31">41</xref>
            ]. Given a graph  = (, ), where  is the set of nodes and  ⊆  × 
represents the edges, each node  ∈  has an associated feature set . The aim of a GNN is to
learn a node embedding ℎ for each node  ∈  by aggregating information from its neighbors
m() =  () (︁
h(− 1), h(− 1),
          </p>
          <p>︁)
a() = ⨁︁ (︁
()</p>
          <p>{m() :  ∈ ()}︁)
h() = () (︁
h(− 1), a())︁
where m() is the message from node  to node  at layer , ⨁︀ denotes the aggregation
function, and () is the update function. Initially, ℎ0 is set to the node’s feature vector .
GNNS enables the efective processing of graph-structured data by iteratively updating node
embeddings, capturing the local graph structure around each node. This formulation facilitates
the application of GNNs in various tasks within the ECC, where the spatial relationships between
devices are critical.</p>
          <p>GNN ( ) = {h() :  ∈ ,  ∈ N}
Where  is the graph with features  for each node  ∈  , formally:  = (, , { :  ∈
 }). By leveraging the unique capabilities of GNNs, it is possible to develop more sophisticated
and adaptive services that can dynamically respond to the complex and distributed nature of
the ECC.
(1)
(2)
(3)
(4)</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Contribution</title>
        <p>
          In this section, we discuss an approach for performing many-agent reinforcement learning
leveraging GNN in pulverized systems. In this context, as in similar many-agent scenarios, we
apply the pattern known as centralized learning and distributed execution [
          <xref ref-type="bibr" rid="ref37">47</xref>
          ]. This is because,
at a large scale and with only partial observability, having a global view of the task helps in
better learning collective tasks.
        </p>
        <p>Learning dynamics We base our approach on the model of pulverization, learning under
the assumption that there exists a set of components that can act according to local logics –
here, we are not concerned with the type of actions. Specifically, as described in Algorithm 1,
for each time step , a graph  represents the connectivity between devices. Each node in
the graph corresponds to a local device, and edges represent the communication or influence
between devices. Each node  ∈ , has an associated observation  creating a decorated
graph  that represents. This same graph can be passed to a macro program  (e.g., aggregate
computing) to encode global system information (e.g., areas of high consumption in the case
of energy distribution, areas of the system more allocated for task distribution). The macro
program aggregates local information to provide a global perspective, which is crucial for the
GNN to make informed decisions. Formally, an evaluation of a macro program  with a graph
 produces a new graph , where for each node, associate the previous local state and
the evolution of the macro program:  = (,  ()). This graph is then passed to a GNN
that will compute the actions a ∈   to be taken. These actions will be executed in the
environment (i.e., the various logical components will execute what they need to), transitioning
to a new collective graph +1 and obtaining a reinforcement signal R . The reinforcement
signal R  can be split into local and global components. The local component is derived from
individual device metrics (e.g., battery consumption), while the global component is computed
from aggregate metrics (e.g., average consumption in a region) that can be also computed by
another macro program  R.</p>
        <p>
          Algorithm 1 GNN-Based many-agent reinforcement learning in ECC
1: Initialize: Local device observations o0
2: for each time step  do
3: Create decorated graph  with observations o
4: Apply macro program  to  to produce aggregate observations 
5: Input  into GNN to compute actions a ∈ 
6: Execute actions a in the environment
7: Transition to new graph +1 and obtain reinforcement signal R
8: Calculate local and collective reinforcement signals
9: Update GNN with new graph +1, observations o+1, and reinforcement signals R
10: end for
Implementation discussion This section proposes some hints for the future concrete
implementation of the algorithm described above and in Algorithm 1. First, it is worth noting
that the proposed solution is completely agnostic to the choosen many-agent reinforcement
learning algorithm. The only constraint is that such algorithm must support the centralized
learning and distributed execution model. On the one hand, a classic approach is to implement
the DQN algorithm [
          <xref ref-type="bibr" rid="ref38">48</xref>
          ] to approximate a function that, for each observation, provides the
most valuable action (i.e., value based methods). On the other hand, policy based algorithms
exists. For instance, PPO [
          <xref ref-type="bibr" rid="ref39">49</xref>
          ] is based on a simple actor and critic method. In this framework,
each agent has an actor network which is used to select the current action, while a central
entity has a critic network used to reward each action. Thus, this method fits the centralized
learning and distributed execution contraint. Moreover, a study shows the efectiveness of PPO
in multi-agent cooperative settings [
          <xref ref-type="bibr" rid="ref40">50</xref>
          ]. To avoid issues, such as the curse of dimensionality
and the exponential growth of interactions between agents, related to the large number of
agents present in the reference context, the described approaches can be integrated with modern
solutions that make learning more stable. For example, mean-field reinforcement learning [
          <xref ref-type="bibr" rid="ref41">51</xref>
          ]
integrates both DQN and actor-critic approaches by ensuring that each agent is influenced not
by the individual actions of neighbors but by their average. Additionally, also the GNN can be
distributed in the system in such a way that it can be executed in a distributed manner without
the need for a central cloud [
          <xref ref-type="bibr" rid="ref33">43</xref>
          ]. Furthermore, thanks to the new self-organizing approach of
the self-organising coordination regions [
          <xref ref-type="bibr" rid="ref42">52</xref>
          ] pattern, these learning points are not necessarily
known a priori and can also change over time, therefore enabling continual learning [
          <xref ref-type="bibr" rid="ref43">53</xref>
          ]. This
distributed approach is further enhanced by self-organizing patterns, allowing for dynamic
adaptation and continual learning in changing environments.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Intelligent Pulverized System</title>
      <p>In this section, we explore how the concept of intelligent collective services can be applied to
various levels of pulverization, building upon the proposed model to enhance system eficiency
and adaptability. In particular, in the following, we focus on three main aspects: (i)
reconfiguration: how reconfiguration policies can be used to optimize the deployment of pulverized
components at runtime; (ii) communication: how communication protocols can be learned
to optimize the exchange of information between devices; and (iii) scheduling: how adaptive
scheduling policies can be used to optimize the execution of pulverized components.</p>
      <sec id="sec-4-1">
        <title>4.1. Reconfiguration</title>
        <p>Discussion In the ECC, reconfiguration is a key component. Specifically, with pulverization,
reconfiguration allows for the ofloading of one or more components of a logical device (right
side of Figure 1) from one host to another to meet certain constraints. These constraints can
be expressed through reconfiguration rules, which can be either local (e.g., the host’s battery
drops below a certain threshold) or global (e.g., defined via aggregate computing) to preserve
global coherence and prevent oscillatory conditions. While these approaches can be efective in
simple scenarios, they have limitations in more complex situations: (i) they can only represent
relatively simple rules; (ii) it is challenging to define all rules a priori, as they can be numerous
and complex; and (iii) the rules are fixed, so if the system changes over time, it cannot adapt.
Scheduling</p>
        <p>Smart
Communication
Intelligent Components</p>
        <p>Reconfiguration
Intelligent Scheduling manage the intelligent execution of the behavior, Smart Communication supports
the  component in eficient communication and neighborhood definition, and
Intelligent Components
Reconfiguration manage the Intelligent and opportunistic relocation of the pulverized components,
improving the system’s performance.</p>
        <p>Vision</p>
        <p>
          we believe that an interesting direction is to learn these rules through the proposed
many-agent reinforcement learning approach. Specifically, using online reinforcement learning
techniques can help manage potential changes or specific cases that were not captured
beforehand. In the literature, some works have attempted to explore this field [
          <xref ref-type="bibr" rid="ref25 ref44 ref45">35, 54, 55</xref>
          ]. However,
they focus on entire microservices. Our approach ofers the following advantages:
1. it is possible to operate more granular, not being forced to relocate the entire logical
device, but rather being able to act at the level of individual components;
2. using macroprogramming and graph neural networks, it is possible to integrate
neighborhood information, thereby enriching the knowledge of individual devices regarding
the global state and objective of the system;
3. it is possible to define complex reconfiguration rules and to learn new rules not known at
design time.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Communication</title>
        <p>Discussion</p>
        <p>in collective systems, communication represents a crucial aspect: thanks to
device communication and coordination, a global common goal can be achieved. In this sense,
the  component of the pulverization model is responsible for managing the
neighborhoodbased communication between devices. For this reason, understanding what, when, and with
whom communicate is crucial for efective and eficient coordination. Consequently, as shown
in Figure 2, supporting the  component with intelligent communication can be a key aspect to
improve the system performance.</p>
        <p>Learning communication protocols is a well-known problem in the literature. Initially,
approaches based on heuristics and meta-heuristics were proposed [56, 57]. However, they often
fail to guarantee scalability and adaptability over time, which are essential qualities for systems
operating in dynamic and large-scale environments. To overcome these limitations, there has
been a significant shift towards utilizing multi-agent reinforcement learning techniques. MARL
has shown promise in addressing the complexities associated with learning and optimizing
communication protocols among multiple agents [58, 59]. In addition to MARL, Graph Neural
Networks have been employed to capture relationships between neighboring agents [60, 61].
Vision we believe the proposed many-agent reinforcement learning approach, combined with
GNNs, holds great promise for learning communication protocols in the context of pulverized
systems. This approach allows for:</p>
        <sec id="sec-4-2-1">
          <title>1. optimizing the amount of information exchanged in the network;</title>
          <p>2. optimizing the set of neighbors to which a message is sent at a given time step;</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>3. optimizing the message exchange frequency.</title>
          <p>Thus, avoiding large bandwidth consumption and energy waste for individual devices.</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Scheduling</title>
        <p>Discussion in the pulverization model, no fixed scheduling policy is generally predefined
for calculating the new state (i.e., executing the behavior). This approach allows developers
considerable flexibility in choosing how and when to execute the primary behavior. The
scheduling policy must be adaptable to the underlying infrastructure layer. For instance, when
computations are entirely ofloaded to the cloud, there are no significant power constraints.
In contrast, when computations are performed on mobile devices like smartphones, power
consumption becomes a critical factor.</p>
        <p>Beyond simple local rules (e.g., reducing evaluation frequency when battery power is low),
scheduling may also be influenced by the collective results of ongoing computations. For
example, in a crowd scenario, if the situation becomes less crowded, the overall frequency
of evaluations may be reduced. Early work in this area, such as programmable distributed
scheduling leveraging aggregate computing [62], involved defining scheduling rules based on
the collective outcome of computations. However, these approaches assumed a fixed scheduling
policy that did not adapt to changing environmental dynamics. Recent studies [23] have shown
that simple variations of Q-learning can improve collective computation by converging more
rapidly to collective structures. However, these methods relied on centralized learning and
Q-tables, making them unsuitable for scaling to the complexity of diverse scenarios.
Vision our vision of embracing many-agent reinforcement learning and graph networks
ofers several advantages:
1. creating distributed global scheduling independent of the number of devices;
2. developing distributed policies influenced by global information or larger neighborhoods;
3. enabling more complex scheduling functions using neural networks instead of simple
tables.</p>
        <p>We believe this is a promising direction, as recent work has successfully applied this approach
to specific task allocation problems [ 63]. Applying it to the pulverization model would enable
its use in various collaborative scenarios, opening the possibility of creating services that
intelligently optimize certain Quality of Service (QoS) metrics based on the collective layer
proposed by the model.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Applicability</title>
        <p>Our approach comprehensively addresses the challenge of deploying and dynamically
reconfiguring applications in the ECC injecting intelligence in these processes (cf. Section 4). Typically,
application scenarios in the context of the IoT, edge computing, and swarm robotics can take
advantages by their deployment in infrastructures like the ECC, improving both functional and
non-functional aspects.</p>
        <p>For instance, in smart city and wearable technology scenarios, our system ofers a dual
benefit: extending battery life and minimizing the overall power consumption of deployed
systems. Intelligent reconfiguration policies can be automatically devised and inferred to better
relocate the component’s execution over the infrastructure, extending the device’s battery life
and reducing the 2 of the system, otherwise dificult or even impossible with traditional
approaches like optimization algorithm or rule-based systems (cf. Section 4.1). Similarly,
determining the optimal execution policy for each component involves a complex interplay
of factors, including the target device, environmental conditions, and specific application
constraints. All these aspects may (and usually do) change frequently over time making it
dificult to predict a priori. In a rescue scenario, our intelligent system can dynamically adapt to
the emergency by increasing the computational frequency of the nearby devices, providing more
updated status about the emergency conditions, while reducing the computational frequency
on the farthest devices not close to the hot spot (cf. Section 4.3). Additionally, via the proposed
approach, intelligent communication patterns can be automatically envisioned determining
where the communication should occur (i.e., the physical place where the communication is
implemented), and who this communication should involve to efectively transfer the minimal
but efective information for the rescuers to promptly intervene in the emergency area (cf.
Section 4.2).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this paper, we discussed a vision for creating and deploying intelligent services within
the Edge-Cloud Continuum (ECC). Specifically, we proposed an idea called intelligent
pulverized systems (IPS), which combines the pulverization model with graph-based networks
and many-agent reinforcement learning. This solution possesses the necessary characteristics
to adapt to modern ECCs, including scalability, the ability to encode collective information,
and the definition of components in an architecture-agnostic manner. We believe that this
approach is essential for maximizing the potential of ECCs to create opportunistic and intelligent
applications.</p>
      <p>
        In the near future, we plan to leverage state-of-the-art many-agent (like mean-field
reinforcement learning [
        <xref ref-type="bibr" rid="ref41">51</xref>
        ]) algorithms to efectively validate this approach in real-world scenarios,
such as smart cities and beyond. By doing so, we aim to demonstrate the practical applicability
and benefits of ISP in dynamically changing environments.
systems, Clust. Comput. 22 (2019) 71–91. URL: https://doi.org/10.1007/s10586-018-2821-8.
doi:10.1007/S10586-018-2821-8.
[14] A. Taherkordi, F. Eliassen, Towards independent in-cloud evolution of cyber-physical
systems, in: 2014 IEEE International Conference on Cyber-Physical Systems, Networks, and
Applications, CPSNA 2014, Hong Kong, China, August 25-26, 2014, IEEE Computer Society,
2014, pp. 19–24. URL: https://doi.org/10.1109/CPSNA.2014.12. doi:10.1109/CPSNA.2014.
12.
[15] W. Z. Khan, E. Ahmed, S. Hakak, I. Yaqoob, A. Ahmed, Edge computing: A survey, Future
Gener. Comput. Syst. 97 (2019) 219–235. URL: https://doi.org/10.1016/j.future.2019.02.050.
doi:10.1016/J.FUTURE.2019.02.050.
[16] R. Newton, M. Welsh, Region streams: functional macroprogramming for sensor networks,
in: A. Labrinidis, S. Madden (Eds.), Proceedings of the 1st Workshop on Data Management
for Sensor Networks, in conjunction with VLDB, DMSN 2004, Toronto, Canada, August 30,
2004, volume 72 of ACM International Conference Proceeding Series, ACM, 2004, pp. 78–87.
      </p>
      <p>URL: https://doi.org/10.1145/1052199.1052213. doi:10.1145/1052199.1052213.
[17] R. Casadei, G. Fortino, D. Pianini, W. Russo, C. Savaglio, M. Viroli, A development approach
for collective opportunistic edge-of-things services, Inf. Sci. 498 (2019) 154–169. URL:
https://doi.org/10.1016/j.ins.2019.05.058. doi:10.1016/J.INS.2019.05.058.
[18] A. Azzara, D. Alessandrelli, S. Bocchino, M. Petracca, P. Pagano, Pyot, a macroprogramming
framework for the internet of things, in: Proceedings of the 9th IEEE International
Symposium on Industrial Embedded Systems, SIES 2014, Pisa, Italy, June 18-20, 2014,
IEEE, 2014, pp. 96–103. URL: https://doi.org/10.1109/SIES.2014.6871193. doi:10.1109/
SIES.2014.6871193.
[19] C. Pinciroli, G. Beltrame, Buzz: A programming language for robot swarms, IEEE Softw.</p>
      <p>33 (2016) 97–100. URL: https://doi.org/10.1109/MS.2016.95. doi:10.1109/MS.2016.95.
[20] J. Beal, D. Pianini, M. Viroli, Aggregate programming for the internet of things, Computer
48 (2015) 22–30. URL: https://doi.org/10.1109/MC.2015.261. doi:10.1109/MC.2015.261.
[21] G. Aguzzi, M. Viroli, L. Esterle, Field-informed reinforcement learning of collective
tasks with graph neural networks, in: IEEE International Conference on Autonomic
Computing and Self-Organizing Systems, ACSOS 2023, Toronto, ON, Canada, September
25-29, 2023, IEEE, 2023, pp. 37–46. URL: https://doi.org/10.1109/ACSOS58161.2023.00021.
doi:10.1109/ACSOS58161.2023.00021.
[22] G. Aguzzi, R. Casadei, M. Viroli, Towards reinforcement learning-based aggregate
computing, in: M. H. ter Beek, M. Sirjani (Eds.), Coordination Models and Languages - 24th
IFIP WG 6.1 International Conference, COORDINATION 2022, Held as Part of the 17th
International Federated Conference on Distributed Computing Techniques, DisCoTec
2022, Lucca, Italy, June 13-17, 2022, Proceedings, volume 13271 of Lecture Notes in
Computer Science, Springer, 2022, pp. 72–91. URL: https://doi.org/10.1007/978-3-031-08143-9_5.
doi:10.1007/978-3-031-08143-9\_5.
[23] G. Aguzzi, R. Casadei, M. Viroli, Addressing collective computations eficiency:
Towards a platform-level reinforcement learning approach, in: R. Casadei, E. D. Nitto,
I. Gerostathopoulos, D. Pianini, I. Dusparic, T. Wood, P. R. Nelson, E. Pournaras, N.
Bencomo, S. Götz, C. Krupitzer, C. Raibulet (Eds.), IEEE International Conference on Autonomic
Computing and Self-Organizing Systems, ACSOS 2022, Virtual, CA, USA, September
19Programming Languages and Operating Systems, Virtual Event, USA, April 19-23, 2021,
ACM, 2021, pp. 135–151. URL: https://doi.org/10.1145/3445814.3446700. doi:10.1145/
3445814.3446700.
[56] C. V. Goldman, J. S. Rosenschein, Emergent coordination through the use of cooperative
state-changing rules, in: B. Hayes-Roth, R. E. Korf (Eds.), Proceedings of the 12th National
Conference on Artificial Intelligence, Seattle, WA, USA, July 31 - August 4, 1994, Volume 1,
AAAI Press / The MIT Press, 1994, pp. 408–413. URL: http://www.aaai.org/Library/AAAI/
1994/aaai94-062.php.
[57] C. V. Goldman, S. Zilberstein, Optimizing information exchange in cooperative
multiagent systems, in: The Second International Joint Conference on Autonomous Agents
&amp; Multiagent Systems, AAMAS 2003, July 14-18, 2003, Melbourne, Victoria, Australia,
Proceedings, ACM, 2003, pp. 137–144. URL: https://doi.org/10.1145/860575.860598. doi:10.
1145/860575.860598.
[58] C. Zhu, M. Dastani, S. Wang, A survey of multi-agent deep reinforcement learning with
communication, Auton. Agents Multi Agent Syst. 38 (2024) 4. URL: https://doi.org/10.1007/
s10458-023-09633-6. doi:10.1007/S10458-023-09633-6.
[59] J. N. Foerster, Y. M. Assael, N. de Freitas, S. Whiteson, Learning to communicate with
deep multi-agent reinforcement learning, in: D. D. Lee, M. Sugiyama, U. von Luxburg,
I. Guyon, R. Garnett (Eds.), Advances in Neural Information Processing Systems 29:
Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016,
Barcelona, Spain, 2016, pp. 2137–2145. URL: https://proceedings.neurips.cc/paper/2016/
hash/c7635bfd99248a2cdef8249ef7bfbef4-Abstract.html.
[60] A. Agarwal, S. Kumar, K. P. Sycara, M. Lewis, Learning transferable cooperative behavior
in multi-agent teams, in: A. E. F. Seghrouchni, G. Sukthankar, B. An, N. Yorke-Smith
(Eds.), Proceedings of the 19th International Conference on Autonomous Agents and
Multiagent Systems, AAMAS ’20, Auckland, New Zealand, May 9-13, 2020, International
Foundation for Autonomous Agents and Multiagent Systems, 2020, pp. 1741–1743. URL:
https://dl.acm.org/doi/10.5555/3398761.3398967. doi:10.5555/3398761.3398967.
[61] J. Jiang, C. Dun, T. Huang, Z. Lu, Graph convolutional reinforcement learning, in:
8th International Conference on Learning Representations, ICLR 2020, Addis Ababa,
Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL: https://openreview.net/forum?id=
HkxdQkSYDB.
[62] D. Pianini, R. Casadei, M. Viroli, S. Mariani, F. Zambonelli, Time-fluid field-based
coordination through programmable distributed schedulers, Log. Methods Comput. Sci. 17 (2021).</p>
      <p>URL: https://doi.org/10.46298/lmcs-17(4:13)2021. doi:10.46298/LMCS-17(4:13)2021.
[63] C. Jian, Z. Pan, L. Bao, M. Zhang, Online-learning task scheduling with gnn-rl scheduler
in collaborative edge computing, Cluster Computing 27 (2023) 589–605. URL: http://dx.
doi.org/10.1007/s10586-022-03957-w. doi:10.1007/s10586-022-03957-w.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Khalyeyev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bures</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hnetynka</surname>
          </string-name>
          ,
          <article-title>Towards characterization of edge-cloud continuum</article-title>
          ,
          <source>CoRR abs/2309</source>
          .05416 (
          <year>2023</year>
          ). URL: https://doi.org/10.48550/arXiv.2309.05416. doi:
          <volume>10</volume>
          .48550/ARXIV.2309.05416. arXiv:
          <volume>2309</volume>
          .
          <fpage>05416</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Moreschini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pecorelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Naz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hästbacka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Taibi</surname>
          </string-name>
          ,
          <article-title>Cloud continuum: The definition</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>131876</fpage>
          -
          <lpage>131886</lpage>
          . URL: https://doi.org/10.1109/ACCESS.
          <year>2022</year>
          .
          <volume>3229185</volume>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2022</year>
          .
          <volume>3229185</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Rosendo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Costan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Valduriez</surname>
          </string-name>
          , G. Antoniu,
          <article-title>Distributed intelligence on the edge-tocloud continuum: A systematic literature review</article-title>
          ,
          <source>J. Parallel Distributed Comput</source>
          .
          <volume>166</volume>
          (
          <year>2022</year>
          )
          <fpage>71</fpage>
          -
          <lpage>94</lpage>
          . URL: https://doi.org/10.1016/j.jpdc.
          <year>2022</year>
          .
          <volume>04</volume>
          .004. doi:
          <volume>10</volume>
          .1016/J.JPDC.
          <year>2022</year>
          .
          <volume>04</volume>
          . 004.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Satyanarayanan</surname>
          </string-name>
          ,
          <article-title>Pervasive computing: vision and challenges</article-title>
          ,
          <source>IEEE Wirel. Commun</source>
          .
          <volume>8</volume>
          (
          <year>2001</year>
          )
          <fpage>10</fpage>
          -
          <lpage>17</lpage>
          . URL: https://doi.org/10.1109/98.943998. doi:
          <volume>10</volume>
          .1109/98.943998.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G. D.</given-names>
            <surname>Abowd</surname>
          </string-name>
          ,
          <article-title>Beyond weiser: From ubiquitous to collective computing</article-title>
          ,
          <source>Computer</source>
          <volume>49</volume>
          (
          <year>2016</year>
          )
          <fpage>17</fpage>
          -
          <lpage>23</lpage>
          . URL: https://doi.org/10.1109/
          <string-name>
            <surname>MC</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <volume>22</volume>
          . doi:
          <volume>10</volume>
          .1109/
          <string-name>
            <surname>MC</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <volume>22</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Friedewald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Raabe</surname>
          </string-name>
          ,
          <article-title>Ubiquitous computing: An overview of technology impacts</article-title>
          ,
          <source>Telematics Informatics</source>
          <volume>28</volume>
          (
          <year>2011</year>
          )
          <fpage>55</fpage>
          -
          <lpage>65</lpage>
          . URL: https://doi.org/10.1016/j.tele.
          <year>2010</year>
          .
          <volume>09</volume>
          .001. doi:
          <volume>10</volume>
          .1016/J.TELE.
          <year>2010</year>
          .
          <volume>09</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Park</surname>
          </string-name>
          , S. Kim,
          <string-name>
            <given-names>Y.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jung</surname>
          </string-name>
          ,
          <article-title>Lired: A light-weight real-time fault detection system for edge computing using LSTM recurrent neural networks</article-title>
          ,
          <source>Sensors</source>
          <volume>18</volume>
          (
          <year>2018</year>
          )
          <article-title>2110</article-title>
          . URL: https://doi.org/10.3390/s18072110. doi:
          <volume>10</volume>
          .3390/S18072110.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <article-title>Deepcrash: A deep learning-based internet of vehicles system for head-on and single-vehicle accident detection with emergency notification</article-title>
          ,
          <source>IEEE Access 7</source>
          (
          <year>2019</year>
          )
          <fpage>148163</fpage>
          -
          <lpage>148175</lpage>
          . URL: https://doi.org/10.1109/ACCESS.
          <year>2019</year>
          .
          <volume>2946468</volume>
          . doi:
          <volume>10</volume>
          .1109/ ACCESS.
          <year>2019</year>
          .
          <volume>2946468</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Many-agent reinforcement learning</article-title>
          ,
          <source>Ph.D. thesis</source>
          , University College London (University of London), UK,
          <year>2021</year>
          . URL: https://ethos.bl.uk/OrderDetails.do?uin=uk.bl.ethos.
          <volume>830054</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Casadei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pianini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Placuzzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viroli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weyns</surname>
          </string-name>
          ,
          <article-title>Pulverization in cyber-physical systems: Engineering the self-organizing logic separated from deployment</article-title>
          ,
          <source>Future Internet</source>
          <volume>12</volume>
          (
          <year>2020</year>
          )
          <article-title>203</article-title>
          . URL: https://doi.org/10.3390/fi12110203. doi:
          <volume>10</volume>
          .3390/FI12110203.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Casadei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Fortino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pianini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Placuzzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Savaglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viroli</surname>
          </string-name>
          ,
          <article-title>A methodology and simulation-based toolchain for estimating deployment performance of smart collective services at the edge</article-title>
          ,
          <source>IEEE Internet Things J</source>
          .
          <volume>9</volume>
          (
          <year>2022</year>
          )
          <fpage>20136</fpage>
          -
          <lpage>20148</lpage>
          . URL: https://doi.org/10. 1109/JIOT.
          <year>2022</year>
          .
          <volume>3172470</volume>
          . doi:
          <volume>10</volume>
          .1109/JIOT.
          <year>2022</year>
          .
          <volume>3172470</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Casadei</surname>
          </string-name>
          ,
          <article-title>Macroprogramming: Concepts, state of the art, and opportunities of macroscopic behaviour modelling</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <volume>275</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>275</lpage>
          :
          <fpage>37</fpage>
          . URL: https://doi.org/10.1145/3579353. doi:
          <volume>10</volume>
          .1145/3579353.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U. S.</given-names>
            <surname>Tim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Gadia</surname>
          </string-name>
          ,
          <source>Cloud and iot-based emerging services 23</source>
          ,
          <year>2022</year>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>20</lpage>
          . URL: https://doi.org/10.1109/ACSOS55765.
          <year>2022</year>
          .
          <volume>00019</volume>
          . doi:
          <volume>10</volume>
          .1109/ACSOS55765.
          <year>2022</year>
          .
          <volume>00019</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>G.</given-names>
            <surname>Ananthanarayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bodík</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chintalapudi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Philipose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ravindranath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sinha</surname>
          </string-name>
          ,
          <article-title>Real-time video analytics: The killer app for edge computing</article-title>
          ,
          <source>Computer</source>
          <volume>50</volume>
          (
          <year>2017</year>
          )
          <fpage>58</fpage>
          -
          <lpage>67</lpage>
          . URL: https://doi.org/10.1109/
          <string-name>
            <surname>MC</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <volume>3641638</volume>
          . doi:
          <volume>10</volume>
          .1109/
          <string-name>
            <surname>MC</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <volume>3641638</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gruteser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Govindan</surname>
          </string-name>
          ,
          <article-title>Real-time trafic estimation at vehicular edge nodes</article-title>
          , in: J. Zhang,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Maggs</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the Second ACM/IEEE Symposium on Edge Computing</source>
          , San Jose / Silicon Valley,
          <string-name>
            <surname>SEC</surname>
          </string-name>
          <year>2017</year>
          , CA, USA, October
          <volume>12</volume>
          -
          <issue>14</issue>
          ,
          <year>2017</year>
          , ACM,
          <year>2017</year>
          , pp.
          <volume>3</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>3</lpage>
          :
          <fpage>13</fpage>
          . URL: https://doi.org/10.1145/3132211.3134461. doi:
          <volume>10</volume>
          .1145/3132211.3134461.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S.</given-names>
            <surname>Tuli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Basumatary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Buyya</surname>
          </string-name>
          , Edgelens:
          <article-title>Deep learning based object detection in integrated iot, fog and cloud computing environments</article-title>
          , CoRR abs/
          <year>1906</year>
          .11056 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1906</year>
          .11056. arXiv:
          <year>1906</year>
          .11056.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Navarro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fernández-Isla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Borraz</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Alonso,</surname>
          </string-name>
          <article-title>A machine learning approach to pedestrian detection for autonomous vehicles using high-definition 3d range data</article-title>
          ,
          <source>Sensors</source>
          <volume>17</volume>
          (
          <year>2017</year>
          )
          <article-title>18</article-title>
          . URL: https://doi.org/10.3390/s17010018. doi:
          <volume>10</volume>
          .3390/S17010018.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xia</surname>
          </string-name>
          , J. Ma, J. Cao,
          <string-name>
            <given-names>S. U.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Zomaya</surname>
          </string-name>
          ,
          <article-title>From insight to impact: Building a sustainable edge computing platform for smart homes</article-title>
          ,
          <source>in: 24th IEEE International Conference on Parallel and Distributed Systems</source>
          , ICPADS 2018, Singapore,
          <source>December 11-13</source>
          ,
          <year>2018</year>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>928</fpage>
          -
          <lpage>936</lpage>
          . URL: https://doi.org/10.1109/PADSW.
          <year>2018</year>
          .
          <volume>8644647</volume>
          . doi:
          <volume>10</volume>
          .1109/PADSW.
          <year>2018</year>
          .
          <volume>8644647</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>I.</given-names>
            <surname>Kovacevic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Harjula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Glisic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lorenzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ylianttila</surname>
          </string-name>
          ,
          <article-title>Cloud and edge computation ofloading for latency limited services</article-title>
          ,
          <source>IEEE Access 9</source>
          (
          <year>2021</year>
          )
          <fpage>55764</fpage>
          -
          <lpage>55776</lpage>
          . URL: https: //doi.org/10.1109/ACCESS.
          <year>2021</year>
          .
          <volume>3071848</volume>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2021</year>
          .
          <volume>3071848</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kwak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. Chong,</surname>
          </string-name>
          <article-title>DREAM: dynamic resource and task allocation for energy minimization in mobile cloud systems</article-title>
          ,
          <source>IEEE J. Sel. Areas Commun</source>
          .
          <volume>33</volume>
          (
          <year>2015</year>
          )
          <fpage>2510</fpage>
          -
          <lpage>2523</lpage>
          . URL: https://doi.org/10.1109/JSAC.
          <year>2015</year>
          .
          <volume>2478718</volume>
          . doi:
          <volume>10</volume>
          .1109/JSAC.
          <year>2015</year>
          .
          <volume>2478718</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zafer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. K.</given-names>
            <surname>Leung</surname>
          </string-name>
          ,
          <article-title>Online placement of multi-component applications in edge computing environments</article-title>
          ,
          <source>IEEE Access 5</source>
          (
          <year>2017</year>
          )
          <fpage>2514</fpage>
          -
          <lpage>2533</lpage>
          . URL: https://doi.org/10. 1109/ACCESS.
          <year>2017</year>
          .
          <volume>2665971</volume>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2017</year>
          .
          <volume>2665971</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>X.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Ant-man: towards agile power management in the microservice era</article-title>
          , in: C.
          <string-name>
            <surname>Cuicchi</surname>
          </string-name>
          , I. Qualters, W. T. Kramer (Eds.),
          <source>Proceedings of the International Conference for High Performance Computing</source>
          ,
          <article-title>Networking, Storage and Analysis</article-title>
          ,
          <source>SC</source>
          <year>2020</year>
          , Virtual Event / Atlanta, Georgia, USA, November 9-
          <issue>19</issue>
          ,
          <year>2020</year>
          , IEEE/ACM,
          <year>2020</year>
          , p.
          <fpage>78</fpage>
          . URL: https://doi.org/10.1109/SC41405.
          <year>2020</year>
          .
          <volume>00082</volume>
          . doi:
          <volume>10</volume>
          .1109/ SC41405.
          <year>2020</year>
          .
          <volume>00082</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hu</surname>
          </string-name>
          , D. Cheng, Y. He,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pancholi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delimitrou</surname>
          </string-name>
          , Seer:
          <article-title>Leveraging big data to navigate the complexity of performance debugging in cloud microservices</article-title>
          , in: I. Bahar,
          <string-name>
            <given-names>M.</given-names>
            <surname>Herlihy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Witchel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Lebeck</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS</source>
          <year>2019</year>
          , Providence, RI, USA, April
          <volume>13</volume>
          -
          <issue>17</issue>
          ,
          <year>2019</year>
          , ACM,
          <year>2019</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>33</lpage>
          . URL: https://doi.org/10.1145/3297858.3304004. doi:
          <volume>10</volume>
          .1145/3297858.3304004.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Kannan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mars</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tang</surname>
          </string-name>
          , Prophet:
          <article-title>Precise qos prediction on non-preemptive accelerators to improve utilization in warehouse-scale computers</article-title>
          , in: Y.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Temam</surname>
          </string-name>
          , J. Carter (Eds.),
          <source>Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS</source>
          <year>2017</year>
          ,
          <article-title>Xi'an, China</article-title>
          , April 8-
          <issue>12</issue>
          ,
          <year>2017</year>
          , ACM,
          <year>2017</year>
          , pp.
          <fpage>17</fpage>
          -
          <lpage>32</lpage>
          . URL: https://doi.org/ 10.1145/3037697.3037700. doi:
          <volume>10</volume>
          .1145/3037697.3037700.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>K.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Adaptive resource eficient microservice deployment in cloud-edge continuum</article-title>
          ,
          <source>IEEE Trans. Parallel Distributed Syst</source>
          .
          <volume>33</volume>
          (
          <year>2022</year>
          )
          <fpage>1825</fpage>
          -
          <lpage>1840</lpage>
          . URL: https://doi.org/10.1109/TPDS.
          <year>2021</year>
          .
          <volume>3128037</volume>
          . doi:
          <volume>10</volume>
          .1109/TPDS.
          <year>2021</year>
          .
          <volume>3128037</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schaul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hessel</surname>
          </string-name>
          , H. van Hasselt,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lanctot</surname>
          </string-name>
          , N. de Freitas,
          <article-title>Dueling network architectures for deep reinforcement learning</article-title>
          , in: M.
          <string-name>
            <surname>Balcan</surname>
            ,
            <given-names>K. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Weinberger</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 33nd International Conference on Machine Learning</source>
          ,
          <string-name>
            <surname>ICML</surname>
          </string-name>
          <year>2016</year>
          , New York City, NY, USA, June 19-24,
          <year>2016</year>
          , volume
          <volume>48</volume>
          <source>of JMLR Workshop and Conference Proceedings, JMLR.org</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1995</fpage>
          -
          <lpage>2003</lpage>
          . URL: http://proceedings.mlr.press/v48/wangf16.html.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>G.</given-names>
            <surname>Nieto</surname>
          </string-name>
          , I. de la Iglesia,
          <string-name>
            <given-names>U.</given-names>
            <surname>Lopez-Novoa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Perfecto</surname>
          </string-name>
          ,
          <article-title>Deep reinforcement learning techniques for dynamic task ofloading in the 5g edge-cloud continuum</article-title>
          ,
          <source>J. Cloud Comput</source>
          .
          <volume>13</volume>
          (
          <year>2024</year>
          )
          <article-title>94</article-title>
          . URL: https://doi.org/10.1186/s13677-024-00658-0. doi:
          <volume>10</volume>
          .1186/ S13677-024-00658-0.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <article-title>Deep reinforcement learning-based task scheduling in iot edge computing</article-title>
          ,
          <source>Sensors</source>
          <volume>21</volume>
          (
          <year>2021</year>
          )
          <article-title>1666</article-title>
          . URL: https://doi.org/10.3390/ s21051666. doi:
          <volume>10</volume>
          .3390/S21051666.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. R.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <article-title>Joint user scheduling and content caching strategy for mobile edge networks using deep reinforcement learning</article-title>
          ,
          <source>in: 2018 IEEE International Conference on Communications Workshops, ICC Workshops</source>
          <year>2018</year>
          ,
          <string-name>
            <given-names>Kansas</given-names>
            <surname>City</surname>
          </string-name>
          ,
          <string-name>
            <surname>MO</surname>
          </string-name>
          , USA, May
          <volume>20</volume>
          -24,
          <year>2018</year>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . URL: https://doi.org/10.1109/ICCW.
          <year>2018</year>
          .
          <volume>8403711</volume>
          . doi:
          <volume>10</volume>
          .1109/ICCW.
          <year>2018</year>
          .
          <volume>8403711</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>D.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Resource management at the network edge: A deep reinforcement learning approach</article-title>
          , IEEE Netw.
          <volume>33</volume>
          (
          <year>2019</year>
          )
          <fpage>26</fpage>
          -
          <lpage>33</lpage>
          . URL: https://doi.org/10. 1109/MNET.
          <year>2019</year>
          .
          <volume>1800386</volume>
          . doi:
          <volume>10</volume>
          .1109/MNET.
          <year>2019</year>
          .
          <volume>1800386</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Long</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>A comprehensive survey on graph neural networks</article-title>
          ,
          <source>IEEE Trans. Neural Networks Learn. Syst</source>
          .
          <volume>32</volume>
          (
          <year>2021</year>
          )
          <fpage>4</fpage>
          -
          <lpage>24</lpage>
          . URL: https: //doi.org/10.1109/TNNLS.
          <year>2020</year>
          .
          <volume>2978386</volume>
          . doi:
          <volume>10</volume>
          .1109/TNNLS.
          <year>2020</year>
          .
          <volume>2978386</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <article-title>Swarm inverse reinforcement learning for biological systems</article-title>
          , in: Y.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Kurgan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>E. R.</given-names>
          </string-name>
          <string-name>
            <surname>Dougherty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kloczkowski</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          (Eds.),
          <source>IEEE International Conference on Bioinformatics and Biomedicine</source>
          ,
          <string-name>
            <surname>BIBM</surname>
          </string-name>
          <year>2021</year>
          , Houston, TX, USA, December 9-
          <issue>12</issue>
          ,
          <year>2021</year>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>274</fpage>
          -
          <lpage>279</lpage>
          . URL: https: //doi.org/10.1109/BIBM52615.
          <year>2021</year>
          .
          <volume>9669656</volume>
          . doi:
          <volume>10</volume>
          .1109/BIBM52615.
          <year>2021</year>
          .
          <volume>9669656</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>E. I.</given-names>
            <surname>Tolstaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Paulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Pappas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <article-title>Learning decentralized controllers for robot swarms with graph neural networks</article-title>
          ,
          <source>in: Proc. of the Conf. on Robot Learning</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>671</fpage>
          -
          <lpage>682</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>E.</given-names>
            <surname>Tolstaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Paulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Pappas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <article-title>Learning decentralized controllers for robot swarms with graph neural networks</article-title>
          ,
          <source>in: Proc. of the Conf. on Robot Learning</source>
          , PMLR,
          <year>2020</year>
          , pp.
          <fpage>671</fpage>
          -
          <lpage>682</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>W.</given-names>
            <surname>Gosrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mayya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Paulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Coverage control in multi-robot systems via graph neural networks</article-title>
          ,
          <source>in: Proc. of the Int. Conf. on Robotics and Automation</source>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>8787</fpage>
          -
          <lpage>8793</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICRA46639.
          <year>2022</year>
          .
          <volume>9811854</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gilmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Schoenholz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Riley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Dahl</surname>
          </string-name>
          ,
          <article-title>Neural message passing for quantum chemistry</article-title>
          ,
          <source>in: Proc. of the 34th International Conference on Machine Learning</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1263</fpage>
          -
          <lpage>1272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          , E. Zaroukian,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dorothy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Basak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Asher</surname>
          </string-name>
          ,
          <article-title>Survey of recent multi-agent reinforcement learning algorithms utilizing centralized training</article-title>
          ,
          <source>arXiv preprint arXiv:2107.14316</source>
          (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2107.14316.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>V.</given-names>
            <surname>Mnih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kavukcuoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Antonoglou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wierstra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Riedmiller</surname>
          </string-name>
          ,
          <article-title>Playing atari with deep reinforcement learning</article-title>
          ,
          <source>CoRR abs/1312</source>
          .5602 (
          <year>2013</year>
          ). URL: http: //arxiv.org/abs/1312.5602. arXiv:
          <volume>1312</volume>
          .
          <fpage>5602</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>J.</given-names>
            <surname>Schulman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wolski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Klimov</surname>
          </string-name>
          ,
          <article-title>Proximal policy optimization algorithms</article-title>
          ,
          <source>CoRR abs/1707</source>
          .06347 (
          <year>2017</year>
          ). URL: http://arxiv.org/abs/1707.06347. arXiv:
          <volume>1707</volume>
          .
          <fpage>06347</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Velu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Vinitsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Bayen</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Wu,</surname>
          </string-name>
          <article-title>The surprising efectiveness of PPO in cooperative multi-agent games</article-title>
          , in: S. Koyejo,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Belgrave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Oh (Eds.),
          <source>Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems</source>
          <year>2022</year>
          , NeurIPS
          <year>2022</year>
          , New Orleans, LA, USA, November 28 - December 9,
          <year>2022</year>
          ,
          <year>2022</year>
          . URL: http://papers.nips.cc/paper_files/paper/2022/hash/ 9c1535a02f0ce079433344e14d910597-Abstract-Datasets_and_Benchmarks.html.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Wang,
          <article-title>Mean field multi-agent reinforcement learning</article-title>
          , in: J. G. Dy,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Krause (Eds.),
          <source>Proceedings of the 35th International Conference on Machine Learning</source>
          ,
          <string-name>
            <surname>ICML</surname>
          </string-name>
          <year>2018</year>
          , Stockholmsmässan, Stockholm, Sweden,
          <source>July 10-15</source>
          ,
          <year>2018</year>
          , volume
          <volume>80</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>5567</fpage>
          -
          <lpage>5576</lpage>
          . URL: http://proceedings.mlr.press/v80/yang18d.html.
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>R.</given-names>
            <surname>Casadei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pianini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viroli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Natali</surname>
          </string-name>
          ,
          <article-title>Self-organising coordination regions: A pattern for edge computing</article-title>
          , in: H.
          <string-name>
            <surname>R. Nielson</surname>
          </string-name>
          , E. Tuosto (Eds.),
          <source>Coordination Models and Languages - 21st IFIP WG 6</source>
          .1 International Conference, COORDINATION 2019,
          <article-title>Held as Part of the 14th International Federated Conference on Distributed Computing Techniques</article-title>
          ,
          <source>DisCoTec</source>
          <year>2019</year>
          ,
          <string-name>
            <given-names>Kongens</given-names>
            <surname>Lyngby</surname>
          </string-name>
          , Denmark, June 17-21,
          <year>2019</year>
          , Proceedings, volume
          <volume>11533</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2019</year>
          , pp.
          <fpage>182</fpage>
          -
          <lpage>199</lpage>
          . URL: https: //doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -22397-7_
          <fpage>11</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -22397-7\_
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>K.</given-names>
            <surname>Khetarpal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riemer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Rish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Precup</surname>
          </string-name>
          ,
          <article-title>Towards continual reinforcement learning: A review and perspectives</article-title>
          ,
          <source>J. Artif. Intell. Res</source>
          .
          <volume>75</volume>
          (
          <year>2022</year>
          )
          <fpage>1401</fpage>
          -
          <lpage>1476</lpage>
          . URL: https://doi.org/10. 1613/jair.1.13673. doi:
          <volume>10</volume>
          .1613/JAIR.1.13673.
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Suh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delimitrou</surname>
          </string-name>
          ,
          <article-title>Sinan: Ml-based and qos-aware resource management for cloud microservices</article-title>
          , in: T.
          <string-name>
            <surname>Sherwood</surname>
            ,
            <given-names>E. D.</given-names>
          </string-name>
          <string-name>
            <surname>Berger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Kozyrakis (Eds.),
          <source>ASPLOS '21: 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems</source>
          , Virtual Event, USA, April
          <volume>19</volume>
          -
          <issue>23</issue>
          ,
          <year>2021</year>
          , ACM,
          <year>2021</year>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>181</lpage>
          . URL: https://doi.org/10.1145/3445814.3446693. doi:
          <volume>10</volume>
          .1145/ 3445814.3446693.
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delimitrou</surname>
          </string-name>
          ,
          <article-title>Sage: practical and scalable ml-driven performance debugging in microservices</article-title>
          , in: T.
          <string-name>
            <surname>Sherwood</surname>
            ,
            <given-names>E. D.</given-names>
          </string-name>
          <string-name>
            <surname>Berger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Kozyrakis (Eds.),
          <source>ASPLOS '21: 26th ACM International Conference on Architectural Support for</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>