<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Axes of Characterizing Generative Systems: A Taxonomy of Approaches to Expressive Range Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Josiah Boucher</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gillian Smith</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yunus Doğan Telliel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Worcester Polytechnic Institute (WPI))</institution>
          ,
          <addr-line>100 Institute Road, Worcester MA</addr-line>
          ,
          <country country="US">U.S.A</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Seeking to leverage Expressive Range Analysis (ERA)- a method of characterizing generative systems- for analysis of generative AI (GAI) systems and their outputs, this paper categorizes the approaches that have been used for ERA. We present a taxonomy with five axes of characterization that may be applied to ERA methodologies: content agnostic vs semantic; quantitative vs qualitative; product vs process; objective vs subjective; and automated vs manual. While ERA has traditionally been limited to the domain of Procedural Content Generation (PCG) in video game development, we recognize parallels between PCG and GAI and hope to see an expansion of the application of ERA into the domain of GAI. Serving this goal, these axes provide metrics through which to categorize, compare, and explore approaches.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Expressive Range Analysis</kwd>
        <kwd>Procedural Content Generation</kwd>
        <kwd>Generative AI</kwd>
        <kwd>Evaluating Generative Systems</kwd>
        <kwd>Taxonomy</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Expressive Range Analysis (ERA) is a method of
characterizing the nature and shape of a generative model in terms
of its outputs. What sorts of outputs is a generator capable
of? What impact does changing inputs of a generator have
on its outputs? How do biases manifest in the generative
outputs, and how do the inputs influence these biases? ERA
is well suited for answering these questions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. ERA comes
from the domain of Procedural Content Generation (PCG)
in video game development: the practice of leveraging
algorithmic methods to design and produce artifacts for use in
games across a variety of contexts, such as levels or maps
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        We recognize common threads between PCG and
generative AI (GAI) systems- defined here as the process of using
generative technologies such as large language models to
produce content, often using natural language prompts as
human-provided input- including high-level functionality,
motivations, use cases, and shortcomings of both domains
(for details, see Section 2). Because of these similarities,
we seek to expand and leverage ERA—which has not only
proven useful for recognizing bias within, and categorizing
generative outputs of PCG systems, but also resembles some
existing approaches for GAI analysis [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]—for analysis of GAI
system outputs, especially text-to-X, large language model
(LLM)-based applications (such as the text-to-text Chat GPT
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], text-to-image Stable Difusion [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and text-to-speech
ElevenLabs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
      </p>
      <p>
        The application of ERA in the domain of GAI is not a
oneto-one translation from how the method is used in PCG. Use
of GAI tools commonly requires more complicated, varied,
and linguistic inputs than PCG, which tends to operate from
numerical randomization as a starting point for much of its
generation. Because these inputs have a major impact on
the outputs of GAI systems [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], selection of these inputs
is an important consideration for applying ERA to GAI—
11th Experimental Artificial Intelligence in Games Workshop, November
19, 2024, Lexington, Kentucky, USA.
* Corresponding author.
† These authors contributed equally.
$ jdboucher@wpi.edu (J. Boucher); gmsmith@wpi.edu (G. Smith);
ydtelliel@wpi.edu (Y. D. Telliel)
      </p>
      <p>0009-0002-6865-9031 (J. Boucher); 0000-0002-2765-7702 (G. Smith);
0000-0001-5651-7349 (Y. D. Telliel)
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License
Attribution 4.0 International (CC BY 4.0).
e.g. some decision points for input selection could include
whether to handcraft inputs or sample the latent space, and
whether the inputs should be restricted in a way that targets
a specific domain of outputs (i.e. a prompt to output a
haiku would be very diferent from one that might produce a
cooking recipe). Furthermore, ERA requires the application
of metrics to large quantities of content outputted from the
system in question. What metrics are used and how they
are determined is a major component of ERA, often directly
determining the value provided by the analysis. Because
GAI outputs tend to occupy a broad domain of applications
and mediums, unique challenges arise when considering
these metrics, further complicating the application of ERA
in this context.</p>
      <p>
        Responding to 1) these parallels between PCG and GAI,
and 2) the massive increase of scope presented by GAI, this
paper presents a taxonomy for categorizing approaches to
ERA. This taxonomy includes vocabulary to better
distinguish examples from existing work in PCG, as well as a
compass to guide further exploration of using this method
in the rapidly expanding domain of GAI. The goal of this
taxonomy is to push the boundaries of what ERA may be
used for and how it may be applied.
This work operates in the intersection of procedural content
generation and generative AI. Guzdial provides a valuable
bridge between these domains with the lens of human-AI
interactive generation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. While Guzdial uses this lens to
frame PCG and GAI as essentially the same process, we view
PCG as a broader term describing the process of content
generation at its highest level, and GAI as a more descriptive
term identifying a subset of generative practices that use
specific technologies. Framing GAI as a subset of PCG—or
using Guzdial’s lens of human-AI interactive generation
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]—broadens the scope of understanding for both domains
and allows the application of evaluation methods of PCG
for analysis of GAI systems.
      </p>
      <p>
        Particularly, we are interested in leveraging expressive
range analysis for evaluation of GAI systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Withington et al. list expressive range analysis as one of 12 "features"
used for comparison of PCG systems [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. We consider ERA
a promising method of GAI analysis compared to alternative
PCG-evaluation practices due to its flexibility in applying
a broad range of analysis metrics to its evaluated systems,
particularly highlighting its strength of identifying bias and
system tendencies- including as influenced by user input
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Withington et al. identifies weaknesses with current
evaluation methods for PCG systems, suggesting they could
be mitigated by diverse research frameworks and promoting
the reuse of methodology where possible [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This paper
seeks to answer this call by broadening the applicability of
ERA.
      </p>
      <p>
        We seek to do so by presenting a taxonomic framework
for categorizing approaches used for ERA, drawing from
similar work in PCG such as Togelius et. al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and Smith’s
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] taxonomies of PCG, as well as Withington et al.’s modern
taxonomy of available evaluation approaches [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        Further outlining the connection between GAI and PCG,
we consider the motivations and use-cases for such
systems. PCG and GAI are both used to automatically generate
large amounts of content, and to increase variety of content
[
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. PCG is also used for assistive tools, helping with
tasks such as level creation [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and support tools have
emerged to make these systems easier to understand [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
LLM applications like ChatGPT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and Stable Difusion
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] are promoted for their potential to reduce labor costs,
enable new business models, and increase access to content
creation– Cook claims all of these as common motivations
for studying AI in games [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. We also acknowledge the
use of machine learning in PCG [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ] as a point that
highlights the connection between PCG and GAI.
      </p>
      <p>
        We also recognize parallel claims of GAI and PCG
capabilities—and shortcomings thereof—that may provide
useful points of interest for continuing this investigation.
Developing a useful PCG generator often takes just as much,
or more, efort than hand-crafted alternatives [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], despite
the allure of rapid content production. Furthermore, the
variety content generators are able to produce can be
limited, as players are often capable of identifying patterns
between generated content [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. While Generative AI
offers potential to expand the type of content that can be
generated—such as more complex generative audio and
music used to increase variety and interactivity of game music
compared to non-generative methods [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], as well as the
range of users that can make use of generators—we also
view the drawbacks of PCG as rich sites of investigation in
GAI. ERA has proven a useful method for identifying such
weaknesses in the domain of PCG, and we hope to see its
efectiveness applied to GAI as well.
      </p>
      <p>
        GAI technologies have prompted significant concern—i.e.
regarding its social and environmental impact, underlying
politics of AI design, implementation, and advocacy [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
Additional concerns include embedded systemic biases, risk
of plagiarism and misinformation, and human and
environmental costs [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Jiang et al. find that artists, in particular,
identify harms to professional reputation, intellectual
property, and financial risk [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Because of GAI’s recognized
potential for harm, we value tools and methods for
understanding and analyzing these systems and their outputs.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3. The Problem Space</title>
      <p>
        Summarizing the process of Expressive Range Analysis, this
form of inquiry operates in three steps [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]:
• 1) Start with a set of images or similar data (i.e.
levels, maps, etc.) produced by the generator being
analyzed.
• 2) Determine useful metrics for the goal of analysis
(i.e. linearity or leniency of a level) and apply them
to the data.
      </p>
      <p>• 3) Produce a visualization of the metered data space.</p>
      <p>
        Two challenges arise when considering this method for
GAI-generators as opposed to PCG-generators. First,
producing a set of data to analyze tends to be more complicated.
GAI text-to-X applications require linguistic input, leading
to larger variance of the input space in terms of both
quantity and meaning. Further complicating the issue, PCG tools
are typically custom-made for specific use-cases [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], where
GAI tools (such as ChatGPT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] or Google’s Gemini [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ])
are commonly presented as general-purpose—while random
inputs could be provided in place of handcrafted ones or the
latent space could be randomly sampled, these approaches
would provide undesirable outputs, lacking the targeted
focus of a particular use-case. The general-purpose nature of
GAI systems adds noise to the output space when looking
at these tools for specific use-cases, since outputs that don’t
ift the specific use-case become irrelevant. Guzdial
summarizes this aspect as GAI tool developers seeking to expand
the possible valid outputs for a particular tool to include all
possible valid outputs for every tool [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This noise presents
additional challenges: how do you narrow the scope of the
output to only match the relevant use-case? What prompts
do you provide as an input to produce a suitable set of data
to categorize the generator as a whole? Guzdial’s
humancentered input alignment [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] may prove a useful attribute
of consideration to respond to this challenge.
      </p>
      <p>
        The second challenge that arises from considering ERA
for use in analyzing GAI systems comes from the metrics of
analysis. ERA can be used alongside any metrics, depending
on what a generative system is being analyzed for. These
metrics must be applied to large quantities of data, in some
cases encompassing the entire generative output of a system
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Furthermore, many examples of ERA employ automatic
processing of data, but not all data and research goals
beneift from automatable metrics- which tend to be quantitative
in nature. The challenge here is twofold: how do you
determine semantically meaningful metrics for increasingly
variable data? And, accounting for practical workload and
feasibility, how do you apply these metrics to increasingly
multitudinous data?
      </p>
      <p>
        We treat producing an intended result as the motivation
behind creating and using generators. While some
generators ofer promises of increased speed and eficiency, those
claims are not always accurate [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]—despite this,
generators are valued by many as useful tools. We recognize the
purpose of ERA as evaluating the success of that motivation.
Using ERA is akin to asking: has this this generator
successfully captured what it was designed to capture? With
this framing in mind, our motivation for this paper is to
facilitate the answering of that question in a greater variety
of contexts. The challenge in the case of GAI is, in short:
how?
      </p>
    </sec>
    <sec id="sec-3">
      <title>4. Framework</title>
      <p>Here, we present a taxonomy of ERA methods, seeking to
provide language and framing to better address the
challenges of applying ERA to a greater variety of contexts—
especially GAI systems. We have defined five axes for
characterizing ERA methodologies: quantitative vs qualitative,
product vs process, objective vs subjective, content agnostic
vs semantic, and automated vs manual.</p>
      <p>
        We considered three factors as a basis for defining these
axes: the origin and definition of ERA, the adopted practice
of this method, and its potential for further expansion and
refinement. We considered the origin of ERA [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], identifying
decision points in how the method may be applied. We
also considered how ERA has been adopted in research
endeavors and sought to challenge assumptions that have
become commonplace. Finally, this taxonomy is also a result
of our eforts to address challenges we faced in applying
ERA to GAI. We hope to further refine these axes and their
definitions in future work.
      </p>
      <sec id="sec-3-1">
        <title>4.1. Quantitative vs Qualitative</title>
        <p>
          The metrics that are applied to the generator-produced
data may be quantitative, qualitative, or varying degrees of
mixed-method. Many traditional ERA approaches utilize
quantitative metrics, as this approach is often more suitable
for automatically applying metrics to data and producing
a descriptive visualization. Smith and Whitehead use two
metrics for applying ERA to Launchpad- linearity and
leniency [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Both of these metrics are quantitative, because
they are described, measured, and depicted numerically.
Interestingly, Kreminski et al. [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] identify determining
quantitative metrics as an essential step of ERA, which is
consistent with the examples of the method’s initial
presentation [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]- it is therefore unsurprising that Qualitative
metrics are relatively under-utilized in ERA. This
taxonomy challenges this assumption—which is present across
many ERA-focused research endeavors—instead suggesting
an expansion of valid data collection eforts.
        </p>
        <p>While there is not a strong foundation of qualitative
metrics applied to ERA in PCG, some hypothetical qualitative
metrics could come from methods like user surveys or
interviews, such as those found in play-testing eforts to evaluate
elements of a game like challenge or engagement. These
metrics could be visually represented with tools such as
word clouds, afinity charts, or heat-maps that highlight the
frequency of thematic elements within the data, for example.</p>
      </sec>
      <sec id="sec-3-2">
        <title>4.2. Product vs Process</title>
        <p>
          Because our goal is to evaluate if a generator is successful
in capturing what it was designed to do, it is interesting to
consider both the product of a generator and the process
of producing that output. The product of a generator is its
output- an image, the answer to a question, or a video game
level are all examples of this. Analyzing the product is useful
for identifying what a generator is capable of- in terms of
quality, variety, etc.- and what sort of biases are present in
the outputs. Smith and Whitehead [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] present a
productfocused approach, as their linearity and leniency metrics
consider only the levels produced by the generative system.
Withington’s [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] approach is also product-focused, since
this work considers the diferences between outputs of a
generator rather than considering the process of producing
those outputs.
        </p>
        <p>
          The process entails the experience of producing that
output. Common narratives of generative systems sell them as
faster and more eficient than manual alternatives.
Looking at the process allows us to evaluate those claims, using
metrics such as the time it takes to get a desirable output
or how many iterations of prompts/generations it takes to
produce such an output. Kreminski et al. [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] arguably
present a process-focused approach, because their
motivation for analysis seeks to evaluate the usability experience
of a generative system. However, the metrics used are
primarily focused on the product of the generator. A more
process-centric example would include metrics that
evaluate procedural aspects such as the time it takes to generate
an output, or the number of generative attempts before a
user finds a suitable option.
        </p>
        <p>
          Shaker et. al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] highlight generator reliability as one
important piece of PCG evaluation. Looking at this aspect
provides a mixed-method approach that tends to use product
to reveal something about the process.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>4.3. Objective vs Subjective</title>
        <p>
          This axis is concerned with the nature of the analysis
metric—is a given metric verifiable based on factual
evidence, or does it vary with perspective, based in emotion or
opinion? Smith and Whitehead’s two metrics for Launchpad
again provide a useful example: Linearity, as described in
their paper, is an objective metric that describes the factual
"profile" of platforming levels (how well the geometry of
the level fits a straight line) in the evaluated system [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
Leniency, however, is identified as a subjective score based
on an intuitive sense of how lenient components of a level
are towards a player [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Subjective metrics are relatively
underrepresented in PCG, but may prove useful for
evaluating GAI systems– especially as they are applied in creative
domains. We consider user experience analysis a useful
point of reference for how subjective metrics may be used
in data analysis and visualization [
          <xref ref-type="bibr" rid="ref27 ref28">27, 28</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>4.4. Content Agnostic vs Semantic</title>
        <p>
          This axis is concerned with the data itself and the value
sought from analysis. Semantic approaches seek to find
meaning within the context of the inputs and outputs of the
generator. Smith and Whitehead [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] present a semantic
approach, as their metrics are intended to allow comparison of
the generated content. Lucas and Volz [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] provide another
example of a semantic approach, as theirs is also intended
to compare generated content.
        </p>
        <p>
          Content agnostic approaches seek to find meaning in
the generator, regardless of its inputs or outputs- though
the inputs and outputs may be useful for analysis. While
Kreminski et al. [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] evaluate the product of the
generator, their interests lie in the users of the generative
systemanswering questions such as how thoroughly they explore
the generative range- rather than finding meaning from
the generative output itself, making their approach
contentagnostic. Withington’s exploration of quality-diversity
algorithms [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] presents another content-agnostic example
because the focus is not on the details of specific outputs,
but rather measuring the diferences between them.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>4.5. Automated vs Manual</title>
        <p>
          This axis is concerned with the process of applying metrics
to the data. Automated processes are typically conducted
computationally, making them especially suitable for
quantitative analysis metrics and often desirable for the promise
of reduced processing time or efort. Smith and Whitehead
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], Kreminski et al. [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], and Kybartas et al. [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ], among
others, all present automated approaches for applying ERA
metrics to their data.
        </p>
        <p>
          Manual processes require human labor for applying
analysis metrics to each piece of data—while this tends to be
more time-consuming, it also allows closer scrutiny of
elements that are dificult to capture without direct human
intervention. Manual processes are often unsuitable for
the large quantities of data that ERA typically processes,
but methods such as crowd sourced photogrammetry—e.g.
as used for identifying information about wildlife
populations [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]—may hypothetically be leveraged to manually
process such data. Human subject experiments, such as
those commonly used in narrative generation projects [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ],
are another example of a manual approach.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Discussion</title>
      <p>
        Framing GAI under the lens of PCG- or both as the same
process, e.g. through the lens of human-AI interaction
generation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]- incorporates this new technology into an
established domain of research that has ERA as a design-focused
method for evaluation. In connecting GAI and PCG, though,
we have identified a need for both an improved vocabulary
for describing ERA and an expanded potential scope for ERA
that challenges existing methods. Thus this taxonomy
further expands the potential range of analytical applications
for ERA, providing language to better describe and imagine
its use-cases and identify its historical gaps. This expansion
responds to existing weak-points in PCG evaluation—such
as those identified by Withington et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]—by increasing
the diversity and re-usability of ERA as a framework for
generative analysis.
      </p>
      <p>
        Further, this taxonomy allows for better description of the
lfexibility and space occupied by broadly applicable aspects
of evaluation, such as those presented in Guzdial’s
humanAI interaction generation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. For example, Guzdial’s call for
human-centered input alignment considers the relationship
between valid system inputs and user-preferences
regarding those inputs. Using the vocabulary from our taxonomy,
this is a process-focused, content-agnostic, subjective
approach, because meaning is found according to individual
perceptions without concern for the generative output. Such
a metric could be applied to data quantitatively or
qualitatively, using either an automated or manual approach.
Guzdial’s adaptability similarly considers the process of
generation and user perceptions [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and has similar axis
placement to human-centered input alignment– though it
is objective rather than subjective, since adaptability was
a predetermined aspect of the generative process, rather
than a variable expressed by human interpretation. Novelty,
however, is an objective, product-focused metric because it
is based in the observed possible generative outputs.
      </p>
      <p>It is our hope that using the vocabulary of this taxonomy
can provide some clarity to the research community on
how they are using ERA for evaluation, as well as identify
potential new approaches for ERA.</p>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusions, Limitations, &amp;</title>
    </sec>
    <sec id="sec-6">
      <title>Future Work</title>
      <p>This paper explores an avenue for exploring the
intersection of PCG and GAI, using expressive range analysis as a
common method for analyzing generative systems and their
outputs. We present a taxonomic framework for
categorizing existing and imagined ERA inquiries, hoping to allow
more efective navigation of this space and leverage PCG
tools for study of GAI technologies and systems.</p>
      <p>This taxonomy opens interesting possibilities for future
applications of ERA. What do qualitative applications of
ERA look like? How would manually applying metrics to
data compare to more commonly applied automated
processes? These are relatively underexplored areas of ERA,
and the possibility of applying this method to generative
systems makes these questions more compelling.</p>
      <p>The scope of this paper is limited to theory. This thread
of research would benefit from a more complete, systematic
review of ERA projects that places existing work on the
axes of this taxonomy and further inform the chosen axes.
As an extension of this research, we also see value in
performing ERA on GAI tools using diferent combinations of
axis placement—especially including qualitative and manual
approaches.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This material is based upon work partially supported by the
National Science Foundation (NSF) under Grant No
DGE1922761. Any opinions, findings, and conclusions or
recommendations expressed in this material are those of the
authors and do not necessarily reflect the views of the NSF.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Whitehead</surname>
          </string-name>
          ,
          <article-title>Analyzing the expressive range of a level generator</article-title>
          ,
          <source>in: Proceedings of the 2010 workshop on procedural content generation in games, 2010</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Smith,</surname>
          </string-name>
          <article-title>Understanding procedural content generation: a design-centric analysis of the role of pcg in games</article-title>
          ,
          <source>in: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>917</fpage>
          -
          <lpage>926</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Deckers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <article-title>Manipulating embeddings of stable difusion prompts</article-title>
          ,
          <source>arXiv preprint arXiv:2308.12059</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] OpenAI, Chatgpt, https://chat.openai.com/,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. AI</given-names>
            ,
            <surname>Stable</surname>
          </string-name>
          <string-name>
            <surname>difusion</surname>
          </string-name>
          , https://stability.ai/stable-image,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6] ElevenLabs, Elevenlabs, https://elevenlabs.io/,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Korzynski</surname>
          </string-name>
          , G. Mazurek,
          <string-name>
            <given-names>P.</given-names>
            <surname>Krzypkowska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kurasinski</surname>
          </string-name>
          ,
          <article-title>Artificial intelligence prompt engineering as a new digital competence: Analysis of generative ai technologies such as chatgpt</article-title>
          ,
          <source>Entrepreneurial Business and Economics Review</source>
          <volume>11</volume>
          (
          <year>2023</year>
          )
          <fpage>25</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Guzdial</surname>
          </string-name>
          ,
          <article-title>Human-AI interaction generation: A connective lens for generative AI and procedural content generation</article-title>
          ,
          <source>in: Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI-24)</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>O.</given-names>
            <surname>Withington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cook</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tokarchuk</surname>
          </string-name>
          ,
          <article-title>On the evaluation of procedural level generation systems</article-title>
          ,
          <source>in: Proceedings of the 19th International Conference on the Foundations of Digital Games</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Togelius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. N.</given-names>
            <surname>Yannakakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. O.</given-names>
            <surname>Stanley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Browne</surname>
          </string-name>
          ,
          <article-title>Search-based procedural content generation: A taxonomy and survey</article-title>
          ,
          <source>IEEE Transactions on Computational Intelligence and AI in Games</source>
          <volume>3</volume>
          (
          <year>2011</year>
          )
          <fpage>172</fpage>
          -
          <lpage>186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Shaker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Togelius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <article-title>Procedural content generation in games (</article-title>
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hendrikx</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Meijer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Van Der</given-names>
            <surname>Velden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Iosup</surname>
          </string-name>
          ,
          <article-title>Procedural content generation for games: A survey, ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 9 (</article-title>
          <year>2013</year>
          )
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Liapis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. N.</given-names>
            <surname>Yannakakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Togelius</surname>
          </string-name>
          ,
          <article-title>Sentient sketchbook: computer-assisted game level authoring (</article-title>
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cook</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Colton</surname>
          </string-name>
          , Danesh:
          <article-title>Interactive tools for understanding procedural content generators</article-title>
          ,
          <source>IEEE Transactions on Games</source>
          <volume>14</volume>
          (
          <year>2021</year>
          )
          <fpage>329</fpage>
          -
          <lpage>338</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cook</surname>
          </string-name>
          , Optimists at heart:
          <article-title>Why do we research game ai?</article-title>
          ,
          <source>in: 2022 IEEE Conference on Games (CoG)</source>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>560</fpage>
          -
          <lpage>567</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Summerville</surname>
          </string-name>
          ,
          <article-title>Expanding expressive range: Evaluation methodologies for procedural content generation</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment</source>
          , volume
          <volume>14</volume>
          ,
          <year>2018</year>
          , pp.
          <fpage>116</fpage>
          -
          <lpage>122</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Summerville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Snodgrass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guzdial</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Holmgård</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Hoover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Isaksen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nealen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Togelius</surname>
          </string-name>
          ,
          <article-title>Procedural content generation via machine learning (pcgml)</article-title>
          ,
          <source>IEEE Transactions on Games</source>
          <volume>10</volume>
          (
          <year>2018</year>
          )
          <fpage>257</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Short</surname>
          </string-name>
          , T. Adams,
          <article-title>Procedural generation in game design</article-title>
          , CRC Press,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>E.</given-names>
            <surname>Short</surname>
          </string-name>
          , Bowls of Oatmeal and Text Generation, https://emshort.blog/
          <year>2016</year>
          /09/21/ bowls-of-oatmeal-and
          <string-name>
            <surname>-</surname>
          </string-name>
          text-generation/,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Plut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pasquier</surname>
          </string-name>
          ,
          <article-title>Generative music in video games: State of the art, challenges, and prospects</article-title>
          ,
          <source>Entertainment Computing</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>100337</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Bender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gebru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McMillan-Major</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shmitchell</surname>
          </string-name>
          ,
          <article-title>On the dangers of stochastic parrots: Can language models be too big?</article-title>
          ,
          <source>in: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency</source>
          , Association for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          , pp.
          <fpage>610</fpage>
          -
          <lpage>623</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <article-title>Generative ai considered harmful</article-title>
          ,
          <source>in: Proceedings of the 5th International Conference on Conversational User Interfaces</source>
          ,
          <source>Association for Computing Machinery</source>
          , New York, NY, USA,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>H. H.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Brown</surname>
          </string-name>
          , J. Cheng, M. Khan,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Workman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Flowers</surname>
          </string-name>
          , T. Gebru,
          <article-title>Ai art and its impact on artists</article-title>
          ,
          <source>in: Proceedings of the 2023 AAAI/ACM Conference on AI</source>
          ,
          <string-name>
            <surname>Ethics</surname>
          </string-name>
          , and Society, Association for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          , pp.
          <fpage>363</fpage>
          -
          <lpage>374</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>G. AI</surname>
          </string-name>
          , Gemini, https://gemini.google.com/,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kreminski</surname>
          </string-name>
          , I. Karth,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mateas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wardrip-Fruin</surname>
          </string-name>
          ,
          <article-title>Evaluating mixed-initiative creative interfaces via expressive range coverage analysis</article-title>
          .,
          <source>in: IUI Workshops</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>34</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>O.</given-names>
            <surname>Withington</surname>
          </string-name>
          ,
          <article-title>Illuminating super mario bros: qualitydiversity within platformer level generation</article-title>
          ,
          <source>in: Proceedings of the 2020 Genetic and Evolutionary Computation Conference Companion</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>223</fpage>
          -
          <lpage>224</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bangor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. T.</given-names>
            <surname>Kortum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>An empirical evaluation of the system usability scale, Intl</article-title>
          .
          <source>Journal of Human-Computer Interaction</source>
          <volume>24</volume>
          (
          <year>2008</year>
          )
          <fpage>574</fpage>
          -
          <lpage>594</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Djamasbi</surname>
          </string-name>
          ,
          <article-title>Eye tracking and web experience, AIS Transactions on Human-Computer Interaction 6 (</article-title>
          <year>2014</year>
          )
          <fpage>37</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Lucas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Volz</surname>
          </string-name>
          ,
          <article-title>Tile pattern kl-divergence for analysing and evolving game levels</article-title>
          ,
          <source>in: Proceedings of the Genetic and Evolutionary Computation Conference</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>170</fpage>
          -
          <lpage>178</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Kybartas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Verbrugge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lessard</surname>
          </string-name>
          ,
          <article-title>Tension space analysis for emergent narrative</article-title>
          ,
          <source>IEEE Transactions on Games</source>
          <volume>13</volume>
          (
          <year>2020</year>
          )
          <fpage>146</fpage>
          -
          <lpage>159</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Wood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. W.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Beltran</surname>
          </string-name>
          ,
          <article-title>Accuracy and precision of citizen scientist animal counts from drone imagery</article-title>
          ,
          <source>PloS one 16</source>
          (
          <year>2021</year>
          )
          <article-title>e0244040</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sanghrajka</surname>
          </string-name>
          , E. Lang,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <article-title>Generating quest representations for narrative plans consisting of failed actions</article-title>
          ,
          <source>in: Proceedings of the 16th International Conference on the Foundations of Digital Games</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>