<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Collaborative Comic Generation: Integrating Visual Narrative Theories with AI Models for Enhanced Creativity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yi-ChunChen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ArnavJhala</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Generative AI System, Human-AI interaction, Comic Generation, Visual Narrative Theories</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>North Carolina State University</institution>
          ,
          <addr-line>Computer Science, Raleigh, NC 27606</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This study presents a theory-inspired visual narrative generative system that integrates conceptual principles-comic authoring idioms-with generative and language models to enhance the comic creation process. Our system combines human creativity with AI models to support parts of the generative process, providing a collaborative platform for creating comic content. These comic-authoring idioms, derived from prior human-created image sequences, serve as guidelines for crafting and refining storytelling. The system translates these principles into system layers that facilitate the creation of comics through sequential decision-making, addressing narrative elements such as panel composition, story tension changes, and panel transitions. Key contributions include the integration of machine learning models into the human-AI cooperative comic generation process, the deployment of abstract narrative theories into AI-driven comic creation, and a customizable tool for narrative-driven image sequences. This approach improves narrative elements in generated image sequences and brings engagement of human creativity in an AI-generative process of comics. We open-source the code at https://github.com/RimiChen/Collaborative_Comic_Generation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Despite their various names, comics, manga, and visual stories represent a dominant form of
storytelling that spans cultures and age groups. Authors combine their creative storytelling ideas
with textual expressions and graphical representations to convey intricate narratives through
multi-modal panel sequences. Recently, generative AI, a trending topic in AI creativity, has
explored the automatic generation of narratives and their synthesis with visual representations,
simulating an activity traditionally rooted in human creativity. These studies lead to an exciting
research question: how can AI collaborate with human creativity to create image sequences,
such as comics and visual stories?</p>
      <p>Nowadays, most generative AI models handle almost the entire process of creating image
sequences, leaving little room for human authors to modify the details during generation.
Traditionally, however, creating visual stories or comics relies heavily on authors’ familiarity
with narrative idiom1s],[principles of visual storytelli2n],ga[nd their skills in translating
narrative content into visual representa3t].ions [</p>
      <p>
        We propose a solution that balances these approaches, allowing AI models to assist in
generating visual stories while also providing space for human creativity. Narrative theories
and idioms are important tools for authors, helping them convey stories with clarity and making
visual narratives more engaging for readers. These elements, derived from existing comics, serve
as underlying schemes. We aim to leverage these narrative elements by integrating them into
the generative process, enhancing the results. Current studies on generating image sequences
or visual storytelling include text-to-image syn4t]h, ecshisar[acter consistency algorithms
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and narrative structuring mod6e,l7s].[Although most synthesis methods allow users to
provide some input as seeds, such as captions or base images, authors have limited flexibility in
the creative process. Furthermore, while narrative idioms like grammar and patterns of visual
storytelling are widely discussed in the analysis of image sequences, they are rarely incorporated
into AI-driven comic generation. Even in studies that integrated rule-based methodologies
to allow narrative theories to influence visual storyt8,el9l,i1n0g, 1[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], authors’ creativity
remained largely excluded from the generative process.
      </p>
      <p>Building on prior studies, we propose a comic-generating system that integrates a
humanin-the-loop process, enabling authors to customize the generation by leveraging multiple AI
models. Our system combines AI-driven narrative development and art generation, reducing
the efort needed to create dynamic visuals. A sentiment analysis model also guides the story
arc, serving as a plotline framework.</p>
      <p>Second, our system dynamically applies narrative idioms during the generation process.
The layer-wised customizations allow authors to edit the image sequence iteratively, selecting
diferent rules to refine the results. This human-in-the-loop process ensures that authors can
adjust based on the generated image sequences. Third, rather than treating the comic panel as
a single entity, we decompose it into multiple layers: background, foreground, compositional
layer, and symbol layer. This approach enables authors to customize specific details without
altering the entire panel.</p>
      <p>Finally, our system provides a Graphical User Interface (GUI) and an Application Programming
Interface (API). These interfaces allow authors to extend and customize the comic generation
process, creating engaging and varied content. Additionally, the system supports testing various
machine learning models in comic generation. This approach balances AI and human creativity,
making the comic creation process more flexible and eficient.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Methodology</title>
      <p>The system is structured to a human-in-loop workflow as Fig1.ure</p>
      <p>The system comprises six modules: the graphical interface, a container of models, the pool of
visual sets and data, an image sequence model, a generator, and a renderer.
• Graphical Interface: This is the primary means of interaction for human authors. It
allows them to input base images, launch and apply ML models to the current image
sequence, import scripts for customized editing layers, and perform simple operations
like selection and dragging to modify images.</p>
      <p>• Container of Models: This module houses various ML models that can be applied during
generation.
• Pool of Visual Sets and Data: This repository contains visual elements and data sets
the system can draw upon to create the image sequences.
• Image Sequence Model: This model organizes the graph model for linking the elements
of the image sequence, including sequence information, panel transitions, panel content,
characters, and narrative elements.
• Generator: This component integrates the customized editing layers into a pipeline that
iteratively edits the image sequence by applying the narrative goals of each layer.
• Renderer: This module finalizes the visual representation of the generated sequences,
ensuring they are ready for presentation or further editing.</p>
      <p>Due to the complexity of applying ML models, some customizations of the generating process
rely on scripts that deploy API functions across diferent system modules. The process typically
follows nine steps between these modules:
1. Register Models: After receiving user scripts as input, tChoentainer of Models registers
all imported ML models and initializes them.
2. Register Editing Layers: This step links the interface buttons with class scripts,
dynamically imports user-defined classes, and executes them upon button click.
3. Retrieve Data: If user-customized scripts include functions to retrieve information, the</p>
      <p>Generator communicates with thPeool of Visual Sets and Data to obtain necessary data.
4. Apply Editing: The customized editing results reflect the corresponding changes to the
data nodes in thIemage Sequence Model.
5. Compose Visual Panels: After updating theImage Sequence Model with the desired
changes, theRenderer is triggered to start composing the visual results of comic panels.</p>
      <p>The updated graph model will also be passed toRtenhdeerer.
6. Map Semantic to Visual: The Renderer maps the data nodes to visual elements in the
Pool of Visual Sets and Data and forms the multi-layered panels. These layers include
background, foreground, composition, symbol, and any other user-customized layers.
7. Render: The multi-layered image panel is passed to the interface to support user selection
and drag functions.
8. Modify Results: The interface inputs user interactions to updaItmeatgheeSequence</p>
      <p>Model and repeats the rendering steps.
9. Get Result: The generated and modified image sequence is presented on the interface
for users to view.</p>
      <sec id="sec-3-1">
        <title>2.1. Build-in and Extendable Elements:</title>
        <p>To demonstrate the system’s capabilities, we incorporated a built-in model, the action causal
graph, and a visual symbol set as the default sources to support the plotline of the visual
representations. These components are extendable through the APICsoonftatihner of Models
and thePool of Visual Sets and Data, respectively. Detailed documentation will be provided in
the API subsections.</p>
        <sec id="sec-3-1-1">
          <title>2.1.1. Common Symbols</title>
          <p>
            In human-created comics, authors often use abstract symbols to visualize ideas such as
atmosphere, motion, and emotions, thereby exaggerating characters’ reactions. These symbols
include emojis to emphasize emotions, speed lines to show movement, explosion shapes to
represent collisions, cross shapes to denote anger, and many other1s.pTraesbelnets examples
from the Manga109 dataset, a collection of 109 Japanese manga titles published in commercial
magazines [
            <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
            ]. These examples serve as references for the symbols or emojis used in our
default set.
          </p>
        </sec>
        <sec id="sec-3-1-2">
          <title>2.1.2. Action Causal Network</title>
          <p>The Action Causal Graph is a directed graph model where each node represents an action that a
character might perform, and the links indicate the causal relationships between actions and
possible reactions. For example, ”Fall” links to ”Fly,” ”Jump,” and ”Run,” while ”Dizzy,” ”Collide,”
and ”Hit” link to ”Run.” Consecutive nodes form action pools for plot planning. The current
version of the default action set built into the system includes a small group of common daily
actions. The detailed method for expanding the graph and its scalability will be discussed in
subsequent subsections.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. AI-Assisted Editing Layers and Narrative Theories:</title>
        <p>We implement three ML-driven editing layers in the generator to demonstrate the system’s
image sequence pipeline. The first layer uses a stable difusion model for modifying visual
elements, while the second layer combines the PAD emotion model with a semantic analysis
language model to guide plotline decisions.</p>
        <sec id="sec-3-2-1">
          <title>2.2.1. Difusion Model for Visual Elements</title>
          <p>
            We employed a pre-trained stable difusion model developed by the CompVis group to implement
one of the editing layer1s4,[
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. This high-resolution image synthesis model transforms input
text descriptions into detailed and coherent images based on latent difusion. The model allows
users to adjust visual elements like characters or scenes.
          </p>
          <p>The generated panel sequences are rendered through our pipeline, where information is
ifrst updated in the Image Sequence Model before being rendered. Any changes in the visual
representation of an entity will be reflected in its semantic nodes, ensuring character
consistency throughout the panel sequence. For example, if the difusion model alters the visual
representation of character_x, this change will be mirrored in its semantic mapping. Similar
rules apply to other visual elements. Furthermore, our system supports multi-layer rendering,
dividing panel images into background, foreground, compositional, and symbol layers. This
feature enables partial redrawing of the generated panel, ensuring that changes applied to one
entity do not interfere with others.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>2.2.2. Plotline and Narrative Theories</title>
          <p>
            To form a simple plotline, we incorporate narrative idioms and theories. This editing layer aims
to generate content for the image sequence according to a specific narrative arc. We use Cohn’s
narrative gramma1r6[
            <xref ref-type="bibr" rid="ref17">, 17, 18</xref>
            ] to estimate the narrative arc, as the grammar categories indicate
plot changes throughout the comic sequence. After establishing the narrative arc, we target the
characters’ actions to predict story tension.
          </p>
          <p>We use the arousal level concept from the PAD emotion m1o9d,e2l0,[21] and sentiment
labels from the language model to predict arousal scores for character actions, where higher
scores indicate greater story tension. Diferences in arousal scores between consecutive actions
(based on our action causal graph) form a probability distribution for subsequent actions.
This mapping enables the editing layer to select actions probabilistically while maintaining
narrative arc alignment. Detailed explanations of each component are provided in the following
subsections.</p>
          <p>Narrative Grammar</p>
          <p>We formalize the plot generation of new comic sequences using Cohn’s Visual Narrative
Grammar (VNG). Starting with the narrative structure to determine the content’s global
reasoning, we adopt Cohn’s theory, which proposes that coherent image sequences follow a grammar,
organizing their global structure into five categories.</p>
          <p>• Establisher(E): Sets the objects and scenes without involving any action.
• Initial(I): Marks the beginning of a story arc—the starting point of a sequence of actions
or events.
• Prolongation(L): Represents the middle state of the story arc, extending an action.
• Peak(P): Indicates the highest story tension—the climax of an action.</p>
          <p>• Release(R): Releases the tension—the outcome or result of an action.</p>
          <p>The five categories form basic phases through linear ordering:
Phase (Establisher) - Initial(Prolongation) - Peak - (Release)</p>
          <p>The use of parentheses indicates that categories are optional when forming phases. The
categories have diferent levels of importance, ranked from highest to lowest as follows: Peak,
Initial, Release, Establisher, and Prolongation. Additionally, more complex combinatorial
structures can be created through the conjunction of embeddings. Our editing layer generates the
narrative structure using center-embedding, expanding a new tree structure by replacing a
single category with a phase.</p>
          <p>Table2 provides an example of comic sequences (illustrated with simplified icons) following
this grammar structure. In the same scene, the circle character appears in the first panel and
then begins performing actions in the subsequent panels. In the fourth panel, a significant event
occurs, creating a small climax, which leads to the resolution in the final panel. This sequence
follows the narrative strucEt-uI-rLe-P-R.</p>
          <p>The implementation follows the algorithm flow below:</p>
          <p>The input is a comic sequence, either a generated result or an empty sequence with a certain
length. Then, the editing layer creates an object dictionary for grammar phases and then
expands the tree structure. The structure then decides the narrative arc of the comic sequence.
It assigned a grammar phase to a sequence, and we then formed the possible narrative arc based
on the structure.</p>
          <p>Narrative Arc</p>
          <p>The importance ranking of narrative grammar categories mirrors the narrative arc, reflecting
a story’s progression from a calm beginning through a peak of tension in the middle to conflict
resolution at the end. The editing layer projects these categories onto a curve that illustrates
changes in story tension.</p>
          <p>Algorithm 2 Narrative Arc
1: procedure Mapping Narrative Arc with Narrative Structure
2:  ← input the current panel sequence.
3:  ← create value dictionary with VNG phases with assigned tension scores.
4:  = { ∶ 0 ∶ 2 ∶ 4 ∶ 6 ∶ 2}
5: if panels in have assigned grammar phastehen
6: loop: over , apply with the grammar phase
7: elseuse default narrative arc scores
8: end if
9: return 
10: end procedure</p>
          <p>To capture the abstract concept of tension, we assigned each grammar phase a score between
one and ten based on its narrative function. For example, the Peak(P) phase has the highest
tension, while the Release(R) phase somewhat alleviates the tension. We begin by creating
a value dictionary for each grammar phase. The next step is to map these phases to their
respective scores and generate the curve of the narrative arc.</p>
          <p>PAD Emotion State Model and Semantic Analysis</p>
          <p>The PAD emotion state model, developed by Albert Mehrabian and James A. Russell, describes
emotions through three dimensions: Pleasure, Arousal, and Dominance. This model quantifies
emotional states, making it valuable in psychology, user experience design, and AI. Specifically,
the Arousal dimension measures how energized or calm one feels, reflecting the activation level
of an emotion. We adapt this concept to model narrative momentum—story tension.</p>
          <p>
            Considering the image sequence-generating system’s future expansion and possible
integration with a more extended narrative, we employed a sentiment analysis model for sentences–
Roberta base model for emotion classification, fine-tuned on the GoEmotions da2t2a]s.eWte[
then project the sentiment labels in the RoBERTa model to the PAD model’s emotion labels,
dividing the emotions into high, medium, and low arousal levels. By computing the feature
vectors through the BERT mod2e3l][of the two label sets, we can measure the Euclidian
distance between the labels, estimating the possible arousal level scores for the sentiment labels.
The mapping process follows the flow below:
Algorithm 3 Mapping Sentiment Labels with Emotion Labels
1: procedure Estimating Arousal Levels
2:  ← emotion labels from the PAD emotion state model, where [high, medium, low]
arousal level maps to [
            <xref ref-type="bibr" rid="ref1">1, 0, -1</xref>
            ].
3:  ← sentiment class labels from the RoBERTa model.
4:  ← Distance matrix for distance between lab elasnind labels i n.
5: loop: over , compute the distance to each elemen t, iannd get
6: loop: over each row in , flat the vector except the minimum two elements in the row.
7: loop: over each row in , use /() as the weight multiply withto get
 _ .
8: normalize _ to[
            <xref ref-type="bibr" rid="ref1">−1, 1</xref>
            ]
9: return  _
10: end procedure
          </p>
          <p>Figure2 shows the results, where the blue points represent the emotion labels from the PAD
model, and the orange points represent the sentiment classes in the language model.</p>
          <p>The sentiment analysis model predicts probabilities across various sentiment classes when
using a sentence or word as input. We use these probabilities as weights and then measure
the distance between the input action and the sentiment classes, predicting the possible score
of the input action. This score indicates the story’s tension. We further compute the slope
between two consecutive actions and divide it by the sum of slopes for all subsequent actions.
Finally, we normalize this value to determine the probability of the following action.</p>
          <p>Action Mapping
By mapping arousal scores with actions, the editing layer integrates the narrative arc and actions
by referencing these scores. Using the curve from the narrative arc, the layer calculates changes
along the curve and selects actions that best fit these changes. Additionally, it sets a likelihood
tolerance with the probability of actions, expanding potential narrative diversity. The process
is described below:
Algorithm 4 Narrative Arc Mapping
1: procedure Mapping Actions to Fit Narrative Arc
2:  ← input the current panel sequence.
3:   ← get the action causal graph network.
4:  ← creates a value dictionary for actions according to the arousal scores.
5:  ← get referenced Narrative Arc.
6: loop: over , check the characters’ actions in the next panel and compute the score
diference, according t o
7: Select actions in likelihood f rom and 
8: loop: revise action selection other panels accord ing to
9: return 
10: end procedure</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>2.2.3. Panel Relations</title>
          <p>Panel transitions refer to the changes in content between consecutive panels. McCloud proposed
six categories of transition types to model various aspects of content2]c,hwahnilgee[Cohn
introduced conjunction schemes to capture more complex panel tran1s6it].ioOnusr[system
combines these theories to modify the visual composition of comic panels—how elements are
arranged—and reflect content changes according to the transition types. The narrative goal of
this editing layer is to arrange the panel transitions in the generated comic sequence to create
dynamic viewport changes and increase tension.</p>
          <p>Here is how we map the transitions with panel content changes:
• Action: McCloud’s action-to-action transition indicates changes in actions between
consecutive panels. We use this transition to guide the selection of diferent character
actions in the next panel.
• Scene: McCloud’s scene-to-scene transition indicates changes in scenes between
consecutive panels. We use this transition to guide the change of location where the character’s
action occurs in the next panel.
• Object: The object-to-object transition indicates a shift in focus from one object to
another. We use this transition to guide the focus to diferent objects.
• Addition: Cohn’s additive conjunction involves panels that add information or detail to
the ongoing story. We use this transition to introduce new objects into the panels.
• Alternation: Cohn’s alternating conjunction presents alternative scenarios or actions,
ofering possibilities within the narrative. Compared to scene transitions, we use this
transition to guide panels to alter most elements while maintaining consistent characters.</p>
          <p>In our system’s current version, some complex transitions or conjunctions are not yet
supported but can be added through customizations. For example, the Temporal Conjunction
depicts the progression of time, a rather abstract concept for visual representation. Another
example is the Contrasting Conjunction, which depicts opposing ideas, actions, and emotions;
achieving this requires deep semantic analysis of the narrative. However, we have integrated
causal conjunctions to some extent by using action causal graphs.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Usage Explanations and Showcases</title>
      <p>This section introduces our system’s application programming interface (API), and graphical
user interface (GUI). The API functions are in T3a.bTlehe system was constructed using several
core classes, with tPhaerameter, Layer, andAttributeNode classes being the most crucial. The
Parameter class manages all the registers of models and scriptAs.ttTrhibeuteNode class serves
as the fundamental component of the graph model, representing the entire sequeLnacyee.rThe
class is the parent class for all editing layer scripts. To execute modificataioppnlsy, tmheethods
within thLeayer class must be overridden and then executed by the Generator module.</p>
      <p>The GUI screenshot is Figur3e.The two columns on the left display the image inputs for the
character and scene of the comic sequence. Users can generate these images using the difusion
model or import them from their work. The column on the right side contains buttons that link
to the imported scripts, triggering functions to apply modifications and generate results. The
large area in the middle shows the currently generated result and allows users to select comic
panels and the elements within them.</p>
      <sec id="sec-4-1">
        <title>3.1. Showcases</title>
        <p>Table4 compares two sets of generated results, one before and one after applying the changes
implemented in the current version. The first set, generated with user-imported images and all
editing layers turned of, features content with default panel composition and randomly chosen
characters’ actions. After partial redrawing with the stable difusion model, the second set
demonstrates the results, followed by an editing layer that uses the sentiment language model
to achieve narrative planning. It includes a partially redrawn icon and scene using a difusion
model from an Einstein head icon and a WindowsXP desktop photo as inputs. In addition, based
Parameter()
addModel
importModels()
Models()
addModule()
Generator()
addVisuals()
importModules() None</p>
        <p>Initialized Parameter
object and attribute node
type
Sequence()</p>
        <sec id="sec-4-1-1">
          <title>Parameters</title>
          <p>Basic parameter for GUI
settings
• win_w, win_h
• menu_w
• sequence_len
Selected ML model class</p>
          <p>• module_name
None
None
Customized
layer class
• layer_name</p>
          <p>editing
• attribute_type
• Parameter()
None
Name of target visual set,
file path, or folder path
for image input
• set_name
• path
Initialized Parameter
object</p>
          <p>• Parameter()
The graph model of a
generated sequence</p>
          <p>• Sequence()
The attribute type and
the initialized Parameter
class
• attribute_type
• Parameter()
The parent node’s name,
attribute type, and the
child node (self).</p>
          <p>• parent_node
• attribute_type
• AttributeNode()
GUIInterface()
Layer()</p>
          <p>None
Layer.apply()
AttributeNode()
addAttribute()
The class that registers and manages
all the modules.</p>
          <p>Register the ML models for use in the
generating process.</p>
          <p>Import all the registered models.</p>
          <p>Initialize the Container of Models and
initialize all the registered models.</p>
          <p>Register script for a new editing layer.</p>
          <p>None
Import all the registered editing layers.</p>
          <p>None
Inherit from the AttributeNode class
and initialize the root of the graph
model for the Image Sequence Model.</p>
          <p>Initialize the Generator Module.</p>
          <p>Expand or create the visual set with the
assigned name.</p>
          <p>None
None
None
None
None
None
None
None
None
None</p>
          <p>None
Initilize the GUI and the Renderer.</p>
          <p>The Parent class of all the editing
layers.</p>
          <p>Apply the modifications to the comic
sequence and return the results.</p>
          <p>Sequence()
The Parent class of all the attribute
nodes, including Sequence, Panel,
Character, etc. It is the Node class for
the graph model in the Image Sequence
Model.</p>
          <p>Add an attribute node to the graph
model as a child node of the assigned
parent node.
on the action causal network, it shows a short narrative in which the two characters ate apples
and felt dizzy after eating, then shocked and finally rested in the garden.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Data Availability</title>
        <p>The data and code used in this study are openly available in a GitHub repository. The repository
includes the raw data, processed data, and all scripts necessary to reproduce the analyses
presented in this paper. You can access the repository at
https://github.com/RimiChen/Collaborative_Comic_Generation. The repository is licensed under MIT License, allowing for reuse
and modification with appropriate attribution.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion and Future Work</title>
      <p>This paper presents an extensible system for generating comic-style visual narratives, integrating
narrative theory with human-AI collaboration. We address challenges in visual modification
and plotline planning, balancing automation with user control for customization.</p>
      <p>While the current system efectively integrates abstract narrative theories, it has limitations
in coordinating visual components and diversifying narratives. Separating visual layers, though
beneficial for user modifications, reduces scene and character interaction, limiting cohesive
artwork and rich actions. Additionally, using symbols to represent actions restricts narrative
diversity, requiring user customization to expand content. Integrating advanced language
models could enhance narrative richness</p>
      <p>Future work will focus on refining the integration of narrative and visual components and
exploring advanced models to enhance the system’s capability to handle diverse narratives. We
also plan to conduct user studies to evaluate efectiveness and usability across diferent groups.
In conclusion, our work advances human-AI cooperative visual narrative generation, ofering a
versatile platform for creating engaging experiences.
(2015).
[18] M. I. N. COHN, From visual narrative grammar to filmic narrative grammar: The narrative
structure of static and moving images, in: Film Text Analysis, Routledge, 2016, pp. 106–129.
[19] A. Mehrabian, J. A. Russell, Basic dimensions for a general psychological theory:
Implications for personality, social, environmental, and developmental studies, Journal of
Comparative and Physiological Psychology 55 (1974) 439–449.
[20] J. A. Russell, L. M. Ward, G. Pratt, Afective quality attributed to environments: A factor
analytic study, Environment and behavior 13 (1981) 259–288.
[21] J. A. Russell, A circumplex model of afect, Journal of Personality and Social Psychology
39 (1980) 1161–1178.
[22] S. Lowe, Roberta base model fine-tuned on goemotions for emotion classificat&lt;ihotnt, ps:
//huggingface.co/SamLowe/roberta-base-go_emotio,2n0s2&gt;1. Accessed: 2024-05-25.
[23] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional
transformers for language understanding, in: Proceedings of the 2019 Conference of
the North American Chapter of the Association for Computational Linguistics: Human
Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 4171–4186.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Cohn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jackendof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Holcomb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. R.</given-names>
            <surname>Kuperberg</surname>
          </string-name>
          ,
          <article-title>The grammar of visual narrative: Neural evidence for constituent structure in sequential image comprehension</article-title>
          ,
          <source>Neuropsychologia</source>
          <volume>64</volume>
          (
          <year>2014</year>
          )
          <fpage>63</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>McCloud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <article-title>Understanding comics: The invisible art</article-title>
          , volume
          <volume>106</volume>
          , Kitchen sink press Northampton, MA,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Martens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. EDU</given-names>
            , R. E.
            <surname>Cardona-Rivera</surname>
          </string-name>
          ,
          <string-name>
            <surname>U. EDU</surname>
          </string-name>
          ,
          <article-title>The visual narrative engine: A computational model of the visual narrative parallel architecture</article-title>
          ,
          <source>in: 8th Annual Conference on Advances in Cognitive Systems</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Team</surname>
          </string-name>
          ,
          <article-title>Image sequence creation with microsoft copilhott</article-title>
          ,ps://microsoft.com/copil,ot
          <year>2023</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-25.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , M. Cheng, J.
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Hou</surname>
          </string-name>
          , Storydifusion:
          <article-title>Consistent self-attention for long-range image and video generation</article-title>
          ,
          <source>arXiv preprint arXiv:2405.01434</source>
          (
          <year>2024</year>
          ). URL: https://arxiv.org/abs/2405.014.34
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Jing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Content-aware video2comics with manga-style layout</article-title>
          ,
          <source>IEEE Transactions on Multimedia (TMM) 17</source>
          (
          <year>2015</year>
          )
          <fpage>2122</fpage>
          -
          <lpage>2133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Gunasekara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. M.</given-names>
            <surname>Perera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Adhihetty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. D.</given-names>
            <surname>Kollure</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kodagoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caldera</surname>
          </string-name>
          ,
          <article-title>Generate comic strips using ai</article-title>
          ,
          <source>in: Proceedings of Conference on Transdisciplinary Research in Engineering</source>
          , volume
          <volume>1</volume>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Alves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McMichael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Simões</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Paiva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aylett</surname>
          </string-name>
          ,
          <article-title>Comics2d: Describing and creating comics from story-based applications with autonomous characters</article-title>
          ,
          <source>in: Proceedings of the International Conference on Computer Animation and Social Agents (CASA)</source>
          , Springer,
          <year>2007</year>
          , pp.
          <fpage>67</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Martens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Cardona-Rivera</surname>
          </string-name>
          ,
          <article-title>Discourse-driven comic generation</article-title>
          ,
          <source>in: Proc. International Conference on Computational Creativity</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Martens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Cardona-Rivera</surname>
          </string-name>
          ,
          <article-title>Generating abstract comics</article-title>
          ,
          <source>in: International Conference on Interactive Digital Storytelling</source>
          , Springer,
          <year>2016</year>
          , pp.
          <fpage>168</fpage>
          -
          <lpage>175</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Y.-C. Chen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Jhala</surname>
          </string-name>
          ,
          <article-title>A customizable generator for comic-style visual narrative</article-title>
          ,
          <source>arXiv preprint arXiv:2401.02863</source>
          (
          <year>2023</year>
          ). URLh:ttps://arxiv.org/abs/2401.028.63
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Narita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tsubota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yamasaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Aizawa</surname>
          </string-name>
          ,
          <article-title>Sketch-based manga retrieval using deep features</article-title>
          ,
          <source>in: 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)</source>
          , volume
          <volume>3</volume>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>49</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matsui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Aramaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fujimoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ogawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yamasaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Aizawa</surname>
          </string-name>
          ,
          <article-title>Sketch-based manga retrieval using manga109 dataset</article-title>
          ,
          <source>Multimedia Tools and Applications</source>
          <volume>76</volume>
          (
          <year>2017</year>
          )
          <fpage>21811</fpage>
          -
          <lpage>21838</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rombach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Blattmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lorenz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Esser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ommer</surname>
          </string-name>
          ,
          <article-title>High-resolution image synthesis with latent difusion models, 202a1r</article-title>
          .Xiv:
          <volume>2112</volume>
          .
          <fpage>10752</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rombach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Blattmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lorenz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Esser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ommer</surname>
          </string-name>
          ,
          <article-title>High-resolution image synthesis with latent difusion modelhs</article-title>
          ,ttps://github.com/CompVis/stable-diffus,
          <year>io20n22</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-24.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>N.</given-names>
            <surname>Cohn</surname>
          </string-name>
          , Visual narrative structure,
          <source>Cognitive science 37</source>
          (
          <year>2013</year>
          )
          <fpage>413</fpage>
          -
          <lpage>452</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Cohn</surname>
          </string-name>
          ,
          <article-title>How to analyze visual narratives: A tutorial in visual narrative grammar</article-title>
          , Online: http://www. visuallanguagelab. com/P/VNG_Tutorial.
          <source>pdf [last accessed: 1 March</source>
          <year>2016</year>
          ]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>