<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing adapted print publication accessibility via text-to-image synthesis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rostyslav Zatserkovnyi</string-name>
          <email>zatserkovnyi.rostyslav@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Petro Kutsyk</string-name>
          <email>kutsykpetro@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roksoliana Zatserkovna</string-name>
          <email>zatserkovna.r@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Volodymyr</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter T. Popov</string-name>
          <email>p.t.popov@city.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>City University of London</institution>
          ,
          <addr-line>Northampton Square, London, EC1V 0HB</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lviv University of Trade and Economics</institution>
          ,
          <addr-line>10 Tuhan-Baranovskyi Str., 79008 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Ukrainian Academy of Printing</institution>
          ,
          <addr-line>19 Pid Goloskom Str., 79020 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>One of the most pressing concerns in the field of adapted printed publications - that is, publications with additional supporting features to make them easily accessible by a wide variety of audiences is preparing illustrations that can clearly convey visual information to the viewer. These illustrations need to be created while accounting for the needs of a diverse inclusive audience, whose requirements may be affected by disabilities such as visual impairment. Currently, there is a limited number of illustrators who can appropriately produce a large number of illustrations which satisfy these requirements; therefore, illustrating print publications is a time-consuming and expensive process for non-profit organizations which are responsible for their production. This article proposes a method for enhancing illustrations within print publications, given the source file (such as a PDF file) of a print publication. Based on modern text-to-image generators, this method extracts all illustrations from a print publication; converts them into textual prompts for a modern text-to-image generator; and finally, produces a series of adapted alternatives for each of the chosen illustrations based on the textual prompts. This allows publishers to obtain accessible illustrations for their publication in a manner of minutes, speeding up the adaptation process and enhancing its accessibility.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Accessible publishing</kwd>
        <kwd>artificial intelligence</kwd>
        <kwd>image synthesis</kwd>
        <kwd>information technologies</kwd>
        <kwd>1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, assistive technologies have enabled diverse audiences of readers to freely access
written information, allowing them to work, study, and participate in civil society. One of the
most essential tools for this category of readers is accessible print publications. While people
with disabilities, such as visual impairment, often rely on e-readers; these may not always be
available or preferred. A recent study suggests that 65% of surveyed Americans have recently
read a print book, while only 30% have read an e-book, indicating that print books still remain
the most popular format for general readers [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In a classroom setting in particular, print books
provide unique benefits such as ease of use and improved notetaking [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Thus, by providing
access to appropriately adapted books, educational institutions as well as independent
organizations can make sure that visually impaired people can succeed.
      </p>
      <p>
        Incorrectly adapted printed books can pose various issues for readers with visual
impairment, such as small text size, poor font styling, insufficient contrast, graphics and tables
as well as other design factors. Unlike e-books, where these parameters can be custom-tailored
to the needs of a particular reader, printed publications are inflexible in their design,
exacerbating these issues [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This article focuses on improving the accessibility of illustrations
in particular, proposing a multi-step method to enhance the accessibility of all illustrations
within a given publication:
1. Extract all illustrations from the source file of a future print publication.
2. Convert all illustrations into prompts for a text-to-image synthesizer.
3. Synthesize new, adapted illustrations based on modified versions of these prompts.
4. Review adapted illustrations and re-introduce them into the source file.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>
        Existing studies in the field most often focus on e-book accessibility. Since e-book formatting is
not rigid and preset by the publisher, but rather determined by the e-reader, the focus on
ebook accessibility lies primarily in correctly marking important parts of the publication (such
as chapters or specific phrases which link to footnotes), as well as developing robust, flexible
ereader software that allows users to adjust fonts, text sizes, as well as other parameters [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        However, print books still have notable advantages, especially in the field of education.
Research suggests that readers tend to understand text slightly better when it’s printed rather
than viewed on-screen [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and one study suggests the haptic feedback of a touch screen (or PC
monitor) is different than that of a paper book, providing a less immersive experience [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        When it comes to adapting illustrations within print books, existing research focuses on
multimodal illustrations – for instance, tactile illustrations, which combine visual illustrations
with tactile Braille overlays, which is useful for readers with legal blindness [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. Still, there
remains the issue of creating effective adapted illustrations before their tactile component is
factored in. In recent years, AI algorithms, such as machine learning algorithms and neural
networks, have been trained to produce a variety of media content – and creating such
illustrations from scratch is one potential application of these generative methods [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed methodology</title>
      <sec id="sec-3-1">
        <title>3.1. Converting images to text prompts</title>
        <p>
          In order to synthesize new adapted illustrations for a print publication, our method must first
obtain the inputs – often known as “prompts” – for a well-known image generator such as
Stable Diffusion or Midjourney. Although in technical publications, illustrations often come
with captions describing the image, these captions are not always present, and often offer brief
interpretations which do not capture the full nuance of the captioned image. Thus, the images
must be converted to textual prompts using an AI model known as an image captioner. Acting
as the inverse of common image generators, image captioners are a form of feature extraction
models which convert images into text, functioning at the crossroads between computer vision
and natural language processing [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>
          Our chosen image captioner model is the CLIP Interrogator [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], based on Salesforce’s BLIP
model. This baseline is a multi-task model which is capable of both image understanding and
image generation, and can operate in three possible modes: an unimodal encoder, an
imagegrounded text encoder, and an image-grounded text decoder.
        </p>
        <p>
          Specifically, the model’s captioner is an image-grounded text decoder. Its intent is to generate
synthetic captions   given training images   collected from web datasets. In the BLIP learning
framework, this is combined with the filter, an image-grounded text encoder which removes
texts that are predicted to not match a given image – this is applied to both synthetic captions
  and real captions   found inside training datasets. This is combined with a set {( ℎ,  ℎ)} of
human-annotated images and texts to produce a robust training dataset for a ML algorithm [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]:
 = {(  ,   )} + {(  ,   )} + {( ℎ,  ℎ)}
        </p>
        <p>Our method works primarily with this model’s captioner. We opt to use the
BLIP-ImageCaptioning-Large sub-model, pre-trained on the COCO dataset to produce human-readable
captions for input images. However, human-readable captions are not the most effective way
to produce the inputs for an image generator. An image generator’s input text, known as a
“prompt”, needs to be detailed and describe multiple keywords pertaining to an image: the
subject of the image itself (which is typically produced by an image captioner out of the box),
as well as the style, resolution, color, lightning and other details.</p>
        <p>To that end, we use the CLIP Interrogator to first generate a baseline caption using BLIP,
and then simplify the caption while adding additional keywords which most closely match the
target image. These come from a predefined dataset known as “flavors”, and include keywords
such as “highly detailed”, “sharp focus”, “intricate”, “digital painting” as well as phrases
referring to specific objects and entities located within an image. The Interrogator selects the
most appropriate keywords and phrases from “flavors” dataset by measuring the distance
between a target image and each separate phrase.
(1)</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Synthesizing images from text prompts</title>
        <p>
          After the images within an adapted print publication are converted to text prompts, the next
step is to transfer them to an AI image generator such as Stable Diffusion to obtain a new set of
images designed with accessibility in mind. Since the CLIP Interrogator’s prompt generator has
been designed with Stable Diffusion in mind, its newest stable version, SDXL v 1.0, has been
selected as the image generator of choice [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Its open-source nature means that it can be
deployed on any local machine as part of our overall method.
        </p>
        <p>This image generator is a latent diffusion model, an improvement on traditional diffusion
models. Traditional diffusion models work by first “corrupting” training data, such as images,
by adding noise to their inputs in a step-by-step-process. At each time step, Gaussian noise is
added to a data distribution  0 ~  ( 0) with variance   ∈ (0, 1), resulting in the iterative
process over the distribution of the variable:
 ( 1, … ,   |   ) =</p>
        <p>(  |   −1)
 (  |   −1) =  (  ;
1 −     −1,    )</p>
        <p>This process is called forward diffusion, and concludes once the distribution  is sufficiently
similar to pure Gaussian noise.</p>
        <p>Reverse diffusion is the process of recovering the original image from the resulting noise. The
overall workflow of diffusion models after they have been trained is to generate new images,
given random noise as input. Latent diffusion models perform this process within latent space
– a mathematical representation of data where similar items are grouped.</p>
        <p>Aside from a shortened version of the CLIP Interrogator’s keyword-based output, we append
several keywords designed at simplifying and matching them to a more clear and simplified
style, such as “illustration for children”, “monochrome”, “very low detail” and “no shading”.
This ensures that, while the objects and entities denoted by the keywords are included within
the final image, it remains simplified without obstructing valuable information by noisy
elements. Keywords can be modified or adjusted as needed – for instance, “monochrome” may
be removed should we require a full-color illustration.</p>
        <p>The complete workflow of our adaptation method can be seen in Figure 1.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Discussion</title>
      <p>In Figure 1, we can see the results of this image-to-text generation model applied to one of the
illustrations from a selected Ukrainian adapted textbook. Initially, the caption generated by the
model is a verbose human-readable representation; the CLIP Interrogator modifies this to an
image generator based on matching keywords from a preset list, which is less human-readable,
but more effective as an input for a future image generation step.</p>
      <p>The number of keywords within the text prompt can be modified at will by restricting the
number of phrases returned within the Interrogator, prioritizing those which most closely
match the image. While a single image-to-text conversion is shown on Figure 1, our method is
a batch process. This means that after all images have been extracted from a printed
publication’s source file, they are converted into textual prompts with no additional human
interference, significantly speeding up the captioning process as compared to a human
captioner.</p>
      <p>An example of image generations, which are the final output of our method, can be seen in
Figure 3. Within this example, the leftmost image, also shown separately as Figure 4, is
particularly appropriate for accessible publications, as it creates an image in a simplified cartoon
style with minimal shading that still keeps its core elements (the girls in the foreground and
trees in the background) legible. This can be translated directly into a mixed-format illustration
which combines visual elements with Braille-like tactile dots.</p>
      <p>The software implementation of our method uses the Python-based Jupyter Notebook
environment to integrate several steps of the process: the PyMuPDF library is used to extract
images from a printed publication’s source file; the CLIP Interrogator (internally based on the
PyTorch library &amp; its torchvision extension) is responsible for extracting prompts for images;
while the SDXL 1.0 generator is used to create new illustrated images. On a RTX 4090 GPU, our
pipeline takes ~22 seconds to convert an original illustrated image into four generated variants
ready to be reviewed by a publication’s editor, meaning that the entirety of a print publication’s
illustrations can be regenerated within hours.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>This article describes a method for adapting illustrations within the source file print publication,
with no or minimal human supervision, that can significantly speed up the process of making
a publication more accessible to a wide audience of readers, such as people with vision
impairment. This method enables both educational and volunteer organizations to produce
high-quality illustrations for an adapted print publication.</p>
      <p>Potential areas for further research include improvements to the prompts used within our
image generation step – for instance, adding support for multi-colored, yet clean and simplified
illustrations. The image generation model itself also has potential for improvement; as AI image
generation is a rapidly developing field of research, our machine learning pipeline can be
periodically revisited to make use of the newest models and techniques.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Faverio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Perrin</surname>
          </string-name>
          ,
          <article-title>Three-in-ten Americans now read e-books</article-title>
          , Pew Research Center,
          <year>2022</year>
          . URL: https://www.pewresearch.org/short-reads/
          <year>2022</year>
          /01/06/three-in
          <article-title>-ten-americansnow-</article-title>
          <string-name>
            <surname>read-</surname>
          </string-name>
          e-books/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Amirtharaj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Raghavan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Arulappan</surname>
          </string-name>
          ,
          <article-title>Preferences for printed books versus E−books among university students in a Middle Eastern country</article-title>
          ,
          <source>Heliyon</source>
          (
          <year>2023</year>
          )
          <article-title>e16776</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.heliyon.
          <year>2023</year>
          .e16776.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zatserkovnyi</surname>
          </string-name>
          et al.,
          <article-title>Application for Determining the Usability of Adapted Textbooks by People with Low Vision</article-title>
          , in
          <source>: Proceedings of the 2023 IEEE 18th International Conference on Computer Science and Information Technologies (CSIT)</source>
          , Lviv, Ukraine,
          <fpage>19</fpage>
          -
          <lpage>21</lpage>
          October
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .1109/csit61576.
          <year>2023</year>
          .
          <volume>10324055</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Salmerón</surname>
          </string-name>
          et al.,
          <article-title>Reading comprehension on handheld devices versus on paper: A narrative review and meta-analysis of the medium effect and its moderators</article-title>
          ,
          <source>Journal of Educational Psychology</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Enhancing</given-names>
            <surname>Reach</surname>
          </string-name>
          :
          <article-title>The Fundamentals of eBook Accessibility</article-title>
          , Ingram Content Group,
          <year>2024</year>
          . URL: https://www.ingramcontent.com/publishers-blog/fundamentals-of-ebookaccessibility .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mangen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Weel</surname>
          </string-name>
          ,
          <article-title>The evolution of reading in the age of digitisation: an integrative framework for reading research</article-title>
          , Literacy (
          <year>2016</year>
          ) Vol.
          <volume>50</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>116</fpage>
          -
          <lpage>124</lpage>
          . doi:
          <volume>10</volume>
          .1111/lit.12086.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Valente</surname>
          </string-name>
          et al.,
          <article-title>Comprehension of a multimodal book by children with visual impairments</article-title>
          ,
          <source>British Journal of Visual Impairment</source>
          (
          <year>2023</year>
          )
          <volume>42</volume>
          (
          <issue>2</issue>
          ), 026461962311720. doi:
          <volume>10</volume>
          .1177/02646196231172071.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zebehazy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wilton</surname>
          </string-name>
          ,
          <article-title>Graphic Reading Performance of Students with Visual Impairments and Its Implication for Instruction and Assessment</article-title>
          ,
          <source>Journal of Visual Impairment &amp; Blindness</source>
          (
          <year>2021</year>
          ) Vol.
          <volume>115</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>215</fpage>
          -
          <lpage>227</lpage>
          . doi:
          <volume>10</volume>
          .1177/0145482x211016918.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kuzmin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pavlova</surname>
          </string-name>
          .
          <source>Analysis of Artificial Intelligence Based Systems for Automated generation of Digital Content, Computer Systems and Information Technologies</source>
          (
          <year>2024</year>
          ) no.
          <issue>1</issue>
          , pp.
          <fpage>82</fpage>
          -
          <lpage>88</lpage>
          . doi:
          <volume>10</volume>
          .31891/csit-2024
          <source>-1-10</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Udo</surname>
          </string-name>
          , T. Koshinaka,
          <string-name>
            <surname>Image Captioners Sometimes Tell More Than Images They See</surname>
          </string-name>
          , arXiv.org,
          <year>2023</year>
          . URL: https://arxiv.org/abs/2305.02932.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <article-title>pharmapsychotic/clip-interrogator: Image to prompt with BLIP and CLIP</article-title>
          , GitHub,
          <year>2024</year>
          . URL: https://github.com/pharmapsychotic/clip-interrogator.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          et al.,
          <article-title>BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, arXiv</article-title>
          .org,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2201.12086.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Stable Diffusion</surname>
            <given-names>XL</given-names>
          </string-name>
          <source>- SDXL 1.0 Model</source>
          ,
          <year>2024</year>
          . URL: https://stablediffusionxl.com/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>