<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IRCDL</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Paving the Way for Personalized Museums Tours in the Metaverse</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alex Falcon</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beatrice Portelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ali Abdari</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Serra</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Naples Federico II</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Udine</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>20</volume>
      <fpage>22</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>Museums play a central role in the preservation and communication of human history. With the advent of powerful and accessible Virtual Reality technologies, new Metaverse Museums started being developed, creating new possibilities for discovering and experiencing knowledge from all over the world. In anticipation of this technology becoming more widely adopted, we must prepare tools to aid future visitors in finding and navigating the museums and exhibitions which are more relevant to their current interests. In this light, Deep Learning methods could be of great use for modeling Metaverse museums, retrieving the most relevant ones, and creating personalized tours of the artifacts contained within them. In this paper, we present our research project, “Personalized Museum Tours in the Metaverse” led by the Artificial Intelligence Laboratory of Udine (AILAB Udine), detailing its objectives, proposed methodology, and future directions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Metaverse</kwd>
        <kwd>Cross-modal Understanding</kwd>
        <kwd>Multimedia</kwd>
        <kwd>Virtual Museums</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and Motivation</title>
      <p>Museums are important institutions in our society, which are open to the public and conserve,
acquire, exhibit, and communicate testimonies of man and their environment. As society changes
and transforms in time, museums also have renewed themselves, adapting their objectives,
roles, and the way they interact with the public. In particular, nowadays it is not unusual for
them to leverage and incorporate digital technologies to ofer a more immersive experience for
the users, for example through downloadable guides, QR-codes next to the artworks to access
detailed descriptions, and even online exhibitions.</p>
      <p>The peak expression of this evolution would be a completely Digitized Museum, or a Museum
in the Metaverse, accessible through Augmented Reality (AR) technologies or even completely
built-in Virtual Reality (VR). There are several reasons why Metaverse Museums could be a
beneficial evolution of this institution: wider access to history and knowledge from all over the
world; a reduction in carbon emissions caused by people physically traveling to the museum,
especially for long-distance travels; a safer way to visit exhibitions in case of healthcare crises
(e.g., the recent COVID-19 pandemic); a more accessible way to visit exhibitions for people with
physical impairment or other restrictions that prevent them from safely visiting museums in
person. Additionally, a museum in the Metaverse allows for more flexibility and interactivity in
the contents presented to the visitors. For example, it could contain recreations of historical
places and artifacts that no longer exist (e.g., project Rekrei1), and provide access to places
too delicate for physical interaction. Consider the Lascaux Cave in France, closed due to the
paintings’ deterioration from light, air changes, and carbon monoxide from visitors’ breath2.</p>
      <p>However, if Metaverse Museums were to become more common and accessible, it could be
dificult for users to choose among all the possible interesting exhibitions and navigate the
immense amount of artifacts and information. This is an area where Artificial Intelligence (AI)
and Deep Learning (DL) techniques can be leveraged to alleviate the cognitive load on the user
and enhance the discoverability of the artworks.</p>
      <p>Our project aims to develop a multimedia system for the creation of personalized Metaverse
museum tours based on the user interests, formulated through a textual query. Realizing such a
system entails an in-depth analysis of heterogeneous types of data, including visual (3D scenes,
paintings, statues, etc) and textual (artifact descriptions and user queries). Notably, this project
deals with many innovative aspects: the analysis of Metaverse museums and their artifacts, the
estimation of their relevance to the user interests, and the automatic generation of personalized
tours, blending the artifacts’ descriptions with the desires of the user and vast amounts of
external knowledge.</p>
      <sec id="sec-1-1">
        <title>1https://rekrei.org/about 2https://archeologie.culture.gouv.fr/lascaux/en/preservation-lascaux-hill</title>
        <p>I’d like to see
paintings of people
playing together
on a sunny day.</p>
        <p>At museum M1, you will find 5
beautiful paintings relevant to your
query. First, you should see ..</p>
        <p>At museum Mk, you will find 2
beautiful paintings relevant to your
query. First, you should see ..</p>
        <p>Metaverse Museums</p>
        <p>M1 M2 M3 M4 … Mn
2</p>
        <p>UndeQrustearnyding
3</p>
        <p>Metaverse</p>
        <p>Analysis &amp;
Query-Artifact</p>
        <p>Relevance
top-k museums
M1 M2 … Mk
4</p>
        <p>Personalized
Tour Creation</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Personalized Metaverse Museums Tour Creation: Project</title>
    </sec>
    <sec id="sec-3">
      <title>Description</title>
      <p>The research project stems from the need for tools to understand, represent, and navigate
complex scenes such as digitized Metaverse museums. The final objective is to develop a system
to create personalized Metaverse museum tours, given a user query. Figure 2 presents an
overview of the proposed system.</p>
      <p>The user interacts with the system by formulating a query in natural language. The Query
Understanding module analyzes the user query and represents it by means of a compact
representation wherein its semantics are preserved. This representation will serve as guidance
for the Metaverse Analysis &amp; Query-Artifact Relevance module. In fact, this module
processes all the available Metaverse museums and ranks them according to their relevance to
the user query. After a subset of the most relevant museums is identified, the Personalized
Tour Creation module outputs for each of them a suggested itinerary covering the relevant
artifacts and explaining how they relate to the user desires.</p>
      <sec id="sec-3-1">
        <title>2.1. Query Understanding: Modelling the User Query</title>
        <p>
          The advancements in natural language processing made in the past few years enabled interactive
applications to understand the semantics of user-generated queries. These include
wordlevel techniques, such as word2vec [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], GloVe [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], and fastText [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], which compute a vector
representation for each word, hence using the words as a unit of meaning; and sentence-level
techniques, such as BERT [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], T5 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], and GPT [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], which use sentences as a unit of meaning.
        </p>
        <p>In practice, to have a better understanding of the user query and its implications, in our
project we will employ a combination of both word- and sentence-level techniques.</p>
        <p>
          Notably, the user queries may be objective and ask for specific authors or art schools, but
they also may be much more subjective and focus on feelings or contain very general and
imprecise prompts (e.g., “happy people playing on a beach” or “mysterious doctors performing
an autopsy”). Therefore, it becomes crucial to integrate these two types of information into
the training procedure. To achieve this goal, we plan to do the following: for each painting,
we consider both an objective description of its visual contents and one or more subjective
descriptions capturing the emotions that it sparked in the observer. Then, during training, the
knowledge from both types of descriptions is integrated into the representations of the visual
artifacts and the user queries. Another tool to relate subjective and objective descriptions is
represented by models trained for formality style transfer [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], as they paraphrase a sentence
from one style to another while preserving the semantic meaning. They could be used either to
generate training data or to transform the user query and use it to match more artifacts.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Metaverse Analysis &amp; Query-Artifact Relevance: Modelling the Museum and the Cross-modal Interactions</title>
        <p>To understand the contents of a Metaverse museum and relate them to the user query, both
local (fine-grained, artifact-level) and global (coarse, museum- or room-level) information
are needed. A representation is then computed for each museum by aggregating these two
sources of information. Finally, they are ranked according to their relevance to the user query,
obtaining the final output of this module.</p>
        <p>
          Local Information. To provide the subsequent modules with precise knowledge of the
relevant artifacts present in the room/museum, the artifacts must first be identified and analyzed.
To automatize this process, several well-known and widely used object detectors in the vision
community can be leveraged, such as YOLO [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and Faster R-CNN [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Once the artifacts are
localized, multiple methodologies can be followed to obtain the needed local information.
        </p>
        <p>
          A first approach represents each object in the scene with its categorical and spatial attributes
(e.g., size, position, and rotation in the 3D space), as done in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Although these attributes
can create a representation with discriminative qualities, there are two limitations: first, the
obtained representation lacks visual grounding; second, it will not lie in the same embedding
space as the query, hence requiring a finetuning step to align the visual and textual spaces.
        </p>
        <p>
          Another possible approach uses a jointly trained vision-language model (e.g., CLIP [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]) to
extract a representation for each artifact. Interestingly, since such an approach involves learning
a joint vision-language embedding space, the artifact representation will lie in the same joint
embedding space as its own description, hence making it easier to understand whether it is
relevant to the query. We followed a similar approach with positive results in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], where we
developed a system to rank digital apartments according to a textual query.
Global Information. To obtain the global information, a naive solution could aggregate
the pieces of local information obtained in the previous step, e.g., by using a Variational
Autoencoder [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] to represent a 3D scene through its objects as done in our previous work [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]
about Metaverse scene retrieval. Also, knowledge graphs can aggregate the local information
by discovering relationships among the artifacts in a room (or museum) [
          <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
          ]. An
alternative method, diverging from the aggregation of local information, involves the use of 3D
Convolutional Neural Networks (CNNs, e.g., [
          <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
          ]) to capture finer-grained spatial relations
between the artifacts. Spatial attention plays a key role in reducing the amount of redundancy
in the learned representations due to the presence of large empty spaces such as walls, halls,
etc. However, the obtained representations would not lie in the same embedding space of the
queries, hence requiring further finetuning to align the two spaces.
        </p>
        <p>Query Relevance and Ranking. The final step performed by this module consists of ranking
the museums according to their relevance to the user query. Two important aspects should
be considered during this step. First, a relevance score should be computed for each artifact,
since few of them are likely to be highly relevant to the query, and those few will impact the
ifnal relevance of the museum as well as becoming the key elements of the personalized tour.
Second, the relevance value for the museum is influenced by several global aspects, including
the number of relevant artifacts per room (or within the museum), and how the relevance values
are distributed (e.g. if there are many low-relevance artifacts or only a few highly relevant
ones). Moreover, further constraints could be added, such as ranking the museums based on
how much the relevant artifacts are clustered within the museum’s rooms. The output of this
step is a ranking of the museums, where for each museum a few artifacts are highlighted as the
most relevant ones for the query.</p>
      </sec>
      <sec id="sec-3-3">
        <title>2.3. Personalized Tour Creation: Presenting Suggestions to the User</title>
        <p>
          For each of the museums selected by the previous module, the Personalized Tour Creation
module creates a textual summary of the relevant artifacts and motivation on their relevance to
the user query and interests. To put together this information, we will use a Large Language
Model (LLM) such as ChatGPT, or open-source alternatives such as LLaMa [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] and the recent
Phi2 [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. These models are impressive tools to automatically create coherent and well-written
summaries. Moreover, since they are pretrained on large amounts of textual data from the
Web, they also have external knowledge of the artifacts which may be useful in adding more
insights to the personalized tour. Finally, due to the model’s prompt-based nature, the tour
can be personalized by creating specialized prompts. For instance, the summary obtained by
prompting ChatGPT with “create a summary of why the artifacts The Annunciation by Leonardo
da Vinci, The Annunciation by Benvenuto Tisi, and The Annunciation by Sandro Botticelli are
relevant to the user interests defined by ‘I want to see the graceful Holy Mary being announced
the birth of Jesus by the angel’” contains both the motivation for the relevance (e.g., “The angel
Gabriel’s announcement to Mary is depicted with a sense of divine tranquility, emphasizing
Mary’s graceful response to the news”) and some key features of the artifact or the style pursued
by the author (e.g., “Da Vinci’s meticulous attention to detail and his ability to convey emotions
through facial expressions make this artwork a captivating representation of the Annunciation”).
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>2.4. Dataset Creation</title>
        <p>At the time of writing, there are no public datasets containing a collection of Metaverse museums.
Therefore, we decide to create a new dataset as a resource for us and future researchers. In
particular, the dataset should contain a large amount of digital museums, where each room
contains multiple artifacts, including paintings, 3D sculptures, and other items of interest.
Moreover, to enable fine-grained cross-modal applications and to estimate the relevance of each
item to the user query, detailed descriptions of the artifacts are needed. As mentioned in the
previous sections, both subjective and objective textual descriptions are crucial to include in
the dataset, due to the wide range of possible user queries.</p>
        <p>
          To obtain a large amount of digital museums, we resort to the creation of a synthetic dataset.
This process involves several steps: first, collecting a diverse set of artifacts, including 2D and
3D art; second, obtaining meaningful descriptions for the artifacts, both objective and subjective;
third, creating each 3D museum with multiple rooms, each containing several artifacts.
Metaverse Museums Creation. We are currently developing an automatic approach to
create 3D museums, each containing multiple rooms with a variable number of sculptures
and/or paintings. The items are either attached to the walls or placed in key positions in the
room (e.g., corners or center of the room). The items will be grouped in the rooms based on
topic, artist, theme, or other criteria, in order to create realistic thematic exhibitions.
2D Artwork Collection. To collect painting data and descriptions, we consider several public
datasets, focusing on the type of annotations. Image captioning techniques like MPlug [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]
and datasets such as ArtCap [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] provide objective descriptions. However, they both lack
expert insights into interpretation and historical context, and subjective feelings expressed by
observers. For the insights, SemArt [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] ofers contextual details and artistic comments on
over 21k paintings from the Web Gallery of Art (WGA). ArtEmis [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] includes 400k non-expert
annotations about the emotions felt while observing fine art, while ArtELingo [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] adds
multilingual annotations and insights into cultural variations in emotional responses to the same
artwork.
3D Artwork Collection. While there are many popular datasets for 3D objects (e.g., ShapeNet
[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]), large open datasets on 3D statues and sculptures are not readily available. Recently,
Objaverse [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] and ObjaverseXL [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] were released, which contain over 800k and 10M annotated
3D objects respectively. Among these objects, there are also artistic statues and sculptures
which are relevant for our purposes. Unfortunately, the textual annotations are sparse and do
not contain an extensive description of the item’s visual appearance or any subjective comments.
To address this, captioning techniques could be used to obtain objective descriptions, whereas
artistic comments could be scraped (e.g., from the WGA when possible).
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Related work</title>
      <sec id="sec-4-1">
        <title>3.1. Research on Applications for the Metaverse</title>
        <p>
          A widespread range of applications have been designed for the Metaverse, ranging from
entertainment to virtual shopping [
          <xref ref-type="bibr" rid="ref28 ref29">28, 29</xref>
          ] and apartments scouting [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], from education to healthcare
[
          <xref ref-type="bibr" rid="ref30 ref31">30, 31</xref>
          ] and industrial applications [
          <xref ref-type="bibr" rid="ref32 ref33">32, 33</xref>
          ]. Although some Metaverse museums are available
online (e.g., the Museum of Other Realities3, and The Vordun Museum4), the research on them
is still in their infancy [
          <xref ref-type="bibr" rid="ref34 ref35">34, 35</xref>
          ] and many topics on their automatic understanding are widely
        </p>
        <sec id="sec-4-1-1">
          <title>3https://www.museumor.com/</title>
          <p>
            4https://secondlife.com/destination/vordun-museum-and-gallery
unexplored. For instance, the task of retrieving multimedia-enriched scenarios by means of
textual queries was only recently introduced [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. Diferently, in this project, we aim to create
personalized Metaverse museum tours in an automatic way.
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Retrieving and Ranking Artistic Artifacts</title>
        <p>
          The task of retrieving and ranking artifacts according to their relevance to a user query is
similar to the more general text-to-image retrieval task, which has seen great improvements in
recent years, leading to the development of high-performing tools [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Conde and Turgutlu
[
          <xref ref-type="bibr" rid="ref36">36</xref>
          ] developed a direct transposition of CLIP for the art domain, a pretrained model that can be
used to describe images containing artworks or retrieve artworks given a description. Other
interesting approaches include an unsupervised method [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ] to discover the relationships
among diferent artworks and schools only based on visual features, and diferent approaches
employing knowledge graphs to capture the underlying relationships between other artistic
attributes too, such as style, movement, subject, and period [
          <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Creating a Personalized Museum Tour</title>
        <p>
          Suggesting a museum tour or a city itinerary to tourists is a dificult topic which sparked
interest in the research community since before the advent of deep learning. Early approaches
for creating personalized demo tours included three main steps: first, a user model is created by
asking the users to rate artworks; second, an art recommendation step is performed based on
the user model and a set of predefined topics; third, a tour creation tool selects a fixed amount
of artworks from those recommended in the previous step [
          <xref ref-type="bibr" rid="ref38 ref39">38, 39</xref>
          ]. More recent literature on
this topic often focuses on improving or optimizing the itinerary based on several constraints,
e.g., travel costs [
          <xref ref-type="bibr" rid="ref40">40</xref>
          ] or visit duration [
          <xref ref-type="bibr" rid="ref41">41</xref>
          ]. Three main diferences set our work apart from the
previous ones. First, the art recommendation part is done automatically through cross-modal
analysis, letting the user inform the model about their interests via text, therefore avoiding the
imposition of predefined topics. This approach is suitable for modeling user interests, whether
they are described using technical terms or a simpler lexicon. It also helps in learning how to
contextualize the user query within the artwork visuals and textual descriptions. Moreover,
our procedure does not require every artwork to be tagged based on a set of predefined topics,
demanding expert knowledge to create a minimal set of comprehensive tags and extensive
human eforts from experts to tag every artwork featured in a virtual museum, therefore
limiting scalability. Second, the diferent context (Metaverse) makes some of the previously
investigated constraints less relevant, such as optimizing travel costs or even visit duration,
since the user can visit the museum during diferent sessions and across several days from
the comfort of their home. Third, our aim is to not only provide a list of artifacts to the user,
but to inform them, creating a brief summary explaining why the user should visit a certain
museum, while specifying a selection of the most relevant artifacts. We envision LLMs as the
main tool to create a coherent tour and to provide contextual information related to the artifacts.
Nonetheless, external knowledge could also be integrated from knowledge bases, similarly
to [42], which combines LLMs and Wikipedia to generate in-depth descriptions of artworks
including background information about the history of the painting too.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusions</title>
      <p>In this paper, we presented a research project detailing the first steps to realize an innovative
multimedia system for creating personalized museum tours in the Metaverse. This project is
motivated by the increasing attention driven to the Metaverse and its vast possibilities in terms
of interactive and immersive experiences, and the importance of spreading knowledge on fine
arts and cultural heritage across the globe. We provided an overview of the proposed system, and
in-depth description of its components and how they relate to state-of-the-art methodologies in
vision-language understanding. Given the lack of suitable public datasets for experimentation,
we also provided a detailed explanation of the data collection process that is being currently
carried out. We hope that in the near future the dataset will be a valuable resource for researchers
and practitioners working on the intersection of Metaverse and multimedia understanding, and
that this project will spark interest in the computer science and cultural heritage community
fostering collaborative endeavors.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This work was supported by the PRIN 2022 “MUSMA” - CUP G53D23002930006 - “Funded by
EU - Next-Generation EU – M4 C2 I1.1”, and by the Department Strategic Plan (PSD) of the
University of Udine–Interdepartmental Project on Artificial Intelligence (2020-25).
itinerary planning, IEEE Access 8 (2020) 88147–88157.
[42] Z. Bai, Y. Nakashima, N. Garcia, Explain me the painting: Multi-topic knowledgeable art
description generation, in: Proceedings of IEEE/CVF ICCV, 2021, pp. 5422–5432.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          , T. Mikolov,
          <article-title>Distributed representations of sentences and documents</article-title>
          ,
          <source>in: Proceedings of ICML, PMLR</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1188</fpage>
          -
          <lpage>1196</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          , Glove:
          <article-title>Global vectors for word representation</article-title>
          ,
          <source>in: Proceedings of EMNLP</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , É. Grave,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , T. Mikolov,
          <article-title>Bag of tricks for eficient text classification</article-title>
          ,
          <source>in: Proceedings of EACL</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>427</fpage>
          -
          <lpage>431</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proceedings of NAACL-HLT</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>
          ,
          <source>The Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>5485</fpage>
          -
          <lpage>5551</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Narasimhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salimans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , et al.,
          <article-title>Improving language understanding by generative pre-training (</article-title>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tetreault</surname>
          </string-name>
          ,
          <article-title>Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer</article-title>
          ,
          <source>in: Proceedings of NAACL-HLT</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>129</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Redmon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Divvala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <article-title>You only look once: Unified, real-time object detection</article-title>
          ,
          <source>in: Proceedings of IEEE CVPR</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>779</fpage>
          -
          <lpage>788</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <surname>Faster</surname>
          </string-name>
          r-cnn:
          <article-title>Towards real-time object detection with region proposal networks</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>28</volume>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Yan,
          <string-name>
            <given-names>H.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bajaj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Scene synthesis via uncertainty-driven attribute synchronization</article-title>
          ,
          <source>in: Proceedings of IEEE/CVF ICCV</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>5630</fpage>
          -
          <lpage>5640</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          , et al.,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <source>in: Proceedings of ICML</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>8748</fpage>
          -
          <lpage>8763</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Falcon</surname>
          </string-name>
          , G. Serra,
          <article-title>Farmare: a furniture-aware multi-task methodology for recommending apartments based on the user interests</article-title>
          ,
          <source>in: Proceedings of IEEE/CVF ICCV</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>4293</fpage>
          -
          <lpage>4303</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Falcon</surname>
          </string-name>
          , G. Serra,
          <article-title>Metaverse retrieval: Finding the best metaverse environment via language</article-title>
          ,
          <source>in: Proceedings of 1st International Workshop on Deep Multimodal Learning for Information Retrieval</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Castellano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Digeno</surname>
          </string-name>
          , G. Sansaro, G. Vessio,
          <article-title>Leveraging knowledge graphs and deep learning for automatic art analysis</article-title>
          ,
          <source>Knowledge-Based Systems</source>
          <volume>248</volume>
          (
          <year>2022</year>
          )
          <fpage>108859</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Renoust</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nakashima</surname>
          </string-name>
          ,
          <article-title>Contextnet: representation and exploration for painting classification and retrieval in context</article-title>
          ,
          <source>International Journal of Multimedia Information Retrieval</source>
          <volume>9</volume>
          (
          <year>2020</year>
          )
          <fpage>17</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Graham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nießner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>Exploring data-eficient 3d scene understanding with contrastive scene contexts</article-title>
          ,
          <source>in: Proceedings of IEEE/CVF CVPR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>15587</fpage>
          -
          <lpage>15597</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Genova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tagliasacchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pollefeys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Funkhouser</surname>
          </string-name>
          , et al.,
          <article-title>Openscene: 3d scene understanding with open vocabularies</article-title>
          ,
          <source>in: Proceedings of IEEE/CVF CVPR</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>815</fpage>
          -
          <lpage>824</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Stone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Albert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almahairi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Babaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bashlykov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhargava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhosale</surname>
          </string-name>
          , et al.,
          <source>Llama</source>
          <volume>2</volume>
          :
          <article-title>Open foundation and fine-tuned chat models</article-title>
          ,
          <source>arXiv preprint arXiv:2307.09288</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Aneja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bubeck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>César Teodoro Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Del Giorno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Eldan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gopi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gunasekar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Javaheripi</surname>
          </string-name>
          , P. Kaufmann,
          <string-name>
            <given-names>Y. T.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , G. de Rosa,
          <string-name>
            <given-names>O.</given-names>
            <surname>Saarikivi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Santacroce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Singh</given-names>
            <surname>Behl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Taumann</given-names>
            <surname>Kalai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ward</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Witte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y. Zhang, Phi-2:
          <article-title>The surprising power of small language models</article-title>
          , https://www.microsoft.com/en-us/research/blog/ phi-2
          <article-title>-the-surprising-power-of-small-language-</article-title>
          <string-name>
            <surname>models</surname>
            <given-names>/</given-names>
          </string-name>
          ,
          <year>2023</year>
          . Accessed: February 12,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cao</surname>
          </string-name>
          , et al.,
          <article-title>mplug: Efective and eficient vision-language learning by cross-modal skip-connections</article-title>
          ,
          <source>in: Proceedings of EMNLP</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>7241</fpage>
          -
          <lpage>7259</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.-Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Artcap: A dataset for image captioning of fine art paintings</article-title>
          ,
          <source>IEEE Transactions on Computational Social Systems</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>N.</given-names>
            <surname>Garcia</surname>
          </string-name>
          , G. Vogiatzis,
          <article-title>How to read paintings: semantic art understanding with multimodal retrieval</article-title>
          ,
          <source>in: Proceedings of ECCV Workshops</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>P.</given-names>
            <surname>Achlioptas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ovsjanikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Haydarov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Elhoseiny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Guibas</surname>
          </string-name>
          ,
          <article-title>Artemis: Afective language for visual art</article-title>
          ,
          <source>in: Proceedings of IEEE/CVF CVPR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>11569</fpage>
          -
          <lpage>11579</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdelfattah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Alhuwaider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Church</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Elhoseiny</surname>
          </string-name>
          ,
          <article-title>Artelingo: A million emotion annotations of wikiart with emphasis on diversity over language and culture</article-title>
          ,
          <source>in: Proceedings of EMNLP</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>8770</fpage>
          -
          <lpage>8785</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>A. X.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Funkhouser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Guibas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hanrahan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Savarese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Savva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Su</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Shapenet</surname>
          </string-name>
          :
          <article-title>An information-rich 3d model repository</article-title>
          ,
          <source>arXiv preprint arXiv:1512.03012</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>M.</given-names>
            <surname>Deitke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Salvador</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Weihs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Michel</surname>
          </string-name>
          , E. VanderBilt, L. Schmidt,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ehsani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kembhavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <article-title>Objaverse: A universe of annotated 3d objects</article-title>
          ,
          <source>in: Proceedings of IEEE/CVF CVPR</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>13142</fpage>
          -
          <lpage>13153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Deitke</surname>
          </string-name>
          , R. Liu,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wallingford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ngo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kusupati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Laforte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Voleti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Y.</given-names>
            <surname>Gadre</surname>
          </string-name>
          , E. VanderBilt,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kembhavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Vondrick</surname>
          </string-name>
          , G. Gkioxari,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ehsani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <article-title>Objaverse-xl: A universe of 10m+ 3d objects</article-title>
          ,
          <source>arXiv preprint arXiv:2307.05663</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dawson</surname>
          </string-name>
          , et al.,
          <article-title>Data-driven consumer engagement, virtual immersive shopping experiences, and blockchain-based digital assets in the retail metaverse</article-title>
          ,
          <source>Journal of SelfGovernance and Management Economics</source>
          <volume>10</volume>
          (
          <year>2022</year>
          )
          <fpage>52</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>W.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , Vtonshoes:
          <article-title>Virtual try-on of shoes in augmented reality on a mobile device</article-title>
          , in: IEEE ISMAR,
          <year>2022</year>
          , pp.
          <fpage>234</fpage>
          -
          <lpage>242</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>H.</given-names>
            <surname>Laaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Miche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tammi</surname>
          </string-name>
          ,
          <article-title>Prototyping a digital twin for real time remote control over mobile networks: Application of remote surgery</article-title>
          ,
          <source>IEEE Access 7</source>
          (
          <year>2019</year>
          )
          <fpage>20325</fpage>
          -
          <lpage>20336</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Deen</surname>
          </string-name>
          ,
          <article-title>A novel cloud-based framework for the elderly healthcare services using digital twin</article-title>
          ,
          <source>IEEE Access 7</source>
          (
          <year>2019</year>
          )
          <fpage>49088</fpage>
          -
          <lpage>49101</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Farhat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chiementin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chaari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bolaers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Haddar</surname>
          </string-name>
          ,
          <article-title>Digital twin-driven machine learning: ball bearings fault severity classification</article-title>
          ,
          <source>Measurement Science and Technology</source>
          <volume>32</volume>
          (
          <year>2021</year>
          )
          <fpage>044006</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>A.</given-names>
            <surname>Siyaev</surname>
          </string-name>
          , G.-S. Jo,
          <article-title>Towards aircraft maintenance metaverse using speech interactions with virtual objects in mixed reality</article-title>
          ,
          <source>Sensors</source>
          <volume>21</volume>
          (
          <year>2021</year>
          )
          <year>2066</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. C.</given-names>
            <surname>Braud</surname>
          </string-name>
          , et al.,
          <article-title>Vr-driven museum opportunities: digitized archives in the age of the metaverse</article-title>
          ,
          <source>Artnodes</source>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>M. C. Longo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Faraci</surname>
          </string-name>
          ,
          <article-title>Next-generation museum: A metaverse journey into the culture</article-title>
          ,
          <source>Sinergie Italian Journal of Management</source>
          <volume>41</volume>
          (
          <year>2023</year>
          )
          <fpage>147</fpage>
          -
          <lpage>176</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Conde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Turgutlu</surname>
          </string-name>
          ,
          <article-title>Clip-art: Contrastive pre-training for fine-grained art classification</article-title>
          ,
          <source>in: Proceedings of IEEE/CVF CVPR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>3956</fpage>
          -
          <lpage>3960</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>G.</given-names>
            <surname>Castellano</surname>
          </string-name>
          , E. Lella, G. Vessio,
          <article-title>Visual link retrieval and knowledge discovery in painting datasets</article-title>
          ,
          <source>Multimedia Tools and Applications</source>
          <volume>80</volume>
          (
          <year>2021</year>
          )
          <fpage>6599</fpage>
          -
          <lpage>6616</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>L.</given-names>
            <surname>Aroyo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Stash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gorgels</surname>
          </string-name>
          , L. Rutledge,
          <article-title>Chip demonstrator: Semanticsdriven recommendations and museum tour generation</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2007</year>
          , pp.
          <fpage>879</fpage>
          -
          <lpage>886</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Stash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sambeek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Schuurmans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Aroyo</surname>
          </string-name>
          , G. Schreiber,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gorgels</surname>
          </string-name>
          ,
          <article-title>Cultivating personalized museum tours online and on-site,</article-title>
          <source>Interdisciplinary science reviews 34</source>
          (
          <year>2009</year>
          )
          <fpage>139</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Sarkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <article-title>A new point-of-interest approach based on multi-itinerary recommendation engine</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>181</volume>
          (
          <year>2021</year>
          )
          <fpage>115026</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>P.</given-names>
            <surname>Yochum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>An adaptive genetic algorithm for personalized</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>