<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop on Computational Methods in the Humanities, June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Exploring Naming Inventories for Architectural Elements for Use in Multi-modal Machine Learning Applications⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ronja Utescher</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aaron Pattee</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ferdinand Maiwald</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jonas Bruschke</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stephan Hoppe</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sander Münster</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Florian Niebling</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sina Zarrieß</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bielefeld University</institution>
          ,
          <addr-line>Universitätsstraße 25, 33615 Bielefeld</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Friedrich-Schiller-University Jena</institution>
          ,
          <addr-line>Fürstengraben 1, 07743 Jena</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Julius-Maximilians-University of Würzburg</institution>
          ,
          <addr-line>Sanderring 2, 97070 Würzburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Ludwig-Maximilians-Universität München</institution>
          ,
          <addr-line>Geschwister-Scholl-Platz 1, 80539 Munich</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>0</volume>
      <fpage>9</fpage>
      <lpage>10</lpage>
      <abstract>
        <p>Computer vision models are increasingly relevant and useful to Digital History. Next to the increasingly complex neural models, data and data selection are an integral part of this process. In this paper, we examine and extend the data collection practices from a major recent paper in the domain of architectural element classification. We collected an image-text data set for a selection of 56 Baroque landmarks to be analysed in like manner. This diferent architectural domain yielded insights into the transferability of the original model and data collection procedures. Notably, the architectural domain also has an impact on the availability of classes of architectural elements as well as the performance of the models classifying them.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Architecture</kwd>
        <kwd>Art History</kwd>
        <kwd>Computer Vision</kwd>
        <kwd>Machine Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The study of architectural art history has greatly benefited from innovative, computer-aided
approaches in recent years. From high-resolution two-dimensional (2D) photos of building
edifices, to three-dimensional (3D) models of entire structures, these emerging techniques
are laying the foundation for new methodologies in researching architecture [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Provided
the three-dimensional nature of buildings, research projects have appropriately focused on
techniques that produce digital replicas of their forms, such as Structure-from-Motion (SfM)
Photogrammetry and Terrestrial Laser Scanning (TLS) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, one critical aspect in
the study of architecture has largely been dormant since the emergence of these technologies,
namely, the computer-aided description of architectural elements. 2D images and 3D models
were the logical first steps in devising new methodologies for identifying architectural elements,
providing precise calculations of their dimensions and structures, buttressed by the unique
capability to virtually study a building. What normally follows is a traditional text description
of the discoveries achieved by the employment of these techniques, relegating these digital
applications as mere means to a more enlightened end. The lamentable result is a digital
purgatory of 3D models awaiting their fate in repositories or online databases.
      </p>
      <p>
        This paper presents one aspect of a larger project seeking to utilise 2D images and 3D models
as essential components of a search engine for architectural elements. These digital objects can
serve as reference points for future research in architectural art history and archaeology. What
is required is a systematic identification of the elements themselves using text descriptors, in
order that the digital representations of the elements can be eficiently explored. In many ways,
this avenue of research is an evolution of [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which had expert and non-expert annotators
choose phrases from longer textual descriptions with which to index paintings in a collection.
For this purpose, we implement the methodology of recent research [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] as the foundation of
the link between text and digital representation. The work by [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] represents an important step
in integrating 3D models, collections of photographs, and written descriptions of a historic
building using state-of-the-art machine learning methods. Given a multi-modal collection of
cathedrals, the model learns to detect and classify 10 classes of architectural objects, for example
portals and columns. These classes of objects are however limited by which terms are frequently
associated with images in the original source.
      </p>
      <p>
        In this paper, we take a closer look at the first two steps in their workflow [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]; (1) curating
images and their descriptions from large, open-source collections and (2) selecting the vocabulary
of architectural elements for the machine-learning model to classify. We argue that these are
important design decisions that have ramifications for the output of the model down the line.
Both [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and the authors of this paper source their images from Wikimedia Commons. While
Commons is a free and abundant source of images, the indexing and naming it provides for
individual images is comparatively limited and unsystematic.
      </p>
      <p>
        Our case study, similar to [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], is abstract as it is neither limited to a specific building, nor a
design implemented by a specific architect. Rather, it is a collection of baroque monumental
buildings largely built between the late 17th and mid-18th centuries, ranging geographically
from Portugal to Russia, by a large network of architects. In efect, we construct a new collection
which parallels the one introduced in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The majority of the collection consists of 2D images,
though this could be supplemented in future research by existing high-resolution TLS models
of historic buildings in the German city of Dresden, such as the iconic Zwinger.
      </p>
      <p>
        For the purposes of this paper, we define a domain according to architectural style. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] do
not explicitly frame the issue this way, but we will briefly talk about the issue here since it is
important from a historical and architectural perspective. Wikiscenes and WikiscenesBaroque
difer both in architectural style (gothic/baroque) and building function (house of
worship/residence of high-ranking dignitaries). [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] define it more so by function, although the Cathedrals
lean towards the Gothic style. The paper is structured as follows: In section 2, we discuss the
method of the paper, including the Wikiscenes data curation policy and classification model
as well as our modifications to said policy. In section 3, we examine the resulting new dataset,
WikiscenesBaroque. Section 4 details the setup and results of the experiments we conducted in
order to assess the influence of this diferent data on the classification model. Finally, Section 5
summarises the diferent lessons learnt from the case study.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>Our goal is to build a machine learning-based system that labels architectural elements in
images of buildings, of particular architectural styles assumed in the study of architectural art</p>
      <p>Category:Hofburg
Category:Heldenplatz, Vienna</p>
      <p>
        Category:Leopoldinischer Trakt
Description: The northwestern facade of the Leopold Wing of the Hofburg Imperial Palace in Vienna.
history. In Section 2.1, we take a brief look at the uses and requirements of machine learning
methods in art history. Sections 2.2 and 2.3 describe [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]’s and our approach respectively.
      </p>
      <sec id="sec-2-1">
        <title>2.1. Image Classification for Architectural Categories</title>
        <p>The documentation of historical built works often goes hand in hand with large collections of
image materials, especially for landmarks which are popular objects of study. Machine Learning
(ML) methods have the potential to greatly benefit research of these historical landmarks, since
they allow the automatic processing of large amounts of data. These ML applications can take
the form of labelling data according to predefined categories, or in interactive settings like
Image Retrieval. This paper focuses on modelling and data collection in classification settings.</p>
        <p>
          Image and Object Classification models have undergone significant development in the last
10 years. Although there have been a number of new models for Image Classification and other
Vision tasks, most models in image processing use a convolutional neural network (CNN) as
a backbone (cf [
          <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
          ]. Image classification models generally work with a set classification
vocabulary, requiring labelled training data. Datasets in digital history are smaller and more
dificult to obtain. There are standard vocabularies for categorising architectural elements, but
these are not connected to the datasets. Domain specificity in classification models runs the
gamut from models trained on generic vocabularies [
          <xref ref-type="bibr" rid="ref10 ref11 ref9">9, 10, 11</xref>
          ], to domain-specifically trained or
ifne-tuned models, to models trained on specific landmarks. We use an approach that does not
rely on NLP, but on language in terms of the vocabulary which we use to talk about real-world
objects and their visual representations.
        </p>
        <p>
          The method for creating these reconstructions is unsupervised except for the previously
mentioned choice in input images for each model. COLMAP also creates a match between
the 3D space of the model and the 2D images. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] exploited these matches in their image
classification and segmentation, using a loss function which rewards the model for consistent
classification of points in the landmark across diferent images.
2.2. Landmarks and Image Categories in Wikiscenes
[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] use a combination of automatic data collection and manual refinement which utilises noisy,
but freely available data. The authors mine this data for semantic concepts which act as classes
for classification and segmentation models based on [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. This approach bypasses the need for
costly manual annotation and exploits the inherent structure of the data.
        </p>
        <p>
          [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] draw upon Wikimedia Commons as the source of image-and-text data. Any interested
party can contribute images, add them to the appropriate Wikimedia Commons (WM) category.
This serves as an alternative to other annotation paradigms, such as expert annotation or
crowd-sourced annotation. There are ways in which it is noisy; if users utilise the caption
and description fields at all, the content and length tends to vary wildly. These user-based
annotations are more sophisticated than what could reasonably be collected by laypeople in
a crowdworking setting. In this paper, we aim to go into detail about the curation process of
a subset of Wikimedia Commons for a number of Baroque Architectural Objects. Wikimedia
Commons provides us with a large number of user-uploaded images with open-source licensing.
Users are given the opportunity to annotate their images with captions, descriptions, and a
selection of other metadata such as the geolocation and camera specifications.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.3. AE classes in WikiScenes</title>
        <p>
          The classes in the original dataset are decided upon bottom-up; if there is an architectural
element that has a significant number of instances available for training, it is selected as one of
the classes for the model. In other words, [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] compute the frequency of all terms from categories,
descriptions and captions and manually select a number of architectural element classes from
the most frequent terms. Figure 1 showcases examples of two of these classes, facade and tower.
        </p>
        <p>
          [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] use image descriptions as well as Wikimedia Commons categories. On Wikimedia
Commons, users can add descriptions or captions to their images. However, only a small portion
of images have captions/descriptions. This makes Category pages the most comprehensive
source for text describing images. The images of said landmark are organised into subcategories,
sub-subcategories, and so on. Figure 2 shows an example of a Wikimedia Commons Image and
its category tree. In this example, the facade label can be sourced from the image description.
As described in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], the Wikimedia commons categories themselves are a significant source for
concept terms; if a term is used in a category name, all direct member images of the category
can be assigned the term as an architectural element (AE) class.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>
        The original WikiScenes dataset [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] contains data for 99 Gothic cathedrals. We build an
analogous dataset of 56 Baroque landmarks, mostly palaces. We investigate how [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] approach
transfers to modelling architectural elements in a diferent domain.
3.1. Curating Landmarks and Image Categories
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] build two point cloud models per landmark; one for the outside and one for the inside of
the landmark. The inside model is computed from images from the "Interior of [landmark]"
category and its leaf categories, white the outside model uses the "Exterior of [landmark]" and
"Views of [landmark]" categories. These three subcategories are present throughout the original
dataset’s landmarks, however we did not find this to be universal in our collection of Baroque
landmarks. Out of the 56 landmarks, 22 only had one subcategory that matched the original
inside/outside selection process. Overall, there were 38 "interior" categories, and 23 "exterior"
(14 + 9 "view").
      </p>
      <p>We selected 56 palaces based upon their architectural similarities to the Zwinger in Dresden,
in which all exhibit building phases in the late 17th and early 18th centuries. Additionally, the
architects and designers of the selection of buildings all had connections as part of a larger
network in which ideas and designs were shared, as many of the major construction eforts
occurred between 1700 and 1730 A.D. It was during this time that the Great Northern War
raged in which the Baltic and Black Seas, as well as the entire east of Europe served as the
battleground for the various armies. Many palatial architects were active military oficers
involved in constructing bastions and fortresses for and against artillery pieces. As a result,
there was a major exchange of ideas and designs during this period which even translated into
the construction plans of palaces.</p>
      <p>
        Each landmark has its own hierarchy of categories in Wikimedia Commons. We use the set
of all immediate subcategories - the landmark’s category being the root - as basis for coming up
with a list of terms to blacklist. For each of the 432 categories in this manual selection, we select
all its subcategories as well (unless they are in the blacklist). We use the blacklist sparingly, i.e.
including as many images as sensible in this step of the annotation. The blacklist is stamp, in art,
painting, collections, plans, aerial, panoramic, plans, signs, maps, things, history, events, leading to
the exclusion of 53 of 432 subcategories. This list is aimed at excluding objects which are not
part of the landmark, but located in it (collections, paintings, signs) or associated with it (in art,
stamp, things). Also, we exclude non-photographic images (plans, maps) and images with an
atypical perspective (aerial, panoramic). Finally, the last two keywords in the blacklist exclude
historical images and images depicting certain events This allows us to utilise a larger set of
images compared to only collecting the images uploaded in the "Exterior", "Interior" and "Views
of" categories This initial selection yields 85,629 images, while using [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]’s method would yield
28,794 images. From these images, we select instances of the AE classes for our cross-domain
test set as well as our data for the Baroque-specific model (cf. Section 4).
      </p>
      <sec id="sec-3-1">
        <title>3.2. Selecting AE Classes</title>
        <p>
          We followed [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and selected 10 AE classes from the most frequent terms present in the WM
Categories and Descriptions. This led to a diferent set of classes than the original dataset, as
shown in Table 3. Out of the 10 AE classes, only 3 (facade, statue, tower) are present in both
models. Out of the available images, we were able to link 17115 with AE classes for training the
model described in the new domain setting (Section 4.2).
        </p>
        <p>
          In order to evaluate [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]’s original model, we also construct an alternate test set using instances
of AE classes from the original Wikiscenes dataset. As listed in Table 2, the availability of
instances in our Baroque domain is mixed. The classes nave and choir are entirely absent, while
the data contains comparatively many statues and facades.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>
        4.1. Model
In this section, we introduce the image classification model without 3D Loss from [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and use it
to perform experiments in two novel classification setups.
      </p>
      <p>
        The model described in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] performs two general steps. 1) For each landmark, generate a 3D
model from photographs and 2) assemble an architectural vocabulary for identifying elements
of the landmark. The [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] model is based on [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The implementation uses a resnet50 backbone
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] with an ImageNet [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] pretraining. Their model’s backbone extracts both low-level and
high-level image features, which are unified in a pooling layer followed by a Global Cue Injection
(GCI) module. In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the 3D loss provides a small performance boost (3.3% across all images).
In this paper, we generally forego the use of 3D Loss as proposed by [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We argue that using
the baseline classification model is adequate to judge the overall feasibility, since 3D does not
fundamentally change the results, instead giving a small boost in accuracy overall.
      </p>
      <sec id="sec-4-1">
        <title>4.2. Classification Setups</title>
        <p>
          We perform experiments in two classification setups. The baroque domain setting is analogous
to [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]; the cross-domain setting tests the performance of the original model in images from a
diferent architectural style.
        </p>
        <p>
          Baroque Domain In the Baroque Domain task, we train a new image classification model
without 3D loss, using our collection of Baroque Wikimedia Commons images. We use a batch
size of 16 and 10 epochs for training the model. Like [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], we use a 9:1 split on landmark level,
so that the images seen at test time belong to unseen landmarks.
Cross-Domain In this experiment, we evaluate the performance of the original model by [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
on images of the landmarks in our data. We refer to this as the cross-domain setup because a
model trained on Gothic cathedrals is used to classify images of Baroque palaces. We construct
a cross-test set with instances from our 56 Baroque landmarks. In Table 2, we provide statistics
on the distribution of classes in our data. Some classes had no instances or very few in the
Baroque dataset, likely due to the domain itself. For the very common class facade, we randomly
sampled 1500 instances for the test set. Analogously to [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], we evaluate by calculating the
precision per class, as well as the mean average precision (mAP) over all instances and the mAP
over all landmarks (mAP*).
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. Results</title>
        <p>Baroque Domain Figure 3 shows some examples of correct and incorrect judgements of
the model. Note that while (a) has the gold label fountain and there is a fountain in the right
foreground of the image, most of the space in the image is taken up by the facade of the building,
which is the predicted class of the model.</p>
        <p>In Table 3, we report the precision of the model trained on WikiscenesBaroque (WSB) for
each AE class. Similarly to Wikiscenes-baseline model (cf. Table 4), this model achieves a high
precision for the class facade, but also for hall and gallery. Even though the WSB model is
tested on unseen landmarks, its precision on statues is higher than the Wikiscenes baseline
model (45.8% vs. 33.8%). The model performs worst for tower at 30.2%, and under 50% for statue
(45.8%), fountain 45.6% and stair (47.3%).</p>
        <p>
          Cross-domain In Table 4, we report the mean average precision (mAP), for each AE class as
well as across all images. In [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], the baseline model achieves a higher overall mAP on known
landmarks (WS-K) than on images of unknown landmarks (WS-U) with 70.8% vs. 48.3%. The
mAP for the cross-test set is lower than for the WS-U test set, however the diference is much
smaller (48.3% vs. 44.2%). In the cross-test set, the model performs worst for the chapel AE at
20.7%, which is substantially lower than the 60.2% mAP in the WS-K set, but higher than the
10.7% in the WS-U set. On the other hand, the baseline model performs consistently well on
the facade class. For the tower and portal classes, the model performance is very similar in the
cross-test and WS-U sets.
(a) Gold label: fountain, (b) Gold label: portal,
predicted: facade predicted: tower
(c) Gold label: portal,
predicted: portal
(d) Gold label: statue,
predicted: statue
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>
        The automatic annotation of basic architectural elements is beneficial to the digital
documentation of Historical Heritage Sites. Architectural classification models like [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] are adept at
utilising large amounts of already existing, noisily annotated data, taking concepts from the
text annotations and locating them in a 3D model of the landmarks. Approaches like these also
require the attention of the researcher handling the model. In this paper, we have examined the
various stages of data curation and their efect on the model’s performance.
      </p>
      <p>For more specific vocabularies, the text provided in Wikimedia Commons is not a suitable
source for labelled data as there are comparatively few instances which are directly annotated.
Training a model with more classes and more diverse training data is possible in principle.
However, domain-specific models have the advantage of needing less computing power.
Additionally, domain experts can advise during the data curation process and steer the model’s AE
classes towards what is of interest to them.</p>
      <p>The results of our experiments suggest that classification and detection of architectural
elements works best within domain, with a moderate gap to the cross-domain performance.
Images of known landmarks which are seen during training appear to be overall easier to classify,
comparing the WS-K to the WS-U and cross-test sets as listed in Table 4. Architectural elements
and their visual styles can be highly domain-specific, or even very individualised to particular
buildings. For well-documented landmarks, it may well be feasible to have classification models
solely for a particular landmark. However, less-resourced landmarks can still benefit from more
general models.</p>
      <p>
        Besides under-resourced landmarks, we want to point out the issue of more fine-grained
AE classes as a topic for future research. The AE classes which can be mined from resources
like Wikimedia Commons are useful for basic classification of images and segmentation of 3D
models. There are a great number of distinctions of AE classes within the study of architectural
art history, many of which are catalogued in resources like the Art &amp; Architecture Thesaurus
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Building Machine Learning models which can identify more elaborate concepts would
open up more detailed analysis and documentation of the landmarks and would be an interesting
direction for future research.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>
        In summary, this paper as well as [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] suggest that the practice of using online image collections,
in particular Wikimedia Commons, yields data that can be used to train models which in turn
classify architectural elements. We curate our own dataset, focusing on Baroque landmarks
that were constructed in the late 17th and early 18th centuries. The experiments in Section 4
suggest that in our case study, model precision decreases for images of unknown landmarks,
especially if they come from a stylistically diferent domain.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>This work was supported by a grant from the Federal Ministry of Education and Research
(BMBF, grant No. 01UG2120).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.-E.</given-names>
            <surname>Lutteroth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hoppe</surname>
          </string-name>
          ,
          <source>Schloss Friedrichstein 2</source>
          .
          <article-title>0: von digitalen 3D-Modellen und dem Spinnen eines semantischen Graphen, in: Computing art reader: Einführung in die digitale</article-title>
          <source>Kunstgeschichte</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>184</fpage>
          -
          <lpage>198</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sapirstein</surname>
          </string-name>
          ,
          <article-title>Accurate measurement with photogrammetry at large sites</article-title>
          ,
          <source>Journal of Archaeological Science</source>
          <volume>66</volume>
          (
          <year>2016</year>
          )
          <fpage>137</fpage>
          -
          <lpage>145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Lercari</surname>
          </string-name>
          ,
          <article-title>Terrestrial laser scanning in the age of sensing, Digital methods and remote sensing in archaeology (</article-title>
          <year>2016</year>
          )
          <fpage>3</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Passonneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lippincott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Klavans</surname>
          </string-name>
          ,
          <article-title>Relation between agreement measures on human labeling and machine learning performance: results from an art history domain</article-title>
          ,
          <source>in: Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Averbuch-Elor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Snavely</surname>
          </string-name>
          ,
          <article-title>Towers of babel: Combining images, language, and 3d geometry for learning multimodal vision</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>428</fpage>
          -
          <lpage>437</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>Very deep convolutional networks for large-scale image recognition, 2014</article-title>
          . URL: https://arxiv.org/abs/1409.1556. doi:
          <volume>10</volume>
          .48550/ARXIV.1409.1556.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <year>2015</year>
          . URL: https://arxiv.org/abs/1512.03385. doi:
          <volume>10</volume>
          .48550/ARXIV.1512.03385.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Krueger</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <year>2021</year>
          . URL: https://arxiv.org/abs/2103.00020. doi:
          <volume>10</volume>
          .48550/ ARXIV.2103.00020.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.-Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Belongie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hays</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Perona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ramanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Zitnick</surname>
          </string-name>
          ,
          <string-name>
            <surname>Microsoft</surname>
            <given-names>COCO</given-names>
          </string-name>
          :
          <article-title>Common objects in context</article-title>
          ,
          <source>in: European conference on computer vision</source>
          , Springer,
          <year>2014</year>
          , pp.
          <fpage>740</fpage>
          -
          <lpage>755</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Caesar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uijlings</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          ,
          <article-title>COCO-stuf: Thing and stuf classes in context</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Redmon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          , Yolo9000: Better, faster, stronger,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Araslanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <article-title>Single-stage semantic segmentation from image labels</article-title>
          ,
          <source>in: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>4253</fpage>
          -
          <lpage>4262</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>N.</given-names>
            <surname>Araslanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <article-title>Single-stage semantic segmentation from image labels</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>4253</fpage>
          -
          <lpage>4262</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <year>2015</year>
          . URL: https://arxiv.org/abs/1512.03385. doi:
          <volume>10</volume>
          .48550/ARXIV.1512.03385.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <article-title>ImageNet: A Large-Scale Hierarchical Image Database</article-title>
          ,
          <source>in: CVPR09</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>P.</given-names>
            <surname>Harpring</surname>
          </string-name>
          ,
          <article-title>Development of the Getty vocabularies: AAT, TGN, ULAN, and</article-title>
          <string-name>
            <surname>CONA</surname>
          </string-name>
          ,
          <string-name>
            <surname>Art</surname>
            <given-names>Documentation</given-names>
          </string-name>
          :
          <source>Journal of the Art Libraries Society of North America</source>
          <volume>29</volume>
          (
          <year>2010</year>
          )
          <fpage>67</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>