<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Algorithm Comparison for Cultural Heritage Image Classi cation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Radmila Jankovic</string-name>
          <email>rjankovic@mi.sanu.ac.rs</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Mathematical Institute of the Serbian Academy of Sciences and Arts Belgrade</institution>
          ,
          <country country="RS">Serbia</country>
        </aff>
      </contrib-group>
      <fpage>26</fpage>
      <lpage>33</lpage>
      <abstract>
        <p>Digitization represents an important part of the development of online systems. As such it includes, among other, the deployment, categorization and preservation of audio, video and textual contents online. Such process is especially interesting from the perspective of cultural heritage, as it allows the long-term preservation and sharing of culture worldwide. This study observes four classi cation algorithms: (i) the multilayer perceptron, (ii) averaged one dependence estimators, (iii) forest by penalizing attributes, and (iv) the k-nearest neighbor rough sets and analogy based reasoning, before and after attribute classi cation, and compares these with the results obtained from the convolutional neural network. The obtained results show that the best classi cation performance was achieved by the multilayer perceptron, followed by the convolutional neural network.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>With an increased use of digitization in the domain of cultural heritage, it is possible to preserve and promote
the cultural heritage present in every part of the world. Every country worldwide has its own practices, places,
values, objects, and arts that are created throughout history, and that represent its cultural heritage. Through
cultural heritage knowledge is being shared and passed on from generation to generation. Some examples of
cultural heritage include photographs, historical monuments, various types of documents, archaeological sites,
and other.</p>
      <p>There are three pillars of digital cultural heritage: (i) digitization focusing on conversion of objects into digital
form, (ii) access to digital heritage, and (iii) long term preservation of digital objects [IDST12]. Classi cation
is an important part of digitization as it includes building a classi cation model that groups new inputs into
categories, based on the previously available set of data. In terms of cultural heritage, classi cation is particularly
important because it allows the preservation of heritage for future generations. Furthermore, digitization enables
promotion of cultural heritage by using innovative technologies to increase accessibility. Through digitization, a
country's cultural heritage is being promoted globally, thus contributing to cultural diversity.</p>
      <p>Di erent methods for cultural heritage image classi cation have been investigated by various authors. Deep
learning algorithms were used for image classi cation in [LlMMZG17]. In particular, AlexNet and Inception V3
Copyright c 2020 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC
BY 4.0).
convolutional neural networks (CNNs) were used, as well as ResNet and Inception-ResNet-v2 residual networks.
The dataset included architectural cultural heritage divided into 10 categories. The results showed that deep
learning methods perform better than other state-of-the-art methods, particularly when dealing with complex
problems [LlMMZG17]. The performance of deep learning methods has also been investigated in [KHP18], where
CNNs were used to classify images, audio and video data, while Recurrent Neural Networks (RNNs) were used to
classify the text belonging to the cultural heritage of Indonesia [KHP18]. It was observed that the RNN achieved
the highest accuracy, while the CNN obtained good accuracy for image and video classi cation (76%)[KHP18].
Considering other methods, k-nearest neighbor (kNN) classi cation was used to classify cultural heritage images
of 12 monuments and landmarks in Pisa [AFG15], but also to classify and detect alterations on historical buildings
with a high accuracy (92%) [MPAL15]. Various decision tree algorithms including J48, random tree, random
forest and fast random forest were investigated in [GDPR18]. Classi cation was performed in WEKA on a set
consisting of 3D cultural heritage models, and the results showed that the fast random forest achieves the highest
accuracy of 69% [GDPR18]. Di erent types of image classi cation techniques were investigated in [AJ19]. In
particular, naive Bayes and Support Vector Machine (SVM) algorithms are widely used for tangible and movable
cultural heritage, while the intangible cultural heritage is mostly classi ed using SVM, kNN, CNN, decision
trees and Conditional Random Fields - Gaussian Mixture Model (CRF-GMM) algorithms [AJ19]. Tangible and
immovable cultural heritage is most commonly classi ed using CNNs [AJ19].</p>
      <p>The aim of this paper is to compare the performance of several classi cation algorithms for cultural heritage
image classi cation, in particular: (i) the multilayer perceptron (MLP), (ii) averaged one dependence estimators
(AODE), (iii) forest by penalizing attributes (Forest PA), and (iv) the k-nearest neighbor rough sets and analogy
based reasoning (RSeslibKnn). The performance was observed on a full set of attributes as well as on the
reduced set of attributes, and the results were compared. Furthermore, a CNN was also developed for comparison
purposes, as deep learning represents a state-of-the-art technique [LlMMZG17, KHP18] and it is interesting to
observe and compare its performance with the performance of the other algorithms used in this study.</p>
      <p>This paper is organized as follows. Section 2 explains the data and methodology used in this research, while
Section 3 presents the results and discussion. Finally, Section 4 contains concluding remarks.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Data and methodology</title>
      <sec id="sec-2-1">
        <title>Data</title>
        <p>The cultural heritage image classi cation was performed on a public dataset created by [LlMMZG17] and obtained
from Datahub (https://old.datahub.io/dataset/architectural-heritage-elements-image-dataset).
The dataset consists of 10,235 images of size 128 128 pixels, but for the purpose of this study, 4,000
images from 5 out of 10 categories were randomly chosen. These images include altars, gargoyles, domes, columns,
and vaults (Figure 1).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Methodology</title>
        <p>The experiments were performed in WEKA (Waikato Environment for Knowledge Analysis) [WFHP16], a free
data mining software based on Java, while the CNN model was developed in Python v.3.7 with the use of the
Keras library. The experiments were performed on a Windows machine with a 2.3 GHz processor and 8 GB of
RAM.</p>
        <p>In Weka, the feature extraction is performed using feature extraction algorithms integrated inside the
imageFilters package. Three types of lters were applied on the dataset: (i) the edge histogram, (ii) the color
layout, and (iii) the JPEG coe cients. Feature extraction for the CNN was not performed manually, as Python
automatically scans and extracts features from the dataset.</p>
        <p>The edge histogram extracts the MPEG7 edge features from the images. In particular, it detects the directions
of edges in images based on the changes in frequency and brightness [WPP02]. There are ve types of edges:
vertical, horizontal, 45-degree diagonal, 135-degree diagonal, and non-directional [WPP02]. The color layout
feature extraction is performed by dividing the image into 64 blocks and calculating the average color for each
block, using the color layout lter in WEKA. Finally, the JPEG coe cients were extracted by splitting the image
based on di erent frequencies, hence keeping only the most important frequencies [MD13].</p>
        <p>After feature extraction, the dataset consisted of 307 attributes. The rst results were generated on the full
set of attributes, while the second results were generated on the attribute-reduced set of data in order to evaluate
the change in performance before and after attribute selection. Feature selection was performed in WEKA using
the AttributeSelection lter. The attribute search was performed using the best- rst method, and attribute
evaluation using the CFS subset evaluator. After feature selection, the number of attributes was reduced to 89.
The dataset was divided into 70% of images for training and 30% for testing the algorithms. Four algorithms
were tested and compared for the purpose of image classi cation: (i) MLP, (ii) Forest PA, (iii) AODE, and (iv)
RSeslibKnn. Additionally, in order to compare the performance of these algorithms with the state-of-the-art
techniques, a deep learning CNN model was developed and tested.</p>
        <p>The MLP network consists of three layers: an input layer, a hidden layer, and an output layer. The layers
consist of neurons, where the neurons in one layer are connected to the neurons in the next layer [PM19].
The MLP uses a nonlinear activation function, making it suitable for di erent types of problems, without any
assumptions regarding data distribution [GD98]. The AODE is a classi cation technique that calculates the
probability of each class and creates a set of one dependence probability distribution estimators [WBW05]. Is
holds less rigorous independence assumptions than naive Bayes, hence it is suitable for various range of problems.
Forest PA is a relatively new decision forest algorithm developed in 2017 by [AI17] in order to overcome the
limitations of the random forest algorithm. The algorithm works in such a way that it uses the full set of
attributes for forest creation, but it also assigns weights to those attributes that already participated in the
previous decision tree. It generates the bootstrap sample from the training set and creates a decision tree using
the attribute weights [AI17]. RSeslibKnn is the k-nearest neighbor classi er that uses a fast neighbor search,
thus making it appropriate for using on large datasets [WL19]. A distance measure is calculated based on the
weighted sum of distances, which results in creation of the indexing tree [WL19]. The classi cation is performed
by nding the k nearest neighbors in the set. The CNN is a neural network consisting of standard type of layers,
as well as of convolution and pooling layers. It is widely used for image processing, as it is able to automatically
scan and extract features from images. In particular, each CNN is composed of convolutional layers followed
by the pooling layers that reduce the dimensionality of the data. The features extracted through these layers
are then transformed into a vector using the attening layer, and the obtained vector is further distributed to a
dense layer, thus forming a fully-connected network.</p>
        <sec id="sec-2-2-1">
          <title>Parameter con guration</title>
          <p>The MLP consisted of one hidden layer with 50 neurons, a learning rate of 0.5, a momentum of 0.6, and a
batch size of 32. The sigmoid activation function was utilized in the hidden and output layers. All attributes
were standardized. The RseslibKnn included the the city and simple value di erence as a distance measure,
the inverse square distance as the voting method, and the distance based weighting method. The AODE was
con gured with the default WEKA parameters. Lastly, the Forest PA included 30 trees in the forest.</p>
          <p>The CNN con guration involved four convolution layers with 32 neurons in the rst two layers, 64 neurons
in the last two convolution layers, one hidden layer with 128 neurons, and an output layer with 5 neurons. The
convolutional and hidden layers used the hyperbolic tangent activation function, while the output layer used
the softmax activation. The kernel size of the convolutional layers was set to 3 3, while the pooling size in
the pooling layer was set to 2 2. The dropout was set to 0.2. Lastly, in order to avoid over tting, the early
stopping regularization parameter was used with patience set to 3. The number of epochs was set to 50, but the
training stopped after 18 epochs because there was no further improvement in the accuracy. The model used
80% of data for training and the rest of the data for validation.</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Evaluation Metrics</title>
          <p>The classi cation performance can be evaluated using several metrics. In particular, this study observed and
analyzed the values of correctly classi ed instances, precision, recall, F-score, kappa statistics, and ROC area. The
percentage of correctly classi ed instances represents the instances that are correctly classi ed by the algorithm,
while precision shows the fraction of instances that belong to the observed class among the total number of
instances that are classi ed into the observed class by the algorithm. Recall represents the true positive rate of
prediction, while the F-measure shows the classi cation accuracy based on the average value of precision and
recall. The F-measure values should ideally be closer to 1 indicating a better classi cation accuracy. Kappa is the
measure of agreement and can have values in an interval of 0{1, with the values in the range 0.81{1 representing
an almost perfect agreement, while values close to 0 represent poor agreement [SW05]. Lastly, the ROC area
represents the ratio of the true positives and the false positives, and its value should be close to 1, indicating a
perfect prediction [FUW06].
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>All algorithms used in this study were rst tested on the full dataset consisting of 307 attributes. The best
performing algorithm in this case is the MLP with 85% of accuracy obtained, followed by the RSeslibKnn,
AODE, and Forest PA with 82%, 79% and 78%, respectively (Table 1).</p>
      <p>The MLP also obtained the best values in terms of the kappa statistics, precision, recall, and the F-measure,
comparing to other three algorithms (Table 1). In terms of the running time, the fastest algorithm is the Forest
PA with 54.02 seconds.</p>
      <p>After observing the results obtained from the full set of data, the next step involved observing the
performance of the algorithms on the reduced set of attributes. The MLP algorithm again performed the best, correctly
classifying 98.9% of instances (Table 2). Other algorithms obtained lower classi cation accuracy, in particular
80.83%, 80.67% and 78.67% for the AODE, RSeslibKnn and Forest PA, respectively. Observing other
performance measures, the MLP also obtained the highest value of kappa statistics (0.986), followed by AODE (0.760),
RSeslibKnn (0.758), and Forest PA (0.733). As described before in this paper, these values indicate a substantial
to almost perfect agreement [SW05]. Moreover, the MLP obtained the highest value of the F-measure (0.986),
followed by RSeslibKnn (0.811), AODE (0.807), and Forest PA (0.797). Lastly, observing the value of the ROC
area, the results indicate the MLP algorithm performed the best with an obtained value of 0.996, followed by
AODE, Forest PA and RSeslibKnn with ROC area values of 0.965, 0.959 and 0.879, respectively. These results
show a good classi cation power of the observed algorithms, with MLP and AODE performing the best (Table
2), but it should be noted that the MLP algorithm requires much longer running time than the other three
algorithms.</p>
      <p>In order to gain more insights about the power of the observed classi cation algorithms, classi cation
matrices are generated (Table 3). The classi cation matrix shows the number of correctly (and incorrectly) classi ed
instances by class, where the numbers in the diagonal represent accurate classi cations. Before attribute
selection, the algorithms mostly miss-classi ed images of gargoyle, column and vault. In particular, the MLP most
accurately classi ed the dome images, while the highest number of miss-classi cations for the MLP algorithm
is observed for the vault and column images. The AODE most accurately classi ed altar images, while Forest
PA most correctly classi ed the dome images. The RSeslibKnn classi ed altar images most accurately, while the
highest number of wrongly classi ed instances is observed mainly for the column images (Table 3).</p>
      <p>After attribute selection has been applied, the new confusion matrix has been obtained (Table 4). In terms
of miss-classi cations, the MLP and AODE algorithms mostly miss-classi ed column images, the RSeslibKnn
mostly miss-classi ed vault images, while Forest PA mostly miss-classi ed the images of vaults and columns. In
terms of accurate classi cations, the MLP accurately classi ed almost all the images, while the AODE, Forest
PA and RSeslibKnn most accurately classi ed the images of altar and dome (Table 4).</p>
      <p>The previously described results were compared to the results obtained by using a deep learning algorithm, as
it represents the state-of-the-art methodology. For this purpose, the CNN model was developed in Python and the
Algorithm
MLP
AODE
Forest PA
RSeslibKnn
loss and accuracy results through epochs are plotted and presented in Figure 2. The obtained results demonstrate
good accuracy of 93%, with training loss of 0.21 and validation loss of 0.42. Such results are promising as they
clearly demonstrate the potential of deep learning techniques. Furthermore, deep learning allows for detailed
modi cations, thus enhancing the possibility of their application to di erent types of problems.
(a)</p>
      <p>(b)</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>The aim of this paper was to compare the performance of several classi cation algorithms before and after
attribute selection. In particular, MLP, AODE, Forest PA, and RSeslibKnn algorithms were applied on the
dataset consisting of cultural heritage images, and their performance was compared to the performance obtained
by the deep learning algorithm (CNN). Several conclusions can be drawn from this study: (i) the MLP algorithm
obtained the best performance both before and after attribute selection, (ii) attribute selection increases the
classi cation accuracy for all algorithms except the RSeslibKnn, and (iii) the deep learning techniques such as
the CNN, obtain higher accuracy without needing to reduce the number of attributes. Hence, as deep learning
techniques do not require to manually extract features from images, they represent an appropriate method of
image classi cation.</p>
      <sec id="sec-4-1">
        <title>Acknowledgements</title>
        <p>This work was supported by the Serbian Ministry of Education, Science and Technological Development through
Mathematical Institute of the Serbian Academy of Sciences and Arts.
[IDST12]</p>
        <p>Ivanova, K., Dobreva, M., Stanchev, P. and Totkov, G. Access to digital cultural heritage:
Innovative applications of automated metadata generation. Plovdiv University Publishing House "Paisii
Hilendarski", 2012.</p>
        <p>Sim, J. and Wright, C.C. "The kappa statistic in reliability studies: use, interpretation, and sample
size requirements." Physical therapy 85, no. 3 (2005): 257-268.
[LlMMZG17] Llamas, J., Lerones, M. P., Medina, R., Zalama, E. and Gomez-Garc a-Bermejo, J. "Classi cation
of architectural heritage images using deep learning techniques." Applied Sciences 7, no. 10 (2017):
992.
osovi, M., Amelio, A. and Junuz, E. Classi cation Methods in Cultural heritage. In Proceedings of
the 1st International Workshop on Visual Pattern Extraction and Recognition for Cultural Heritage
Understanding co-located with 15th Italian Research Conference on Digital Libraries (IRCDL, pp.
13-24, 2019.</p>
        <p>Won, C.S., Park, D.K. and Park, S.J. E cient Use of MPEG-7 Edge Histogram Descriptor. ETRI
journal 24, no. 1 (2002): 23-30.</p>
        <p>More, N.K. and Dubey, S. JPEG Picture Compression Using Discrete Cosine Transform.
International Journal of Science and Research (IJSR) 2, no. 1 (2013): 134-138.</p>
        <p>Witten, I.H., Frank, E., Hall, M.A. and Pal, C.J. Data Mining: Practical machine learning tools
and techniques. Morgan Kaufmann, 2016.</p>
        <p>Pal, S.K. and Mitra, S. "Multilayer perceptron, fuzzy sets, and classi cation." IEEE Transactions
on neural networks 3 (1992): 683{697.</p>
        <p>Gardner, M.W. and Dorling, S. "Arti cial neural networks (the multilayer perceptron)a review of
applications in the atmospheric sciences." Atmospheric environment 32, no. 14-15 (1998):
26272636.</p>
        <p>Webb, G.I., Boughton, J.R. and Wang, Z. "Not so naive Bayes: aggregating one-dependence
estimators." Machine learning 58, no. 1 (2005): 5-24.
[FUW06]</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [WL19]
          <string-name>
            <surname>Adnan</surname>
            ,
            <given-names>M.N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Islam</surname>
            ,
            <given-names>M.Z.</given-names>
          </string-name>
          "
          <string-name>
            <surname>Forest</surname>
            <given-names>PA</given-names>
          </string-name>
          :
          <article-title>Constructing a decision forest by penalizing attributes used in previous trees</article-title>
          .
          <source>" Expert Systems with Applications</source>
          <volume>89</volume>
          (
          <year>2017</year>
          ):
          <fpage>389</fpage>
          -
          <lpage>403</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Wojna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Latkowski</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Rseslib 3: Library of rough set and machine learning methods with extensible architecture</article-title>
          .
          <source>In Transactions on Rough Sets XXI</source>
          , pp.
          <volume>301</volume>
          {
          <fpage>323</fpage>
          . Springer,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Upadhye</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Worster</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>"Understanding receiver operating characteristic (ROC) curves."</article-title>
          <source>Canadian Journal of Emergency Medicine</source>
          <volume>8</volume>
          , no.
          <issue>1</issue>
          (
          <year>2006</year>
          ):
          <fpage>19</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>