<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Classification Methods in Cultural Heritage</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>C´ osovi´</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Junuz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dzemal Bijedic University, Faculty of Information Technology</institution>
          ,
          <addr-line>Sjeverni logor 12, 88000 Mostar</addr-line>
          ,
          <country country="BA">Bosnia and Herzegovina</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Calabria, DIMES</institution>
          ,
          <addr-line>Via Pietro Bucci 44, 87036 Rende (CS)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of East Sarajevo, Faculty of Electrical Engineering</institution>
          ,
          <addr-line>Vuka Karadˇzi ́ca 30, 71126 Lukavica, East Sarajevo</addr-line>
          ,
          <country country="BA">Bosnia and Herzegovina</country>
        </aff>
      </contrib-group>
      <fpage>13</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>This paper describes relevant classification methods applied to the cultural heritage context. In particular, a categorisation of the classification methods is provided according to tangible and intangible cultural heritage, where movable and immovable objects can be in the focus. A short description of each method is reported for each cultural heritage category in terms of feature representation, classification approach and obtained results. The proposed survey can be useful in the research community of pattern recognition and visual computing for exploring the current literature about the topic. It will hopefully provide new insights for the advancement of knowledge discovery in cultural heritage.</p>
      </abstract>
      <kwd-group>
        <kwd>Pattern recognition</kwd>
        <kwd>Visual computing</kwd>
        <kwd>Cultural heritage</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Classification is the process of labelling data items as belonging to a given class
from a model which is built from a selected set of data [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In particular,
supervised classification aims to learn a model from a training set of data, which
is then used to classify unseen test data. By contrast, the aim of the
unsupervised classification is to compute the labels from the data by grouping them into
meaningful classes. It can be performed using di↵erent optimisation algorithms
which can make assumptions about the model of data.
      </p>
      <p>
        Classification has been adopted in di↵erent domains, which include image
processing and document analysis [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In image processing, supervised
classification is used for identification of images, regions or pixels as belonging to a
given semantic class. By contrast, unsupervised classification is used for
grouping a set of images into meaningful semantic classes, or a set of pixels into
homogeneous image regions according to brightness, colour, or texture. In
document analysis, supervised classification can be employed for identification of
documents based on authorship, typology, language, script, orthography style,
dialect or sub-dialect classes. By contrast, unsupervised classification can be
useful for grouping documents based on similar characteristics, e.g. language, script
or similar content.
      </p>
      <p>
        In recent time, classification in both supervised and unsupervised forms has
played an important role for knowledge discovery in the cultural heritage. In
particular, di↵erent classification methods for images, documents and other kinds
of data have been proposed for supporting the cultural heritage understanding
and preservation. Cultural heritage is an essential part of the everyday life. It
includes old and contemporary pieces of art, buildings, furnitures, monuments,
documents, archeological sites, as well as oral traditions and habits of di↵erent
populations worldwide [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This provides an important motivation for exploring
and analysing the di↵erent state-of-the-art classification methods in the cultural
heritage domain.
      </p>
      <p>In this paper, a survey of relevant works about classification methods for
the cultural heritage is proposed. In particular, a categorisation of the cultural
heritage is first provided, which the classification methods are included in. Then,
for each category, the explored methods are described in terms of feature
representation, classification algorithm and obtained results. This paper provides
a useful guide for understanding classification methods in the cultural heritage
domain, and evaluate their limitations. It will be useful for the proposition of
innovative methods in the state-of-the-art.</p>
      <p>The paper is organised as follows. Section 2 describes the classification
methods included in the di↵erent cultural heritage categories. Finally, Section 3 draws
conclusions about the presented work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Cultural Heritage Methods</title>
      <p>
        The term cultural heritage according to UNESCO encompasses two main
categories [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]: (i) tangible, and (ii) intangible cultural heritage. Tangible cultural
heritage can be categorised into: (i) movable cultural heritage (paintings,
sculptures, coins, manuscripts), (ii) immovable cultural heritage (monuments,
archaeological sites, historical buildings), and (iii) underwater cultural heritage
(shipwrecks, underwater ruins and cities). Intangible cultural heritage has been one
of UNESCO’s priorities in the cultural domain recently, as it promotes cultural
diversity. Through preservation of oral traditions and expressions, ways of life,
traditional crafts and festivals, amongst other activities, humans are
safeguarding their cultural identities.
      </p>
      <p>The classification methods applied to the cultural heritage context can fit to
this categorisation. Figure 1 shows the cultural heritage categorisation.
2.1</p>
      <sec id="sec-2-1">
        <title>Tangible Cultural Heritage</title>
        <p>
          Movable Heritage. In the document analysis context, Esposito et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
proposed the use of an intelligent system for document processing based on machine
        </p>
        <p>Cultural Heritage
Tangible Cultural</p>
        <p>Heritage</p>
        <p>Intangible Cultural</p>
        <p>Heritage
Movable Heritage</p>
        <p>Immovable Heritage
•  Paintings
•  Sculptures
•  Furnitures
•  Wall paintings
•  Documents
•  Historical buildings
•  Monuments
•  Archeological sites
•  Oral traditions and</p>
        <p>expressions
•  Science &amp; habits related</p>
        <p>
          to nature and world
•  Traditional skills
learning techniques and adapted it to the problem of automatically processing
documents in film archives for the COLLATE project. The WISDOM++
system previously developed by users [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] provides a document transformation into
a web-accessible form such as XML format. It is designed with high
adaptivity, real-time user interaction and multi-page document management as main
features. Throughout various steps of document processing, a rule base is built
from a training set using machine learning techniques, hence a highly adaptive
system. The authors use inductive learning techniques throughout the whole
course of the document processing, consisting of document analysis,
classification and understanding, and transformation into a web-accessible format. Each
of these steps uses machine learning as a viable solution to various problems
arising in image processing such as: classifying document components with
respect to the content (separation of text from graphics), finding logical structure
of the document based on layout structure extraction, text extraction from the
relevant part of the document using Optical Character Recognition (OCR), and
the transformation of the page into HTML/XML format. Preliminary results
produced by the research are with limitations relating to the WISDOM++
system upgrade for facilitating color images, as well as images of mixed content,
and further experiments are needed. Also, need for a better integration of OCR
with WISDOM++ is reported to improve the usability by the end user.
        </p>
        <p>Also, INTHELEX (INcremental THEory Learner from EXamples) is a
learning system based on object identity assumption that learns theories from positive
and negative examples and can learn multiple concepts simultaneously. It is a
fully incremental learning system exploiting the induction of hierarchical logic
theories examples. Also, COLLATE is a project for annotation, indexing and
retrieval of digitalized historical archive material with the need to address the
automatic processing of multi-format cultural heritage documents from film archives.</p>
        <p>
          Considering that the authors successfully applied INTHELEX in the paper
document processing domain [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], it was natural to incorporate it to COLLATE to
learn rules for the automatic identification of a wide range of COLLATE
document classes. The identified classes were further used for indexing and retrieval
as well as annotation by users. There are 119 considered COLLATE documents,
belonging to five di↵erent classes (four classes of positive and one class of negative
examples). An experiment was conducted as follows: the first-order descriptions
were generated by the WISDOM++ and are used to run the learning system.
Scanned images were used to identify layout blocks (type and relative position)
hence each document belonging to positive classes was described, on average,
with 144, 215, 269, and 260 literals respectively. The features that are contained
in the document descriptions are: height and width of the layout blocks,
horizontal and vertical position, type of the layout and relative position. Each document
is, at the same time, considered a positive example of the class it belongs to and
a negative example for other classes. Definitions for each class were learnt from
the initial theory that was empty. In order to compare their solution with the
state-of-the-art system, the authors conducted experiments using Progol and
observed accuracy and runtime of both systems confirming there are insignificant
di↵erences between the two in terms of accuracy. The runtime was significantly
shorter using the INTHELEX system contributing to the incremental approach
vs. batch approach in the case of Progol. Hence, the obtained results imply that
the proposed solution is also a feasible solution for the automatic processing of
multi-format cultural heritage documents.
        </p>
        <p>
          In the same context, Brodi´c et al. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] proposed a new document classification
method for discrimination based on the language. The first phase consisted in
extracting a feature representation from the document, where each letter was
codified into a numerical code according to its extension in the text-line area. A
total of four numerical codes were considered, corresponding to grayscale levels of
an image. In this way, the document was represented by a 1-D image from which
vectors of co-occurrence texture analysis features were extracted. In the second
phase, the feature vectors were subjected to a genetic-based classification
approach for discrimination of the corresponding documents in di↵erent languages.
The algorithm represented the documents as a weighted undirected graph, where
nodes were the documents and edges linked each node with its spatially close
and k nearest neighbour nodes in terms of L1 distance among the corresponding
vectors. The weight on each edge corresponded to a similarity value computed
from the distance between the involved nodes. Then, a genetic algorithm was
applied on the graph for detecting the connected components which corresponded
to groups of documents in the same language. The genetic algorithm started
with a population of individuals, each randomly initialised and representing a
possible division of the graph in connected components. Then, variation
operators of uniform crossover and mutation were applied on each individual and
the fitness function of weighted modularity was computed. This procedure was
iterated for a given generation number. At the end, the individual with the
best fitness function was selected as the final solution. Finally, a complete link
clustering strategy for correcting the local optima was employed on the genetic
solution. The experiment was performed on a database of 85 documents given in
French, English, Serbian and Slovenian contemporary languages. The obtained
results showed the discriminatory ability of the introduced document feature
representation as well as the good performances of the genetic-based classifier
versus other competing methods.
        </p>
        <p>
          An extension of the genetic-based classifier was proposed by Brodi´c et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
for dealing with a more complex task of discrimination of languages evolved over
time. The first modification from the baseline algorithm was the introduction of
a parameter in the similarity computation for reducing the sparsity of the
document graph which may occur when representing close languages. The second
modification consisted in managing the presence of isolated nodes of the graph
inside the genetic algorithm, which can be caused by the procedure of graph
construction. The same document image coding was adopted for discrimination
of documents given in languages evolved one into another over time. The
difference was that run-length and local binary patterns instead of co-occurrence
textural features were extracted from the 1-D grayscale image of the document
and combined to create the feature vector. The experiment was carried out on
a database of 50 documents given in Italian Vulgar (1260-1374 AC) and
modern Italian languages. Comparison results with other well-known discrimination
algorithms revealed that the proposed approach overcomes the other methods
as well as the baseline genetic algorithm in discriminating the di↵erent evolved
languages.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], the same feature coding was used with the Naive-Bayes method for
identification of the orthography style of documents based on the evolution of
script or language characteristics over time. In this case, the documents were
stored as images, and a preprocessing phase of bounding box detection and
filling was carried out for finding the position of the letters in the text-line area.
This allowed the coding of the document in a sequence of numerical codes from
which the grayscale 1-D image was produced. Then, run-length and adjacent
local binary patterns were used on the obtained image for the extraction of the
document feature vector. Finally, the Naive-Bayes classifier was employed on
the vector for identification of the orthography of the corresponding document.
The proposed approach was tested on two di↵erent contexts of: (i)
languagebased orthography identification from Serbian historical documents
(classification of Slavonic-Serbian and Serbian languages), and (ii) script-based
orthography identification from Croatian historical documents (classification of old and
new Glagolitic scripts). Results obtained on two databases of digitised
documents showed that the proposed approach is robust to data noise and overcomes
other competing methods in orthography style identification.
        </p>
        <p>
          Also, Naive-Bayes and support vector machine were adopted in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] on the
same feature coding for identification of documents as given in di↵erent Serbian
pronunciations from the Shtokavian dialect. After the document image coding,
di↵erent textural features were extracted from the 1-D image and combined to
create the feature vector for the document. They included: (i) local binary
patterns, (ii) neighbour binary patterns, and (iii) the newly introduced adjacent
neighbour binary patterns. The texture operators were considered with
di↵erent parameter values and tested. Then, both classifiers were used on the feature
vectors for identification of the Serbian pronunciation of the corresponding
documents. The experiment was performed on 50 documents written in ijekavian and
ekavian pronunciations of the Shtokavian dialect in Serbian language. The
obtained results showed that the proposed method overcomes the n-gram approach
in the classification task.
        </p>
        <p>
          In the furniture context, a new technique was proposed in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] for
automatic classification of archaeological pottery sherds image views contained in
two databases, using features based on colour and texture. The classification
task is challenging considering the pottery sherds belonging to di↵erent classes
have hardly no visible di↵erences. Ground truth images indicated by specialists
are used as representatives of each class. Local features based mainly on color
properties (sherds hue, chromaticity, saturation) and local texture variance of
the pottery sherds are extracted for each pixel from both front and back views
of the individual items. Histograms of local features are created in order to focus
on prevailing information and combined with a novel bag-of-words model based
on Reddi multithresholding. In this way, a global sherd descriptor for each sherd
image is created. In order to reduce the global descriptors’ dimensionality as well
as preserve classification accuracy, the authors resorted to the Principal
Component Analysis (PCA) feature selection statistical method as the most balanced
amongst eight considered feature selection methods. This process o↵ers a
tradeo↵ between the computational complexity (size of the feature matrix) and loss of
information on the other side resulting in potential misclassification. The authors
use several machine learning algorithms, namely K-Nearest Neighbour (KNN),
support vector machine, Naive-Bayes, SMO, and Simple Logic for classification
of the reduced global descriptor and have shown that the KNN algorithm is the
best selection for a given method. The classification rate when using ground truth
images as training example for di↵erent classes, or when splitting the dataset
into training and testing parts (40% training, 60% testing) provides comparable
results of about 70%. Additionally, the authors have tested their method on the
ceramic sherd database as well as compared their method in both databases
with four state-of-the-art feature extraction techniques (self-similarity, pyramid
histogram of words–color/gray–geometric blur and a weighted combination of
them) that have demonstrated an improved performance when tested in a generic
image database. In the case of both tests satisfactory results were obtained.
        </p>
        <p>
          Finally, Mensink and van Gemert [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] introduced a new large dataset
composed of 112,039 artistic items from the collection of Rijksmuseum in
Amsterdam, the Netherlands. The items consisted of images and associated metadata
in XML format from portraits, furniture, pictures, miniatures, etc. of ancient
and medieval period, and late 19th century. The dataset is open and freely
available for tasks of art image classification and content-based image retrieval. Also,
di↵erently from previous datasets of specific artistic objects, such as paintings
or vases, the proposed dataset is broad and variegated enough for capturing a
museum-centered view. Starting from the dataset, four open challenges were
proposed, including: (i) prediction of the artist, (ii) prediction of the type of artistic
work, (iii) prediction of the material by which the artistic work is created, and
(iv) prediction of the year of creation. For each challenge, Fisher vectors were
adopted as image feature representation and support vector machine was used
for classification of the artistic images. All material retrieved from the
experiments is available for download, which includes dataset, experimental settings
and image features.
        </p>
        <p>
          Immovable Heritage. To protect, keep and eventually restore tangible
cultural heritage buildings, systematic image collection and classification in order
to correctly interpret and manage the vast information is of utmost importance.
Guiding principles for recording, documentation and information management
for the conservation of heritage buildings were used in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], such that, in their
works the authors created datasets comprising of ten di↵erent categories and
performed classification of the images. A motivation for creating datasets of the
architectural cultural heritage images was the lack of available datasets within
the research community. The dataset encompasses ten di↵erent categories of
Cultural Heritage images, namely Altar, Apse, Bello tower, Column, Dome
(inner), Dome (outer), Flying buttress, Gargoyle (and Chimera), Stained glass, and
Vault, totalling to over 10,000 images. Each image has a main element that is
reflected in the labelling process. One of the authors’ contributions is that the
labeled images can be used as a starting point to establish superior
classification techniques. They claim to have used deep learning techniques based on
Convolutional Neural Network (CNN) that previously have not been used for
classification of cultural heritage images and reported significant accuracy.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], the authors’ main motivation was the classification of three di↵erent
architectural styles of buildings from images of the Mexican cultural heritage.
Raw video content was used for the production of cultural heritage images, hence
they contain undesirable yet not obvious elements besides the object in the focus
that could impair the classification process. Machine learning techniques can
address issues of identification of complex patterns in large sets of data where
the objects in the focus are positioned within a specific scene (amongst other
objects in the image), perhaps given in a di↵erent perspective, etc. The authors
use visual attention predictors (saliency driven content selection approaches) and
a traditional cropped image (centered-content) for training of a CNN in order
to classify di↵erent styles of architecture. Graph based visual saliency is the
model used for saliency prediction. The results show that using visual attention
predictors increases the quality of the cropped data, hence better classification
results are obtained based on saliency driven content selection in comparison to
a traditional cropped image selection. The obtained classification rates and the
authors’ future work course imply that the system for batch processing of raw
video footage could be devised with the contribution of creating a large number
of cultural heritage images.
        </p>
        <p>
          Also, Grilli et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] aim to improve on traditional preservation and
restoration methods of cultural heritage monuments. Such complex tasks require, at
least, documenting and archiving architectural/historical heritage content,
differentiating between various techniques of constructing buildings, and
recognition of previous restoration evidence, if any. Point clouds, amongst other
purposes, are used for 3D modelling of the cultural heritage monuments. Present
challenges in automatic model analysis are in the domain of segmentation and
classification of point clouds. Authors use 2D segmentation of the texture of 3D
models generated based on three archeological case studies conducted in Italy
(Villa Adriana in Tivoli; Cavea walls of the Circus Maximus in Rome; Portico
in Bologna). Downsides of using this technique come from the large amount of
data drawn from di↵erent types of constructing buildings as well as decorative
elements used on the facades. In addition, the cultural heritage sites are from
di↵erent time periods and in di↵erent stages of deterioration, hence the
classification is less ecient. Supervised machine learning (based on decision tree
algorithms) on UV maps, generated from the unwrapped textured 3D models,
is used for classification of the 3D cultural heritage models. Accuracy and
algorithm execution time were used as performance metrics. A training stage is
required for supervised learning and a sucient dataset of manually labeled
images has been used. The authors propose using deep learning to get a better
classification accuracy and hence propose to adequately extend the training set.
        </p>
        <p>
          In the context of archeological sites, ICARE is an emerging digital heritage
platform presently at the early stages of development [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. It is envisioned as
multimodal archiving system intended for archiving digital heritage content from
potential donors and performing semantic queries for searching data archives
from potential requestors. Both donors and requestors are interacting through
the system that could in the future, by means of multimedia, metadata and text
archives, provide believable digital cultural experiences in cultural heritage sites
that are forever lost to humankind. So far the authors have collected images from
cultural heritage sites in the Palmyra region and performed tests of several
machine learning models for visual categorization. The bag of features framework is
a novel technique for visual categorization by the support vector machine
classifier that used the largest number of categories in a simultaneous experiment.
Deep learning using CNN is also used by the authors for visual categorization.
An episodic memory technique was adopted to achieve better classification
results considering the limited image dataset. CNN using the transfer learning
approach provided the best classification accuracy.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Intangible Cultural Heritage</title>
        <p>
          As oral traditions, Michon et al. [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] solved the problem of identification of
Arabic dialects using neural network architectures which allowed to achive
satisfactory results with also a reduced number of training data for learning. In
particular, the task required the discrimination of Modern Standard Arabic
language and four Arabian dialects: (i) Egyptian, (ii) Gulf, (iii) Levantine, and
(iv) North African, which are pretty similar. Three di↵erent runs were
generated for the challenge. The first one was a neural network model characterised
by a Multi-Input CNN, which separately learned lexical, acoustic and phonetic
characteristics of the languages. The second run was characterised by a neural
network CNN-biLSTM where the inputs were speech spectrograms, from which
spatial and sequential characteristics were extracted. The last run included a
binary CNN-biLSTMs. Results obtained by the di↵erent runs were compared with
results obtained by a support vector machine. It was visible that Multi-Input
CNN (first run) obtained the best performances versus the other methods and
architectures.
        </p>
        <p>
          Also, Ma et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] explored the properties of their structural pronunciation
representation for the extraction of linguistic features useful to identify the
dialect or subdialect of Chinese speakers. In the first stage, the experiment involved
a dialect-based speaker classification on data from 19 speakers with di↵erent
dialects and sub-dialects. In the second phase, a new dataset of recordings in the
di↵erent dialects and sub-dialects from an expert dialectologist was created with
minimum speaker di↵erences. Another classification task was performed on the
new dataset which obtained close results than the previous experiment. Finally,
a last dataset of recordings with maximum speaker di↵erences was built.
Different classification experiments were carried out on the original and simulated
datasets, based on both structural and spectral comparison. Results showed that
the structural comparison is able to capture the linguistic features and
classification performances do not depend on the speaker characteristics.
        </p>
        <p>
          In the traditional skills and habits context, Liu et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] proposed a
classification approach for identification of music pieces as belonging to cultural
styles. They were represented by four di↵erent types of features: (i) timbre, (ii)
wavelet coecients, (iii) rhythm, and (iv) characteristics based on musicology.
In particular, the timbral texture was adopted, which was composed of spectral
shape and contrast. Also, the rhytm was characterised by strength, regularity
and tempo. As musicology-based features, the chromogram, chord distribution,
contrast, and pitch interval histogram were adopted. Finally, the first three
moments (mean, variance and skewness) of the normalised histogram of wavelet
coecients in each level were extracted. The experiment was carried out on a
dataset of 1300 music pieces, each represented by the aforementioned feature
sets. Three classifiers were then employed on the dataset: (i) decision tree, (ii)
KNN and (iii) multi-class support vector machine. The obtained results showed
that support vector machine and KNN overcome the decision tree in
classification of the music pieces in six cultural styles: (i) Western classical music, (ii)
Chinese traditional music, (iii) Japanese traditional music, (iv) Indian classical
music, (v) Arabic folk music, and (vi) African folk music. In particular,
support vector machine obtained the highest accuracy of more than 86% when all
features were used.
        </p>
        <p>
          Also, Dimitropoulos et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] presented the i-Treasures Project for
management, extraction and analysis of intangible cultural heritage, which included
traditional songs, dance, pottery and contemporary music pieces. In particular,
the project aimed to build a platform for open access to intangible cultural
heritage, transmission and exchange of knowledge. This was not only focused on
digitization of cultural resources, but also on generation of new knowledge by
exploiting new multisensory-based techniques for exploring the digital content.
The first step in the project was the semantic analysis of the cultural content in
order to create a unified knowledge platform. It was accomplished by extracting
semantics for capturing hidden patterns and relationships among elements of
intangible cultural heritage, which are useful for tracking its evolution over the
di↵erent generations. Image and signal processing methods were employed on
the di↵erent signals for the feature extraction. Finally, a learning environment
was developed using 3D technology of web-based game engines.
        </p>
        <p>
          Finally, Li et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] proposed a method for classification of traditional
Chinese folk songs according to the di↵erent regional styles. A temporal model was
adopted for capturing the temporal structures of the folk songs. In particular,
Conditional Random Fields (CRF) were used for creating the model, whose state
and transfer functions were computed by the Gaussian Mixture Model (GMM)
approach. The experiment involved a dataset of 334 folk songs from the Chinese
regions of Shaanxi, Jiangsu and Hunan. For each song category, the CRF model
was learned from the training set. Then, the GMM was employed for fitting the
audio frame features and estimating the corresponding label sequence for each
CRF. In order to predict the category of a folk song in the test set, for each CRF
corresponding to a specific category (regional style), the posterior probability of
the folk song to belong to the given category was computed. The regional style
class of the folk song was that corresponding to the highest computed posterior
probability. The obtained results showed that the proposed classification method
achieved an improvement of 4.6%-18.13% versus the other competing methods,
i.e. support vector machine, KNN, etc.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>This paper presented a categorisation of classification methods according to
different types of cultural heritage. Hence, tangible and intangible heritage
classification methods were described. In the tangible heritage, movable and immovable
heritage classification methods were analysed. Each method was characterised
by feature representation, classification algorithm and obtained results. Apart
from the di↵erent features, the classification algorithms for the tangible
movable heritage included: (i) rule-based algorithms, (ii) genetic algorithms, (iii)
Naive-Bayes, (iv) support vector machine, and (v) KNN. By contrast, CNN
and decision trees were predominantly used for tangible immovable heritage.
Finally, CNN, decision trees, KNN, CFR-GMM and support vector machine were
adopted for intangible heritage. Table 1 shows an overview of the described
categorisation of classification algorithms. This work can be useful for analysis of the
current literature in the field and for the proposition of new methods overcoming
the limitations of the state-of-the-art approaches.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amelio</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amelio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Classification Methods in Image Analysis with a Special Focus on Medical Analytics</article-title>
          , pp.
          <fpage>31</fpage>
          -
          <lpage>69</lpage>
          . Springer International Publishing (
          <year>2019</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -94030-4 3
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>T.M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferilli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di Mauro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esposito</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Incremental induction of classification rules for cultural heritage documents</article-title>
          .
          <source>In: Proceedings of the 17th International Conference on Innovations in Applied Artificial Intelligence</source>
          . pp.
          <fpage>915</fpage>
          -
          <lpage>923</lpage>
          . IEA/AIE'2004, Springer Springer Verlag Inc (
          <year>2004</year>
          ). https://doi.org/10.1007/b97304
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Brodi´c, D.,
          <string-name>
            <surname>Amelio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Discrimination of di↵erent serbian pronunciations from shtokavian dialect</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>112</volume>
          ,
          <fpage>1935</fpage>
          -
          <lpage>1944</lpage>
          (
          <year>2017</year>
          ). https://doi.org/https://doi.org/10.1016/j.procs.
          <year>2017</year>
          .
          <volume>08</volume>
          .047, knowledgeBased and
          <string-name>
            <given-names>Intelligent</given-names>
            <surname>Information</surname>
          </string-name>
          &amp;
          <source>Engineering Systems: Proceedings of the 21st International Conference, KES-20176-8 September</source>
          <year>2017</year>
          , Marseille, France
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Brodi´c, D.,
          <string-name>
            <surname>Amelio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Recognizing the orthography changes for identifying the temporal origin on the example of the balkan historical documents</article-title>
          .
          <source>Neural Computing and Applications (Nov</source>
          <year>2017</year>
          ). https://doi.org/10.1007/s00521-017-3292-1
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Brodi´c, D.,
          <string-name>
            <surname>Amelio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Milivojevi´c,
          <string-name>
            <surname>Z.N.</surname>
          </string-name>
          :
          <article-title>Clustering documents in evolving languages by image texture analysis</article-title>
          .
          <source>Applied Intelligence</source>
          <volume>46</volume>
          (
          <issue>4</issue>
          ),
          <fpage>916</fpage>
          -
          <lpage>933</lpage>
          (
          <year>Jun 2017</year>
          ). https://doi.org/10.1007/s10489-016-0878-8
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Brodi´c, D.,
          <string-name>
            <surname>Amelio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Milivojevi´c,
          <string-name>
            <surname>Z.N.</surname>
          </string-name>
          :
          <article-title>Language discrimination by texture analysis of the image corresponding to the text</article-title>
          .
          <source>Neural Computing and Applications</source>
          <volume>29</volume>
          (
          <issue>6</issue>
          ),
          <fpage>151</fpage>
          -
          <lpage>172</lpage>
          (
          <year>Mar 2018</year>
          ). https://doi.org/10.1007/s00521-016-2527-x
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dimitropoulos</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manitsaris</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsalakanidou</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikolopoulos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Denby</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kork</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crevier-Buchman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pillot-Loiseau</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adda-Decker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dupont</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tilmanne</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alivizatou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yilmaz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hadjileontiadis</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Charisis</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deroo</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manitsaris</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kompatsiaris</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grammalidis</surname>
          </string-name>
          , N.:
          <article-title>Capturing the intangible an introduction to the i-treasures project</article-title>
          .
          <source>In: 2014 International Conference on Computer Vision Theory and Applications (VISAPP)</source>
          .
          <source>vol. 2</source>
          , pp.
          <fpage>773</fpage>
          -
          <lpage>781</lpage>
          (
          <year>Jan 2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Esposito</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malerba</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semeraro</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferilli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altamura</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>T.M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berardi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceci</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mauro</surname>
            ,
            <given-names>N.D.</given-names>
          </string-name>
          <article-title>: Machine learning methods for automatically processing historical documents: from paper acquisition to xml transformation</article-title>
          .
          <source>In: First International Workshop on Document Image Analysis for Libraries</source>
          ,
          <year>2004</year>
          . Proceedings. pp.
          <fpage>328</fpage>
          -
          <lpage>335</lpage>
          (
          <year>Jan 2004</year>
          ). https://doi.org/10.1109/DIAL.
          <year>2004</year>
          .1263262
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>European-Union</surname>
          </string-name>
          :
          <article-title>The european year of cultural heritage 2018</article-title>
          , https://europa.eu/cultural-heritage/about
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Grilli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dininno</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petrucci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Remondino</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>From 2d to 3d supervised segmentation and classification for cultural heritage applications</article-title>
          . ISPRS - International
          <source>Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences XLII-2</source>
          ,
          <fpage>399</fpage>
          -
          <lpage>406</lpage>
          (
          <year>2018</year>
          ). https://doi.org/10.5194/isprs-archives
          <source>-XLII-2- 399-2018</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kurniawan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salim</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suhartanto</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasibuan</surname>
            ,
            <given-names>Z.A.</given-names>
          </string-name>
          :
          <article-title>E-cultural heritage and natural history framework: an integrated approach to digital preservation</article-title>
          .
          <source>In: International Conference on Telecommunication Technology and Applications</source>
          . pp.
          <fpage>177</fpage>
          -
          <lpage>182</lpage>
          .
          <source>Proc .of CSIT</source>
          vol.
          <volume>5</volume>
          , IACSIT Press (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>The regional style classification of chinese folk songs based on gmm-crf model</article-title>
          .
          <source>In: Proceedings of the 9th International Conference on Computer and Automation Engineering</source>
          . pp.
          <fpage>66</fpage>
          -
          <lpage>72</lpage>
          . ICCAE '17,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2017</year>
          ). https://doi.org/10.1145/3057039.3057069
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Cultural style based music classification of audio signals</article-title>
          .
          <source>In: 2009 IEEE International Conference on Acoustics, Speech and Signal Processing</source>
          . pp.
          <fpage>57</fpage>
          -
          <lpage>60</lpage>
          (
          <year>April 2009</year>
          ). https://doi.org/10.1109/ICASSP.
          <year>2009</year>
          .4959519
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Llamas</surname>
            , J.,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lerones</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Medina</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zalama</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gmez-Garca-Bermejo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Classification of architectural heritage images using deep learning techniques</article-title>
          .
          <source>Applied Sciences</source>
          <volume>7</volume>
          (
          <issue>10</issue>
          ) (
          <year>2017</year>
          ). https://doi.org/10.3390/app7100992
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Minematsu</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qiao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirose</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Dialect-based speaker classification using speaker-invariant dialect features</article-title>
          .
          <source>In: 2010 7th International Symposium on Chinese Spoken Language Processing</source>
          . pp.
          <fpage>171</fpage>
          -
          <lpage>176</lpage>
          (
          <year>Nov 2010</year>
          ). https://doi.org/10.1109/ISCSLP.
          <year>2010</year>
          .5684491
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Makridis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daras</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Automatic classification of archaeological pottery sherds</article-title>
          .
          <source>J. Comput. Cult. Herit</source>
          .
          <volume>5</volume>
          (
          <issue>4</issue>
          ),
          <volume>15</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          :
          <fpage>21</fpage>
          (Jan
          <year>2013</year>
          ). https://doi.org/10.1145/2399180.2399183
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Mensink</surname>
            , T., van Gemert,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The rijksmuseum challenge: Museum-centered visual recognition</article-title>
          .
          <source>In: Proceedings of International Conference on Multimedia Retrieval</source>
          . pp.
          <volume>451</volume>
          :
          <fpage>451</fpage>
          -
          <lpage>451</lpage>
          :
          <fpage>454</fpage>
          . ICMR '14,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2014</year>
          ). https://doi.org/10.1145/2578726.2578791
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Michon</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pham</surname>
            ,
            <given-names>M.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crego</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Senellart</surname>
          </string-name>
          , J.:
          <article-title>Neural network architectures for arabic dialect identification</article-title>
          .
          <source>In: Proceedings of the Fifth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial</source>
          <year>2018</year>
          ). pp.
          <fpage>128</fpage>
          -
          <lpage>136</lpage>
          . Association for Computational Linguistics (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Obeso</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          , Va´zquez,
          <string-name>
            <given-names>M.S.G.</given-names>
            ,
            <surname>Acosta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.A.R.</given-names>
            ,
            <surname>Benois-Pineau</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          :
          <article-title>Connoisseur: Classification of styles of mexican architectural heritage with deep learning and visual attention prediction</article-title>
          .
          <source>In: Proceedings of the 15th International Workshop on Content-Based Multimedia Indexing</source>
          . pp.
          <volume>16</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          :
          <fpage>7</fpage>
          . CBMI '17,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2017</year>
          ). https://doi.org/10.1145/3095713.3095730
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Yasser</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clawson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowerman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Saving cultural heritage with digital make-believe: Machine learning and digital techniques to the rescue</article-title>
          .
          <source>In: Proceedings of the 31st British Computer Society Human Computer Interaction Conference</source>
          . pp.
          <volume>97</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>97</lpage>
          :
          <fpage>5</fpage>
          . HCI '17,
          <string-name>
            <given-names>BCS</given-names>
            <surname>Learning</surname>
          </string-name>
          &amp; Development Ltd.,
          <string-name>
            <surname>Swindon</surname>
          </string-name>
          , UK (
          <year>2017</year>
          ). https://doi.org/10.14236/ewic/HCI2017.97
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>