<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ontology-based descriptions of image collections</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ma. Auxilio Medina</string-name>
          <email>mmedina@uppuebla.edu.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Alfredo Sa´nchez</string-name>
          <email>j.alfredo.sanchez@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. de la Calleja</string-name>
          <email>jdelacalleja@uppuebla.edu.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Benitez</string-name>
          <email>abenitez@uppuebla.edu.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad Polit ́ecnica de Puebla Tercer Carril del Ejido Serrano S/N Juan C. Bonilla</institution>
          ,
          <addr-line>Puebla, M ́exico</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad de las Am ́ericas Puebla Ex-Hacienda Santa Catarina M ́artir S/N San Andr ́es Cholula</institution>
          ,
          <addr-line>Puebla, M ́exico</addr-line>
        </aff>
      </contrib-group>
      <fpage>133</fpage>
      <lpage>140</lpage>
      <abstract>
        <p>Thesaurus have been widely used to provide structured vocabularies to describe books, however, they only provide part of the required knowledge for semantic web contexts. This paper proposes the use of ontologies and standard metadata to model semantic descriptions of image collections that refers to book covers. Ontologies represent keyword-based organizations. The classes and the hierarchical relationships of the ontology allow users to query images by topic with common search engines. The greenBookC collection is used as a test bed. This collection is published and maintained in Greenstone, a suite of software for digital libraries.</p>
      </abstract>
      <kwd-group>
        <kwd>digital libraries</kwd>
        <kwd>ontologies</kwd>
        <kwd>semantic descriptions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Users of digital libraries employ keyword based search engines, browsing
mechanisms or recommendation systems to find relevant information. In these tools,
words are considered as “units of meaning”. However, there is a gap between the
contents and meanings in multimedia data [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>Different approaches exist to organize multimedia collections that vary from
processing low level features to associate concepts to media. The results of these
approaches often require knowledge of experts or specialized users. Most of the
users of digital libraries interact with image collections through search engines
that allow them to query images by title, date, author, format, publisher or with
a feature associated with the file that stores the image. These descriptors are
associated with image content.</p>
      <p>
        Frequently, image collections are also described using a set of terms of a
keyword-based organization system or with natural language descriptions. In
order to extend the implicit syntactic level of these descriptions, semantic
descriptions with predefined structures and control vocabularies are considered richer
structures to store background knowledge [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Our interest is focused on describe images that refers to book covers.
Traditionally, thesaurus have been widely used to provide structured vocabularies to
describe books, however, they only provide part of the required knowledge for
semantic web contexts. In semantic digital libraries (SDLs), materials, tools and
meanings are addressed to offer benefits such as: anyone can use it, knowledge
is accessible from the SDL, resources are available with the modality anytime
anywhere, there are friendly and multi-modal interfaces with multiple connected
devises [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>This paper explores an alternative to associate background knowledge to
image descriptions that refers to book covers. The alternative makes use of
ontologies and metadata standards. The classes and the hierarchical relationships
of ontologies can be exploited by keyword-based search engines. A collection of
images is built as a test bed. The collection is published in Greenstone, a suite
of software for digital libraries.</p>
      <p>The paper is organized as follows. Section 2 presents brief descriptions of
software tools that allow users to construct image collections in the semantic
web. Section 3 explains the ontologies and metadata management. Section 4
describes the test bed collection. Section 5 presents preliminary results. Finally,
Section 6 includes conclusions and suggests future directions of our work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>Open freely distributed software exists around the world to construct image
collections for educational, government and commercial institutions. This section
describes some of the most common tools used in semantic web applications and
depicts some representative collections.</p>
      <p>
        DSpace1 allow users to create repositories of digital content in multiple
formats. This is a widely popular tool that supports the customization of interfaces
to fit different user needs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. “The Spanish Image collection” is an example of a
collection constructed in DSpace. The collection is formed by 1196 JPEG files2.
The images can be query with descriptors such as author, title or date; a list
of categories is used to organize images. The description of images at DSpace is
syntax-based.
      </p>
      <p>
        Fedora commons repository software3 is a general purpose, open-source
digital object repository system under the Apache License. This is a centered
platform that enables storage, access and management of digital content, although
does not support indexing, discovery and delivery application mechanisms [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
The modules offer interoperability and extensibility of data. Image collections
can be constructed upon Fedora such as the “Maryland Map Collection”. This
collection is formed by JPEG files that depicts Maryland, the Chesapeake Bay,
1 Dspace home page is available at http://www.dspace.org/
2 The Spanish Image Collection is available at:
      </p>
      <p>http://dspace.nitle.org/handle/10090/1267
3 Fedora home page is available at: http://fedora-commons.org/. This work is licensed
under a Creative Commons Attribution-Share Alike 3.0 Unported License
and the surrounding region from 1590 to the present. The interface visualizes
images and metadata. The Libraries’ Catalog and the ArchivesUM are used to
organize this collection.</p>
      <p>
        Greenstone4 is a suite of software to build and construct collection of
semantic digital libraries. This is produced by the New Zealand Digital Library Project
at the University of Waikato, and developed and distributed in cooperation with
UNESCO and the Human Info NGO. This is an open-source, multilingual
software issued under the terms of the GNU General Public License. Greenstone
supports a variety of formats that represents text, images, audio and video [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
The “Historic Campus Architecture Project”5 developed by the Council of
Independent Colleges (CIC) in the United States, is an example of an image collection
in Greenstone. In the project, 5000 images of hard copy photographs of 4 x 6
and 8 X 10 inches were transformed into JPEG files. Metadata is used as the
descriptors to explore the collection. Metadata is stored in XML files.
      </p>
      <p>
        After analyzing these software tools, we choose Greenstone to implement
the ontology-based alternative to describe image collections. The features of
Greenstone related with the support of semantic descriptions of images are the
following ones [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]:
– The digital library server runs in different platforms
– Full text is searchable by default
– Servers and harvesters data compliance with the Open Archives Initiative
      </p>
      <p>Protocol for Metadata Harvesting (OAI-PMH)
– Export and ingest data from DSpace
– Collections can be updated anytime without disturbing users
– Dublin Core (DC) is the default metadata format when a new collection is
constructed
3</p>
    </sec>
    <sec id="sec-3">
      <title>Building semantic descriptions</title>
      <p>
        Although thesaurus have been a typical classification system of digital libraries,
they are not able to accomplish the information requirements of the semantic
web. Ontologies have larger representation power than thesaurus [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. An
adaptation of the steps proposed [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] was done to construct a lightweight ontology used
to build semantic descriptions.
      </p>
      <p>The steps to construct semantic descriptions based on metadata and
ontologies that refers to book covers are the following ones:
1. Construction of a description template for book covers. This template
answer questions such as: what kind of information does users want to record
for a particular image, how users query an image collection of book covers,
what format is appropriate to store images, how can images be organized,
4 Greenstone is available at: http://www.greenstone.org/
5 The History Campus Arquitecture Project is available at
http://puka.cs.waikato.ac.nz/cgi-bin/cic/library
what metadata standard is useful to support semantic descriptions, will the
descriptions be used locally or a mechanism of information sharing must be
considered
2. Linking images to ontology classes in such a way that classes can be used as
values for a metadata standard. Hierarchical relationships should be used as
a query expansion mechanism
3. Describing additional domain and background knowledge in metadata
standard elements</p>
      <p>Each case study might use of a different ontology as well as a distinct
metadata standard as input. The goal in this work is to exploit ontological
characteristics as the basis to improve image descriptions and retrieval mechanisms.</p>
      <p>Table 1 shows an excerpt of an ontology that was used to construct semantic
descriptions of book covers that form a collection called greenBookC. This
collection correspond to book covers of physical books of the library of the UPPuebla.
The areas of knowledge proposed by the “Consejo Nacional de Ciencia y
Tecnologa” (CONACYT) are used as the main classes of this ontology.</p>
      <p>The ontology of Table 1 has 8 main classes and 29 subclasses organized in
3 levels. Each class has a label, a level and one or more instances (images). An
additional class is used to hold images of specific types of literature: biographies,
dictionaries and novels. The classes were constructed using knowledge of experts.
Semantic information is represented by metadata attached to each image; in
particular, the dc:subject element is used to support the organization of topics,
this means that the names of the classes are used to fill this element. Table 1
shows the DC elements used to form a semantic description. An institutional
policy help to maintain a controlled vocabulary for the semantic descriptions.</p>
      <p>The hierarchical relationships of the ontology are useful for free text search.
For example, a cover book that belongs to the materials class, it is also considered
as a member of the industrial engineering class.
4
greenBookC: an image collection with semantic
descriptions at Greenstone
This section describes the use of the ontology described in Section 3 to enrich
an image collection of book covers. GreenBookC collection uses existing legal
metadata of Greenstone. Semantic description of images are described with the
unqualified Dublin Core (DC) elements of Table 2 as shown in Figure 1. After a
metadata standard is added, metadata elements need to be filled as illustrated in
Figure 2. These data can be assigned in different languages in order to improve
collection accessibility.</p>
      <p>
        According to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], semantic descriptions are stored in standard metadata
formats and knowledge representation languages such as XML, RDF, RDF-Schema
and OWL. At greenBookC, these descriptions are stored as XML files. These files
can be exported to DSpace or be processed by another XML tools.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Preliminar results</title>
      <p>
        The greenBookC collection is formed by 1504 images of book covers in JPEG
format. JPEG is a standard to represent compressed continuous-tone images [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
The images are normalized as follows: 32 bits for colours, each one 288 per 352
pixels. The size of the files varies between 33 and 98 Kbytes. A document camera
was used to get the images in order to enhance text and graphics.
The create panel at the librarian interface of Greenstone is used to construct
the greenBookC collection. The simple image collection (image-e) from the demo
Greenstone collections was used as a template. The ImagePlugin of Greenstone
processed all the images appropriately on a Windows XP operative system.
      </p>
      <p>greenBookC has a web-based interface to search the images. Descriptors such
as topic, author or the browse list can be used to query the images. Users can
carry traditional keyword-based searches in different DC elements or use the
metadata descriptions whose values comes from the ontology. Distinct languages
can be used to query the collection. At the date, the work has been address to
construct the semantic descriptions and empirically the descriptions have
improve collection accessibility. However, a lot of effort is still required to evaluate
the proposed approached.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>The use of lightweight ontologies to enrich metadata standard is a basic but
simple policy to organize images by content and an alternative to construct
semantic descriptions. Ontologies and metadata enable information sharing
between academic communities. Ontologies provide controlled vocabularies that
help to reduce ambiguity from natural language or keyword based descriptions.</p>
      <p>The visualization of book covers at Greenstone allow users to recognize visual
features of book covers. The use of existing legal metadata enables information
sharing and improves interoperability in digital libraries. Different metadata
standards or additional elements can be used to store relevant information of
book covers as the abstract or the synopsis of a book, the publisher and the
physical location, respectively.</p>
      <p>There are several challenges in the construction of collections for semantic
digital libraries, however there are software tools that can help users to develop
this task successfully. As future work, we plan to represent the classes,
subclasses, individual, properties and restrictions of the ontology formally in order
to integrate reasoning capabilities.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We thank to the staff library at the UPPuebla for their help and cooperation in
the construction process. This work is partially supported by PROMEP grant
Biblioteca Digital Sem´antica de Recursos Educativos /103-5/09/4023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Allemang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hendler</surname>
          </string-name>
          , J.:
          <article-title>Semantic web for the working ontologist</article-title>
          . Morgan Kauffman Publishers (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kruk</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McDaniel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          : Semantic Digital Libraries. Springer-Verlag (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lagoze</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Payette</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilper</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Fedora: an architecture for complex objects and their relationships</article-title>
          .
          <source>International Journal on Digital Libraries V6(2)</source>
          ,
          <fpage>124</fpage>
          -
          <lpage>138</lpage>
          (
          <year>Apr 2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Ricardo</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.R.N.</surname>
          </string-name>
          :
          <article-title>Modern information retrieval</article-title>
          .
          <source>Addison Wesley</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Wielinga</surname>
            ,
            <given-names>B.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreiber</surname>
            ,
            <given-names>A.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wielemaker</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sandberg</surname>
            ,
            <given-names>J.A.C.</given-names>
          </string-name>
          :
          <article-title>From thesaurus to ontology</article-title>
          .
          <source>In: Proceedings of the 1st international conference on Knowledge capture</source>
          . pp.
          <fpage>194</fpage>
          -
          <lpage>201</lpage>
          . K-CAP '
          <fpage>01</fpage>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2001</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/500737.500767
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bainbridge</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nichols</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          :
          <article-title>How to Build a Digital Library</article-title>
          . Morgan Kaufmann, Burlington, MA, 2. edn. (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>