<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A linked open data architecture for contemporary historical archives</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexandre Rademaker</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Suemi Higuchi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dario Augusto Borges Oliveira</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IBM Research</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>FGV/EMAp</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>FGV/CPDOC</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>52</fpage>
      <lpage>64</lpage>
      <abstract>
        <p>This paper presents an architecture for historical archives maintenance based on Open Linked Data technologies and open source distributed development model and tools. The proposed architecture is being implemented for the archives of the Center for Teaching and Research in the Social Sciences and Contemporary History of Brazil (CPDOC) from Getulio Vargas Foundation (FGV).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>amount of textual, iconographic and audiovisual documents; (2) updating the
dictionary DHBB; and (3) prospecting innovative technologies that enable new
uses for CPDOCs collections.</p>
      <p>The advances in technology o er new modes of dealing with digital contents
and CPDOC is working to make all data available in a more
intelligent/semantic way in the near future, o ering swift access to its archives. In collaboration
with the FGV School of Applied Mathematics (EMAp), CPDOC is working on a
project that aims to enhance access to documents and historical records by means
of data-mining tools, semantic technologies and signal processing. At the
moment, two applications are being explored: (1) face detection and identi cation
in photographs, and (2) voice recognition in the sound and audiovisual archives
of oral history interviews. Soon it will be easier to identify people in the
historical images, and link them to the entries in CPDOC archives. Additionally, voice
recognition will help locate speci c words and phrases in audiovisual sources
based on their alignment with transcription { a tool that is well-developed for
English recordings but not for Portuguese. Both processes are based on machine
learning and natural language processing, since the computer must be taught to
recognize and identify faces and words.</p>
      <p>CPDOC also wants its data to constitute a large knowledge base, accessible
using the standards of semantic computing. Despite having become a reference
in the eld of organization of collections, CPDOC currently do not adopt any
metadata standards nor any open data model for them. Trends for data sharing
and interoperability of digital collections pose a challenge to the institution to
remain innovative in its mission of e ciently providing historical data. It is time
to adjust CPDOC's methodology to new paradigms.</p>
      <p>
        In Brazilian scenario many public data is available for free, but very few
are in open format following the semantic web accepted standards. Examples in
this direction are the Governo Aberto SP [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the LeXML [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] and the SNIIC
project 3.
      </p>
      <p>
        In this sense, we present hereby a research project that re ects a change in the
way CPDOC deals with archives maintenance and di usion. The project is an
ongoing initiative to build a model of data organization and storage that ensures
easy access, interoperability and reuse by service providers. The project proposal
is inspired by: (1) Open Linked Data Initiative principles [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]; (2) distributed
open source development model and tools for easy and collaborative data
maintenance; (3) a growing importance of data curating concepts and practices for
online digital archives management and long-term preservation.
      </p>
      <p>
        The project started with an initiative of creating a linked open data version
of CPDOC's archives and a prototype with a simple and intuitive web interface
for browsing and searching the archives was developed. The uses of Linked Open
Data concept are conformed to the three laws rst published by David Eaves [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
and now widely accepted: (1) If it can't be spidered or indexed, it doesn't exist;
3 Sistema Nacional de Informaes e Indicadores Sociais, http://culturadigital.br/
sniic/.
(2) If it isn't available in open and machine readable format, it can't engage; and
(3) If a legal framework doesn't allow it to be repurposed, it doesn't empower.
      </p>
      <p>
        Among the project objectives we emphasize the construction of a RDF [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]
data from data originally stored in a relational database and the construction of
an OWL [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] ontology to properly represent the CPDOC domain. The project
also aims to make the whole RDF data available for download similarly to what
DBpedia does [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>This paper re ects a real e ort grounded in research experience to keep
CPDOC as a reference institution in the eld of historic preservation and
documentation in Brazil.
2</p>
    </sec>
    <sec id="sec-2">
      <title>CPDOC information systems</title>
      <p>2.1</p>
      <sec id="sec-2-1">
        <title>Personal Archives (Acessus)</title>
        <p>This system is composed by personal les from people who in uenced the
political and social scenario of our country. These historical documents, in textual
or audiovisual form, in form of handwritten and printed texts, diaries, letters,
photographs, speeches or memos, represent much more than private memories:
they are the registry of a whole collective memory.</p>
        <p>Currently, more than 200 personal archives from presidents, ministers,
military personal and other Brazil's important public gures compose the CPDOC's
collections. Together, they comprise nearly 1.8 million documents or 5.2 millions
pages. From this, nearly 700 thousands pages are in digital format and the
expectance is to digitize all collections in the next few years. The collection entries
metadata are stored in an information system called Acessus. It can be accessed
through the institution's intranet for data maintenance or by internet for simple
data query. Currently, allowed queries are essentially syntactic, i.e., restricted to
keywords searches linked to speci c database elds de ned in an ad hoc manner.
For those documents that are already digitized, two digital le versions were
generated: one in high resolution aiming long-term preservation and another in
low resolution for web delivery. High resolution les are stored in a storage
system with disk redundancy and restricted access, while low resolution les are
stored in a le server 4 (Figure 1).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Oral History Interviews (PHO)</title>
        <p>The CPDOC collection of Oral History hosts currently more than 6.000 hours of
recording, corresponding to nearly 2.000 interviews. More than 90% of it, video or
audio, are in digital format. For the time being, two kinds of queries are available
in the database: query by subject and query by interviewed. Each interview
record holds a brief technical information and a summary with descriptions of
the interview themes in the order they appear in the record. Almost 80% of
the interviews are transcribed, and to access the audio/video content the user is
requested to come personally to CPDOC.</p>
        <p>Currently, CPDOC is analyzing better ways of making this data available
online, considering di erent aspects such as the best format, use policies, access
control and copyrights.</p>
        <p>As in the case of Acessus, the database actually stores only the metadata
about the interviews, while digitized recorded audios and videos are stored as
digital les in the le servers.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Brazilian Historical-Biographic Dictionary (DHBB)</title>
        <p>The Brazilian Historical-Biographic Dictionary (DHBB) is certainly one of the
main research sources for contemporary Brazilian politicians and themes. It
contains more than 7.500 entries of biographic and thematic nature, i.e., people,
institutions, organizations and events records carefully selected using criteria that</p>
        <sec id="sec-2-3-1">
          <title>4 https://en.wikipedia.org/wiki/File_server.</title>
          <p>measure the relevance of those to the political history for the given period. The
entries are written evenly, avoiding ideological or personal judgments. CPDOC
researchers carefully revise all entries to ensure the accuracy of the information
and a common style criteria.</p>
          <p>The DHBB's database stores few metadata concerning each entry, and the
query is limited to keywords within the title or text.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Current Status</title>
      <p>In this section we summarize the main problems identi ed in CPDOC's current
infrastructure and daily working environment.</p>
      <p>
        As described in Section 2, CPDOC's archives are maintained by three
different information systems based on traditional relational data models. This
infrastructure is hard to maintain, improve and re ne, and the information is
not found or accessed by standard search engines for two reasons mainly: (1) an
entry page does not exist until it is created dynamically by an speci c query;
(2) users are required to login in order to make queries or access the digital
les. Service providers do not access data directly and therefore cannot provide
specialized services using it. Users themselves are not able to expand the queries
over the collections, being limited to the available user interfaces in the website.
Thereupon, data of CPDOC's collections is currently limited to what is called
\Deep Web" [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The maintenance of current di erent information systems is very problematic.
It is expensive, time demanding and ine ective. Improvements are hard to
implement and therefore innovative initiatives are usually postponed. A relational
database system is not easily modi ed, because relational data models must be
de ned a priori, i.e., before the data acquisition's stage. Moreover, changes in
the database usually require changes in system interfaces and reports. The whole
work ow is expensive, time consuming and demands di erent professionals with
di erent skills from interface developers to database administrators.</p>
      <p>Concerning terminology, CPDOC's collections do not follow any metadata
standards, which hinders considerably the interoperability with other digital
sources. Besides, the available queries usually face idiosyncratic indexing
problems with low rates of recall and precision. These problems are basically linked
to the ad hoc indexing strategy adopted earlier to de ne database tables and
elds.</p>
      <p>Finally, data storage is also an issue. Digitized Acessus's documents and Oral
History's interviews are not stored in a single place, but scattered in di erent
le servers. The CPDOC database only stores the metadata and le paths to the
le servers, making it very di cult to ensure consistency between les, metadata
information and access control policies.
4</p>
    </sec>
    <sec id="sec-4">
      <title>The proposal</title>
      <p>As discussed in Section 3, relational databases are often hard to maintain and
share. Also, the idea of having in-house developed and closed source information
systems is being increasingly replaced by the concept of open source systems.
In such systems the responsibility of updating and creating new features is not
sustained by a single institution but usually by a whole community that share
knowledge and interests with associates. In this way the system is kept
upto-date, accessible and improving much faster due to the increased number of
contributors. Such systems are usually compatible with standards so as to ensure
they can be widely used.</p>
      <p>Our objective is to propose the use of modern tools so CPDOC can improve
the way they maintain, store and share their rich historical data. The proposal
focuses on open source systems and a lightweight, shared way of dealing with
data. Concretely, we propose the substitution of the three CPDOC systems by
the technologies described as follows.</p>
      <p>The Acessus data model comprises personal archives that contains one or
more series (which can contain also other series in a strati ed hierarchy) of
digitalized documents or photos. The PHO system data model is basically a
set of interviews grouped according to some de ned criteria within the context
given by funded projects. For instance, a political event could originate a project
which involve interviewing many important people taking part on the event.</p>
      <p>
        Therefore, Acessus and PHO systems can be basically understood as systems
responsible for maintaining collections of documents organized in a hierarchical
structure. In this way, one can assume that any digital repository management
system (DRMS) have all the required functionalities. Besides, DRMS usually
have desirable features that are not present in Acessus or PHO, such as: (1) data
model based on standard vocabularies like Dublin Core [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and SKOS [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]; (2)
long-term data preservation functionalities (tracking and noti cations of changes
in les); (3) ne-grained access control policies; (4) exible user interface for
basic and advanced queries; (5) compliance with standard protocols for repositories
synchronization and interoperability (e.g., OAI-PMH [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]); (6) import and
export functionalities using standard le formats and protocols; and more.
      </p>
      <p>In our proposal the data and les from Acessus and PHO systems are planned
to be stored in an open source institutional repository software such as Dspace 5
or Fedora Commons Framework 6. In this article we assume the adoption of
Dspace with no prejudice of theoretical modeling.</p>
      <p>As to DHBB, its relational model can be summarized to a couple of tables
that store metadata about the dictionary entries (stored in a single text eld of a
given table). The actual dictionary entries are created and edited in text editors
outside the system and imported to it only after being created and revised.</p>
      <p>The nature of its data suggests that DHBB entries could be easily maintained
as text les using a lightweight human-readable markup syntax. The les would</p>
      <sec id="sec-4-1">
        <title>5 http://www.dspace.org/</title>
      </sec>
      <sec id="sec-4-2">
        <title>6 http://www.fedora-commons.org</title>
        <p>be organized in an intuitive directory structure and kept under version control
for coordinated and distributed maintenance. The use of text les 7 is justi ed
by a couple of reasons. They are: easy to maintain using any text editor allowing
the user to adopt the preferred text editor (tool independent); conform to
longterm standards by being software and platform independent; easy to be kept
under version control by any modern version control system 8 since they are
comparable (line by line); and e cient to store information 9.</p>
        <p>The use of a version control system will improve the current work ow of
DHBB reviewers and coordinators, since presently there is no aid system for
this task, basically performed using Microsoft Word text les and emails. The
adoption of such tool will allow le exchanges to be recorded and the process
controlled without the need of sophisticated work ow systems, following the
methodology developed by open sources communities for open source software
maintenance. For instance, Git 10 is specially suited to ensure data consistency
and keeps track of changes, authorship and provenance.</p>
        <p>Many of the ideas here proposed were already implemented as a proof of
concept to evaluate the viability of such environment in CPDOC. Figure 2
illustrates the necessary steps to fully implement our proposal. In the following text
we brie y describe each step.</p>
        <p>
          Step (1) is implemented and the relational database was exported to RDF [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]
using the open source D2RQ [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] tool. The D2RQ mapping language [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] allows
the de nition of a detailed mapping from the current relational model to a graph
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>7 http://en.wikipedia.org/wiki/Text_file</title>
      </sec>
      <sec id="sec-4-4">
        <title>8 https://en.wikipedia.org/wiki/Revision_control.</title>
        <p>9 A text le of a DHBB entry has usually 20% the size of a le DOCX (Microsoft</p>
        <p>
          Word) for the same entry.
10 http://git-scm.com.
model based on RDF. The mapping from the relational model to RDF model
was already de ned using the standard translation from relational to RDF model
sketched out in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. The mapping created so far defers any model improvement
to step (2) described below.
        </p>
        <p>
          Step (2) is planned and represents a re nement of the graph data model
produced in step (1). The idea is to produce a data model based on standard
vocabularies like Dublin Core [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], SKOS [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], PROV [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] and FOAF [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and
well-known conceptual models like [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The use of standard vocabularies will
make the data interchangeable with other models and facilitate its adoption by
service provides and users. It will also help us to better understand the database
model and its semantics. In Section 5 we describe the re nement proposal in
detail.
        </p>
        <p>
          Step (3) is already implemented and deploys a text le for each DHBB entry.
Each text le holds the entry text and metadata 11. The les use YAML [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
and Markdown [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] markup languages to describe the metadata and entry
content. YAML and Markdown were adopted mainly because both languages are
human-readable markups for text les and are supported by almost all static
site generators 12. The use of a static site generator allows DHBB maintainers
to have full control over the deployment of a DHBB browsable version.
        </p>
        <p>Note that step (3) was actually implemented to use the initial version of the
RDF produced in step (1). The code can be easily adapted to use the nal RDF
model produced by step (2).</p>
        <p>In the planned step (4) the digital les and their metadata will be imported
into a DRMS. This step is much more easily implemented using the RDF
produced in step (2) than having to access the original database. It is only necessary
to decide which repository management system will be adopted.</p>
        <p>
          The proposed work ow architecture is presented in Figure 3. Recall that
one of the main goals is to make CPDOC archive collections available as open
linked data. This can be accomplished by providing data as RDF/OWL les for
download and a SPARQL Endpoint [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] for queries. Since data evolve constantly,
CPDOC teams would deliver periodical data releases. Besides the RDF/OWL
les and the SPARQL Endpoint, we believe that it is also important to provide
a lightweight and exible web interface for nal users to browse and query data.
This can be easily done using a static website generator and Apache Solr 13 for
advanced queries. As a modern index solution, Solr can provide much
powerful and fast queries support when compared to traditional relational database
systems. Note that the produced website, the RDF/OWL les and SPARQL
Endpoints are complementary outputs and serve to di erent purpose and users.
        </p>
        <p>Finally, it is vital to stress the main contrast between the new architecture
and the current one. In the current CPDOC architecture the data is stored in
relational databases and maintained by information systems. This means that
11 It is out of this article scope to present the nal format of these les.
12 In this application we used Jekyll, http://jekyllrb.com, but any other static site
generator could be used.
13 http://lucene.apache.org/solr/
any data modi cation or insertion is available in real time for CPDOC website
users. However, this architecture has a lot of drawbacks as mentioned in
Section 3, and also the nature of CPDOC data does not require continuous updates,
which means that the cost of this synchronous modus operandi is not needed.
Usually, CPDOC teams work on projects basis and therefore new collections,
documents and metadata revisions are not very often released.</p>
        <p>The results obtained so far encouraged us to propose a complete data model
aligned with open linked data vocabularies, presented in detail in next section.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Improving Semantics</title>
      <p>More than improving the current infrastructure for storing and accessing the
CPDOC's data, we would like to exploit the semantic possibilities of such rich
source of knowledge. One of the ways to do that is to embed knowledge from
other sources by creating links within the available data. Since much of the data
is related to people and resources with historical relevance, or historical events,
some available ontologies and vocabularies can be used in this task.</p>
      <p>
        The personal nature of the data allows us to use projects that are already
well developed for describing relationships and bonds between people, such as
FOAF [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] (Friend of a Friend) { a vocabulary which uses RDF to describe
relationships between people and other people or things. FOAF permits intelligent
agents to make sense of the thousands of connections people have with each
other, their belongings and historical positions during life. This improves
accessibility and generates more knowledge from the available data.
The analysis of structured data can automatically extract connections and,
ultimately, knowledge. A good example is the use of PROV [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], which provides a
vocabulary to interchange provenance information. This is interesting to gather
information of data that can be structurally hidden in tables or tuples.
      </p>
      <p>The RDF graph model enables also the merging of data content naturally.
The DBpedia project, for instance, allows users to query relationships and
properties associated with Wikipedia resources, and users can link other datasets to
the DBpedia dataset in order to create a big and linked knowledge knowledge
base. CPDOC could link their data to DBpedia and then make their own data
available for a bigger audience.</p>
      <p>
        In the same direction, the use of lexical databases, such as the WordNet [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
and its Brazilian version OpenWordnet-PT [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], will allow us to make
natural language processing of DHBB entries. Named entities recognition and other
NLP tasks can automatically create connections that improve dramatically the
usability of the content. Other resources like YAGO [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] and BabelNet [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] links
Wikipedia to WordNet. The result is an \encyclopedic dictionary" that provides
concepts and named entities lexicalized in many languages and connected with
large amounts of semantic relations. Finally, the SUMO Ontology [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] could also
be used to provide a complete formal de nition of terms linked to WordNet. All
of these lexical resources and ontologies will be further explored when we start
the natural language processing of DHBB entries.
The model presented in Figure 4 depicts that D2RQ was not able to
automatically improve much further the model. D2RQ was able to correctly translate
relations N:M in the relational model, such as entrevista entrevistador
(originally a table in the relational model) to a property that connect directly instances
of entrevista (interview) with instances of entrevistador (interviewer).
Nevertheless, the N:M relation between entrevista and tecnico (technician) was
kept as an intermediary class called tecnico entrevista due to the existence
of an aditional information in this N:M relation, the role (funcao class) of the
interview technician. The relational model also seems to have some
inconsistences. Although the connection of technician and interview is parameterized by
di erent roles, the donator, interviewer and interviewed of an interview are
modeled each one in a speci c table. In this case interviewed, interviewer, donator
and technician are all people that share a lot of common properties like name,
address, etc, and could be modeled as people. These problems are all result of
a \ad hoc" modeling process. The model de ned this way only makes sense for
CPDOC team and it could hardly be useful outside CPDOC.
      </p>
      <p>Figure 5 shows how PHO model can be re ned. The new model uses standard
vocabularies and ontologies, making the whole model much more
understandable and interoperable. In the Figure 5, prov:Activity was duplicated only for
a better representation. The pre xes in the names indicate the vocabularies and
ontologies used: prov, skos, dcterms, dc, geo, and bio. We also de ned a
CPDOC ontology that declares its own classes and speci c ontology links, such as
the one that states that a foaf:Agent is a prov:Agent. In this model, we see
that some classes can be subclasses of standard classes (e.g. Interview), while
some classes can be replaced by standard classes (e.g. LOCALIDADE).</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper we presented a new architecture for CPDOC archives creation
and maintenance. It is based on open linked data concepts and open source
methodologies and tools. We believe that even though CPDOC users would
need to be trained to use the proposed tools such as text editors, version control
software and command line scripts; this architecture would give more control
and easiness for data maintenance. Moreover, the architecture allows knowledge
to be easily merged to collections data without the dependency of database
refactoring. This means that CPDOC team will be much less dependent from
FGV's Information Management and Software Development Sta .</p>
      <p>
        Many proposals of research concerning the use of lexical resources for
reasoning in Portuguese using the data available in CPDOC are being carried out so
as to improve the structure and quality of the DHBB entries. Moreover, the
automatic extension of the mapping proposed in Section 5 can be de ned following
ideas of [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Due the lack of space, we do not present them in this paper.
      </p>
      <p>
        Finally, we aim to engage a wider community and an open-source
development process in order to make the project sustainable. As suggested by one of
the reviewers, we must also learn from experiences of projects like Europeana 14
and German National Digital Library [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
14 http://www.europeana.eu.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Alzira</given-names>
            <surname>Alves</surname>
          </string-name>
          <string-name>
            <surname>Abreu</surname>
          </string-name>
          , Fernando Lattman-Weltman, and Christiane Jalles de Paula.
          <source>Dicionario Historico-Biogra co Brasileiro pos-1930. CPDOC/FGV, 3 edition</source>
          ,
          <year>February 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Oren</given-names>
            <surname>Ben-Kiki</surname>
          </string-name>
          ,
          <article-title>Clark Evans, and Ingy dot Net. Yaml: Yaml ain't markup language</article-title>
          . http://www.yaml.
          <source>org/spec/1</source>
          .2/spec.html.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Michael</surname>
            <given-names>K</given-names>
          </string-name>
          <string-name>
            <surname>Bergman</surname>
          </string-name>
          .
          <article-title>White paper: the deep web: surfacing hidden value</article-title>
          .
          <source>journal of electronic publishing</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ),
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Tim</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <article-title>Relational databases on the semantic web</article-title>
          .
          <source>Technical report, W3C</source>
          ,
          <year>1998</year>
          . http://www.w3.org/DesignIssues/RDB-RDF.html.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          and Richard Cyganiak.
          <article-title>D2R server-publishing relational databases on the semantic web</article-title>
          .
          <source>In 5th international Semantic Web conference, page 26</source>
          ,
          <year>2006</year>
          . http://d2rq.org.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          , Jens Lehmann, Georgi Kobilarov, Soren Auer, Christian Becker, Richard Cyganiak, and
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Hellmann</surname>
          </string-name>
          .
          <article-title>Dbpedia - a crystallization point for the web of data</article-title>
          .
          <source>Web Semantics</source>
          ,
          <volume>7</volume>
          (
          <issue>3</issue>
          ):
          <volume>154</volume>
          {
          <fpage>165</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Dan</given-names>
            <surname>Brickley</surname>
          </string-name>
          and
          <string-name>
            <given-names>Libby</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Foaf vocabulary speci cation</article-title>
          . http://xmlns.com/ foaf/spec/,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Isabel</given-names>
            <surname>Cafezeiro</surname>
          </string-name>
          , Edward Hermann Haeusler, and
          <string-name>
            <given-names>Alexandre</given-names>
            <surname>Rademaker</surname>
          </string-name>
          .
          <article-title>Ontology and context</article-title>
          .
          <source>In IEEE International Conference on Pervasive Computing and Communications</source>
          , Los Alamitos, CA, USA,
          <year>2008</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Kendall</given-names>
            <surname>Grant</surname>
          </string-name>
          <string-name>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Lee</given-names>
            <surname>Feigenbaum</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Elias</given-names>
            <surname>Torres</surname>
          </string-name>
          .
          <article-title>SPARQL protocol for RDF</article-title>
          .
          <source>Technical report, W3C</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Nick</surname>
            <given-names>Crofts</given-names>
          </string-name>
          , Martin Doerr, Tony Gill, Stephen Stead, and Matthew Sti .
          <article-title>De nition of the CIDOC conceptual reference model</article-title>
          .
          <source>Technical Report 5.0</source>
          .4, CIDOC CRM Special Interest Group (SIG),
          <year>December 2011</year>
          . http://www.cidoc-crm.org/ index.html.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. Richard Cyganiak, Chris Bizer, Jorg Garbers, Oliver Maresch, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Becker</surname>
          </string-name>
          .
          <article-title>The D2RQ mapping language</article-title>
          . http://d2rq.org/d2rq-language.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. Valeria de Paiva, Alexandre Rademaker, and Gerard de Melo.
          <article-title>Openwordnet-pt: An open brazilian wordnet for reasoning</article-title>
          .
          <source>In Proceedings of the 24th International Conference on Computational Linguistics</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Governo</surname>
          </string-name>
          do Estado de Sa~o Paulo.
          <article-title>Governo aberto sp</article-title>
          . http://www.governoaberto. sp.gov.br,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>David</given-names>
            <surname>Eaves</surname>
          </string-name>
          .
          <article-title>The three law of open government data</article-title>
          . http://eaves.ca/
          <year>2009</year>
          / 09/30/three-law
          <article-title>-of-open-government-data/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Christiane Fellbaum, editor.
          <source>WordNet: An Electronic Lexical Database</source>
          . MIT Press, Cambridge, MA,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Natalja</surname>
            <given-names>Friesen</given-names>
          </string-name>
          , Hermann Josef Hill,
          <string-name>
            <given-names>Dennis</given-names>
            <surname>Wegener</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Doerr</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Kai</given-names>
            <surname>Stalmann</surname>
          </string-name>
          .
          <article-title>Semantic-based retrieval of cultural heritage multimedia objects</article-title>
          .
          <source>International Journal of Semantic Computing</source>
          ,
          <volume>06</volume>
          (
          <issue>03</issue>
          ):
          <volume>315</volume>
          {
          <fpage>327</fpage>
          ,
          <year>2012</year>
          . http://www. worldscientific.com/doi/abs/10.1142/S1793351X12400107.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Yolanda</given-names>
            <surname>Gil</surname>
          </string-name>
          and
          <string-name>
            <given-names>Simon</given-names>
            <surname>Miles</surname>
          </string-name>
          .
          <article-title>Prov model primer</article-title>
          . http://www.w3.org/TR/2013/ NOTE-prov-primer-
          <volume>20130430</volume>
          /,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. W3C OWL Working Group, editor.
          <source>OWL 2 Web Ontology Language Document Overview. W3C Recommendation. World Wide Web Consortium</source>
          ,
          <volume>2</volume>
          <fpage>edition</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>John</given-names>
            <surname>Gruber</surname>
          </string-name>
          . Markdown language. http://daringfireball.net/projects/ markdown/.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. Dublin Core Initiative.
          <article-title>Dublin core metadata element set</article-title>
          . http://dublincore. org/documents/dces/,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21. Open Data Initiative.
          <article-title>Open data initiative</article-title>
          . http://www.opendatainitiative. org,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Carl</surname>
            <given-names>Lagoze</given-names>
          </string-name>
          , Herbert Van de Sompel, Michael Nelson, and
          <string-name>
            <given-names>Simeon</given-names>
            <surname>Warner</surname>
          </string-name>
          .
          <article-title>The open archives initiative protocol for metadata harvesting</article-title>
          . http://www. openarchives.org/OAI/openarchivesprotocol.html,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. LexML. Rede de informaca~
          <article-title>o informativa e jur dica</article-title>
          . http://www.lexml.gov.br,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>Frank</given-names>
            <surname>Manola</surname>
          </string-name>
          and Eric Miller, editors.
          <source>RDF Primer. W3C Recommendation. World Wide Web Consortium</source>
          ,
          <year>February 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>James McGann</surname>
          </string-name>
          .
          <source>The Think Tank Index. Foreign Policy</source>
          ,
          <year>February 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>Alistair</given-names>
            <surname>Miles</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sean</given-names>
            <surname>Bechhofer</surname>
          </string-name>
          .
          <article-title>Skos simple knowledge organization system reference</article-title>
          . http://www.w3.org/
          <year>2004</year>
          /02/skos/,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <article-title>Roberto Navigli and Simone Paolo Ponzetto</article-title>
          .
          <article-title>BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network</article-title>
          .
          <source>Arti cial Intelligence</source>
          ,
          <volume>193</volume>
          :
          <fpage>217</fpage>
          {
          <fpage>250</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <given-names>Ian</given-names>
            <surname>Niles</surname>
          </string-name>
          and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Pease</surname>
          </string-name>
          .
          <article-title>Towards a standard upper ontology</article-title>
          .
          <source>In Proceedings of the international conference on Formal Ontology in Information Systems-Volume</source>
          <year>2001</year>
          ,
          <article-title>pages 2{9</article-title>
          . ACM,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Fabian M. Suchanek</surname>
            , Gjergji Kasneci, and
            <given-names>Gerhard</given-names>
          </string-name>
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          .
          <article-title>Yago: A Core of Semantic Knowledge</article-title>
          .
          <source>In 16th international World Wide Web conference (WWW</source>
          <year>2007</year>
          ), New York, NY, USA,
          <year>2007</year>
          . ACM Press.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>