=Paper=
{{Paper
|id=None
|storemode=property
|title=Integration of Intelligence Data through Semantic Enhancement
|pdfUrl=https://ceur-ws.org/Vol-808/STIDS2011_CR_T1_SalmenEtAl.pdf
|volume=Vol-808
|dblpUrl=https://dblp.org/rec/conf/stids/SalmenMHCS11
}}
==Integration of Intelligence Data through Semantic Enhancement==
Integration of Intelligence Data through Semantic
Enhancement
Salmen David Malyuta Tatiana Hansen Alan Cronen Shaun Smith Barry
dsalmen@data-tactics.com tmalyuta@data-tactics.com alan.hansen1@us.army.mil shaun.cronen@us.army.mil phismith@buffalo.edu
Data Tactics Corp. Data Tactics Corp., Intelligence and Intelligence and National Center for
City University of New Information Warfare Information Warfare Ontological Research,
York Directorate Directorate University at Buffalo
Abstract—We describe a strategy for integration of data that is As a first step towards meeting these requirements we
based on the idea of semantic enhancement. The strategy introduced in 2009 the Data Representation and Integration
promises a number of benefits: it can be applied incrementally; it Framework (DRIF) [1, 2], which presents minimal barriers to
creates minimal barriers to the incorporation of new data into the incorporation of new data into a data resource, thus
the semantically enhanced system; it preserves the existing data
requiring no heavy pre-processing and no data or data-model
(including any existing data-semantics) in their original form
(thus all provenance information is retained, and no heavy pre- conditioning. DRIF embraces the full spectrum of data
processing is required); and it embraces the full spectrum of data sources, types, models, and modalities, including text, images,
sources, types, models, and modalities (including text, images, audio, and signals, while supporting a variety of integration
audio, and signals). The result of applying this strategy to a given and analytic processes and tools. Details are presented below.
body of data is an evolving Dataspace that allows the application
of a variety of integration and analytic processes to diverse data The Dataspace store of intelligence data which is the subject
contents. We conceive semantic enhancement (SE) as a light- of this communication is the result of applying the DRIF to the
weight and flexible process that leverages the richness of the task of integrating very large heterogeneous primary data
structured contents of the Dataspace without adding storage and
artifacts. As the Dataspace has evolved through time, so it has
processing burdens to what, in the intelligence domain, will be an
already storage- and processing-heavy starting point. SE works incorporated progressively ever larger quantities of data, and
not by changing the data to which it is applied, but rather by also more specific local implementations and data structures
adding an extra semantic layer to this data. We sketch how the used by data analysts, some of which bring their own data
semantic enhancement approach can be applied consistently and semantics. For the purposes that the Dataspace is intended to
in cumulative fashion to new data and data-models that enter the serve, it is vital that no restrictions are imposed either on the
Dataspace. types of source-artifacts and the associated models and media
within the Dataspace, or on the processes by which the
Keywords: integration, intelligence data, ontology, semantic Dataspace is populated (whether by loading structured data
technology.
from a database, by extraction from a text document through
some Natural Language Processing application, by automatic
I. INTRODUCTION analysis of signals, or through inference by a human analyst).
The success of the war fighter and homeland defender in the
Net-Centric Warfare environment is largely defined by the The design of the Dataspace is such that it can incorporate
ability to quickly acquire and efficiently and accurately process hundreds of millions of unstructured documents and similarly
intelligence information from numerous heterogeneous sources large quantities of images, signals data, and other structured
of different structure and modality. Traditional data integration and unstructured primary data artifacts. Each of these artifacts,
approaches fail in the face of the scale, diversity, and when it enters the Dataspace, is represented through a set of
heterogeneity of intelligence data sources and data-models metadata, including labels specifying image type, MIME type,
because they fail to address one or more of the following and so forth, as well as provenance information. Further
requirements: processing may, for example, associate pixels in an image
• Integration must proceed without heavy pre-processing with the name of a person, or a range of characters in an
• Integration must proceed regardless of the data-models unstructured text document with the name of a location, or
used (or not used) in the data sources to be integrated, extract a cell from a database table. The DRIF provides a
• Integration must proceed regardless of the data modality, common framework in which the results of all of these
and without loss or distortion of data, of its associated processes are represented in a unified way, details of which
data semantics, and of data-provenance information, are provided below. As a result, primary data can be utilized
• Integration must involve the ability to incorporate immediately upon entering the Dataspace for a variety of
multiple points of view on the data to be integrated, different kinds of search and more sophisticated processing
including different views of the data, for example on the based thereon. DRIF is not, however, a magic bullet; many
part of different analysts using different analytical tools. issues of data integration at the syntactic level will remain,
arising for example as a result of data formats which do not forms, which we call primary and derived, respectively. The
match, where we will need to normalize the format into an Dataspace is divided into corresponding segments (see Figure
augmented model that will serve as the target of annotations. 1) in a way that supports a comprehensive approach to
This will involve considerable effort to ensure that the needed integration that allows accommodation of the multiple views
actions are performed promptly and consistently whenever of the primary and derived data and of the associated data-
new data comes in. Here, however, we focus exclusively on semantics and metadata which arise for example as a result of
those issues which arise at the stage of what we can loosely the workings of multiple different sorts of analytical tools.
call the ‘representational’ aspects of data integration.
A. Approach to Integration
Some primary artifacts within the Dataspace already Our approach to integrating intelligence data starts with source
incorporate useable semantic content – for instance a artifacts consisting of primary data across a variety of
structured database which incorporates meaningful column representation modalities. This primary data is weakly
headers, or a message with a structured payload incorporating integrated in the sense that indexes are provided to support
meaningful tags. But such content is ad hoc. It is tied to simple (string-based) data search across all primary artifacts.
specific local implementations and typically falls short of what
is needed to secure semantic interoperability of the Some primary data comes with its own native structure, and
implementations involved because of the absence of a further structure will typically be added thorough analytical
common formally coherent approach to semantics and of a processing. The second integration step addresses the need for
common governance process. the unified storage of this structured data to support more
complex structured search across both primary and derived
Moreover, full semantic integration is in any case prevented artifacts.
by the needs of openness of the Dataspace to ever new sorts of
primary data and analytically derived data. It is to compensate Importantly, we here embrace the diversity of domain-specific
for this problem that we have developed our strategy for data-models employed throughout the Intelligence Community
semantic enhancement. We start out from the assumption that while at the same time reaping benefits from an approach that
semantic data enrichment can be achieved only incrementally, is data-model agnostic. This is because the unified
through the step-by-step creation of ontology modules that are representation provided by the DRIF allows analytic
designed in coordinated fashion to work well both with each processing of data in highly diverse primary artifacts
other and with specific bodies of Dataspace content. The associated with different native data-models to be used as
vision is a lightweight, flexible approach comprising an extra targets of cross-artifact analytics. For example, and most
ontology layer that leverages the contents of the Dataspace simply, it is possible to perform unrestricted string search
without adding storage and processing weight to what is an across structured artifacts of highly different sorts. Examples
already storage- and processing-heavy resource. We discuss of more sophisticated analytics include computer-aided data-
the details of semantic enhancement in section IV. First, model harmonization, for example by allowing significant
however, we introduce the DRIF and the Dataspace to which overlap of sets of values of attributes from different databases
the SE strategy will be applied. to be flagged by the analytic process as a potential indication
that the attributes have the same meaning, thereby making it
II. DATA REPRESENTATION AND INTEGRATION possible for the relevant portions of the two databases to be
FRAMEWORK enriched through fusion.
Our starting point is a body of U.S. Department of Defense B. Dataspace Organization
(DoD) intelligence data within what we are here calling the
The organization of the Dataspace is schematically illustrated
Dataspace. The implementation in the specific context upon
in Fig. 1.
which we focus here is engineered around cloud computing
paradigms and is primarily based upon open-source cloud
Segment 0 is a store of primary artifacts, including documents,
software stack components. This cloud computing foundation
images, signals, and analysts’ work products vetted for re-use
leverages advantages of linear scaling and parallel distributed
as input for further processing. The physical implementation
computation when faced with the reality of ever increasing
of Segment 0 may be such that all data is stored internally; or
data volumes and integration processing. All the work
it may be distributed, so that source artifact data may for
described is either deployed or in the final stages of testing
example be either contained in the cloud store or stored
prior to deployment.
externally to the Dataspace and referenced in the cloud store.
Primary data vary widely by nature; they may have different
The Dataspace is built using the Data Representation and
structures (for example of a relational database), or they may
Integration Framework (DRIF), which has been designed to
be unstructured (for example, free text, audio or video files),
represent large quantities of data in a form that is useful to the
and they may be of different modalities (for example they may
end user both for direct inspection and for the application of
be cells of a relational database, audio sequences, assertions of
various kinds of analytics. Representations of source data
an analyst).
artifacts and their contents within the Dataspace are of two
• Segment 3 describes the data-models which support the
Segment 3 – Structured descriptions of models associated with two sets of views just mentioned as well as synoptic
primary and derived artifacts (e.g. attributes of a relational table views (ultimately including SE-based views) of the type
with associated functional dependencies) which can foster harmonization.
D. Where Models and Primary Data Come Together
Segment 1 – Structured Segment 2 – Structured
descriptions of primary descriptions of data (results We believe that the principal contribution of the Dataspace
artifacts of processing of primary endeavor is to resolve certain problems of storage and thus of
(e.g. registration data; data, e.g. through NLP, representation, enrichment, and evolution of large bodies of
relations to other artifacts) image analysis, or other data. The goal is to provide room for both primary data and
data extraction)
the multiple results of processing these data by different
analysts or analytic methods. To achieve this we introduced in
Segment 0 – Primary artifacts (stored in CloudBase, HDFS [3] a strategy for description of data that is designed to enable
and elsewhere) (including documents, images, signals, true data integration across a constantly evolving and highly
databases; as well as analysts’ products, some with built in heterogeneous resource comprehending extremely large
data models) volumes of data. As already recognized at the very beginning
of contemporary high-level research in biomedical ontology
[4], this end can be achieved only if data are exposed in a way
Figure 1. Organization of the Dataspace. Solid line: registration processes; that is independent of their original intended use. This must
curved solid lines: processes that ingest artifacts into the Dataspace, involve some means to represent original data-models at a
including feeding back into the Dataspace analysts’ products – results of the
Dataspace processing ; dashed lines: derivation processes.
level of abstraction that is higher than that of primary data. We
accordingly propose an abstract data-model based on five core
Segment 1 includes primary artifact registration data as well as elements: sign, concept, term, predicate, and statement, which
specifications of relations between artifacts (for example, we believe is sufficient to represent any data-model in these
nesting of an image within a document, or attachment of one terms.
document to another). Segment 1 will include also data
pertaining to the way each derived artifact of Segments 2 and Sign: A sign gi is a string that is the abstracted proxy within
3 is derived from primary artifact(s) in Segment 0. the dataspace for one or more chunks of data used in some
primary artifact with the intention of referring to some
Segment 2 stores the structured data that is either already individual entity (e.g. person, location, organization, object,
present in primary artifacts or derived therefrom through event). Examples include: a sign of the type proper name that
analytic processing resting on data-models represented in is associated with an expression (for example ‘he’ or ‘Dr.
Segment 3. Watkins’ occurring in a document; a label annotating an area
in a pixel array as forming an image of some building; a label
Segment 3 stores the descriptions of the data-models used in annotating a fragment of an audio stream or other signal as
Segment 2. These data-models may include database schemas, recording some explosion event. Each sign is associated with
message formats, or XML schemas. The data-models one or more physical extents within those primary artifacts
themselves are primary artifacts and are thus stored in with which it is associated, which we call mentions (the latter
Segment 0 and registered in Segment 1. are what are elsewhere called tokens). The collection G = {gi}
comprehends all signs extracted from primary data artifacts
The Dataspace is evolving continuously not only because of and changes with the incorporation of new artifacts.
new primary data ingested from the outside, but also because
new artifacts are being created, for example, through analysts’ Concept: A concept ci is (for the purposes of this exposition) a
reports based on processing of existing data. These artifacts string that is used in the Dataspace to represent some general
themselves have a status of new primary artifacts. category or grouping. The purpose of concept strings is to
represent and allow reuse of classifications native to primary
C. Segments as Abstractions Over the Artifacts artifacts. Concepts are taken from data-models registered in
Each of Segments 1-3 is an abstraction over the corpus of Segment 1. Examples of concepts are: the classes of an
primary data artifacts (Segment 0) and supports analytics of a ontology such as UCore SL, the tag set in an XML Schema
particular type: Document (XSD), and the attribute or table names in a
• Segment 1 is a high-level view of the entire artifact relational database. The collection C = {ci} comprehends all
corpus including the relations between the artifacts, but concepts within the Dataspace and changes as new data-
with no reference to their internal contents. models are incorporated.
• Segment 2 is a collection of detailed views of the internal
contents of the artifacts at the level of individual data Term: A term, tij, is an ordered pair of strings , where gi
items. ∈ G and cj ∈ C. Each term results from a process of contextual
disambiguation of a sign, a process which associates a sign
with a concept, as in <123-45-6789, SSN>. The collection T =
{tij} comprehends all terms identified by analytic processing of current phase of evolution of DRIF, the phase of Semantic
primary artifacts. Enhancement (SE). SE, as we conceive it, is a light-weight
and flexible solution that leverages the richness of the native
Predicate: A predicate (by which we mean here always: source data and of any local semantics associated with these
binary relational predicate) pi is a string that is used to data without adding storage and processing weight. The SE
connect terms in accordance with domain and range strategy is compliant with and complements the DRIF.
constraints. Predicates are used in the formation of statements
Artifacts
(as described below). Examples of predicates are: hasSSN,
hasLocation, hasBirthDate. Predicates are derived from data- Database A Database B
models registered in Segment 1, for example from table ID PersonName PersonID Name Address
… …
column headings or from XML tags. The collection P = {pi} 732 Bill 821 William DC
comprehends all predicates within the Dataspace and changes
as new data-models are added. Document X
….Scott performed the database backup…
Statement: A statement si is an ordered triple consisting of a
subject, a predicate, and an object. The collection S = {si} of
statements is recursively defined. At the lowest level,
Structured content
statements are ordered triples consisting of a term, a predicate,
Sign Concept Predicate
and a second term. In higher-level statements, subjects and
objects may be lower-level statements. Examples: <[Bruno, key label key label key label
1 732 1 ID 1 hasName
PersonName] hasSSN [123-45-6789, SSN]> 2 Bill 2 Scientist 2 hasAddress
3 821 3 PersonID 3 sameAs
The five primitives of the DRIF (sign, concept, predicate, 4 William 4 Name 4 knows
term, and statement) define a data reference model which, by 5 DC 5 Address
effectively decoupling data from data-models, can represent 6 Scott 6 DBA
any sort of data-model at the level that is useful for Term
integration. key sign_Key concept_Key
1 [732, ID] 1 1
Fig. 2 schematically illustrates the representation of structured 2 [Bill, Scientist] 2 2
data in accordance with the DRIF for three sample primary 3 [821, PersonID] 3 3
4 [William, Name] 4 4
artifacts, two of them relational databases, the third an 5 [DC, Address] 5 5
unstructured document. The example also shows how data- 6 [Scott, DBA] 6 6
semantics come to be added to the Dataspace in ad hoc
Statement
fashion – here, because an analyst decides to to introduce a
new Concept DBA (meaning: database administrator). key term_Key_Subject predicate_Key term_Key_Object
1 3 [821, PersonID] 1 hasName 4 [William, Name]
Additional Statements establishing relationships between 2 3 [821, PersonID] 2 hasAddress 5 [DC, Address]
Terms using Predicates SameAs and Knows are also included 3 3 [821, PersonID] 3 sameAs 1 [732, ID]]
in the Figure. 4 3 [821, PersonID] 4 knows 4 [Scott, DBA]
The reader familiar with the Resource Description Framework
(RDF/RDFS) may wonder what is different here. RDF Figure 2. Simplified example of structured content derived from 3
primary artifacts.
employs a similar level of abstraction, but it is a language,
while what we are offering here is a specific, albeit still highly
abstract, data-model. This data-model could of course be A. Goals of Semantic Enhancement
specified very easily using the RDF language; but it could be SE is a strategy that is currently being implemented to
specified also using relational database or some other storage improve our handling of the enormous heterogeneity of
technology. Our choice of data-model was motivated further Dataspace content. It is centered on building a flexible and
by the fact that our implementation and security requirements extensible framework of hierarchically organized, controlled
dictated the use of a specific type of cloud storage solution [5, structured vocabularies – called ‘ontologies’ – covering
6] that is both highly scalable and offers highly granular different areas of relevance to intelligence analysis. The
security access controls. framework will be constructed in part by reusing already
existing resources, in part through collaboration with other
III. SEMANTIC ENHANCEMENT defense and military organizations in the creation of new
ontology modules. The ontologies will be used in an
The DRIF focuses on the representational aspects of the incremental process of annotation (or ‘tagging’) of those
Dataspace and on the basic types of data integration that such concepts and predicates already identified in data-models
representation provides. In what follows we describe the within the Dataspace along the lines described in our
discussion of Segment 3 above. The latter amount to what we subtype (or is_a) hierarchy; and second through the
referred to above as ‘ad hoc semantics’. Because the salient progressive incorporation in all nodes of the SE ontologies of
data-models derive from so many heterogeneous sources, they links to relevant synonyms derived through the annotations
use a multiplicity of partially overlapping and partially which will link ontology nodes to the rich collection of
conflicting vocabularies, which it is the task of SE to reconcile corresponding concepts and predicates in other areas of the
by associating co-referring concepts and predicates (strings) Dataspace.
employed within distinct data-models in the Dataspace to
C. The Strategy for Semantic Enhancement
single nodes within the external SE ontologies.
Our strategy is designed to achieve its goals not by changing
To function in the needed way, annotations must be the Dataspace, but rather by adding an extra semantic layer
cumulative, in the sense that our strategy will ensure that tags thereto. The strategy is thus similar to that underlying the
created by different annotators will be consistent with each Universal Core (UCore), which arose out of the National
other. The value of annotations must also be preserved when Information Sharing Strategy supported by multiple U.S.
the SE ontologies change, for example through refinements Federal Government Departments, by the intelligence
created to reflect advances in knowledge, and to this end the community, and by a number of other national and
ontologies must be subject to strict versioning policies. international organizations [7, 8]. Here, a small controlled
vocabulary was provided for multi-community use to associate
Finally, the SE framework must be implemented in such a way simple summary tags to message payloads for purposes of data
that it can serve not merely as a tool of harmonization of the search and integration.
data-models internal to the Dataspace but also in a way that
allows integration with other, external data resources wherever Reflecting the extreme diversity of intelligence data, multiple
common ontologies are used for annotation. subject-matter expert communities will be contributing to the
SE. For the strategy to work and provide useful and efficient
To address these constraints is by no means a simple matter. integration, these multiple distributed teams must use the SE
When data value codifications do not match – for example approach in a consistent fashion. Previous efforts to create a
when we have 1,2,3 in one data source, R, G, B in another data broad-based, multi-community ontological approach to data
source, and RED, GREEN, BLUE in our Color ontology, then integration in defense and intelligence domains have failed
annotation for each source to hierarchy values can be very because the incompatible, and often over-simplistic, views of
labor intensive and require significant SME effort. reality incorporated into legacy databases and data-models led
to incompatible development of ontologies in ways that
B. Sample Benefits of Semantic Enhancement precluded interoperability. Many advocates of semantic
We can see the sorts of benefits that SE will provide already at approaches to data integration have still failed to appreciate
the level of search, where problems arise because of the the tremendous challenges, both technical and human, created
multiple different ways of describing data within the by the entrenched predisposition on the part of ontology
Dataspace. Problems that need to be confronted include: developers to create ontologies each on the basis of their own
potentially idiosyncratic data representations.
1. The need to find data items identified by means of terms
which are narrower or broader in meaning than the terms The solution which we advocate is modeled on the successful
analysts will standardly use when searching; semantic annotation approach pioneered in the field of
2. The need to find data items in documents that are bioinformatics by the Gene Ontology [9]. This approach is
formulated using a language or technical jargon with which now being pursued systematically within the framework of the
analysts are unfamiliar. OBO Foundry [10, 11], which starts out from the idea that the
most effective way to ensure mutual consistency of ontologies
To provide some very simple examples: we know that a given created by multiple independent groups over time and to
package ‘has been shipped with a red label’, but the ensure that these ontologies are maintained in such a way as to
documents that we have pertaining to this package use only keep pace with advances in knowledge is to organize
the word ‘vermillion’; or we need to find references to a ontologies as a collection of modules with discrete (non-
package identified as ‘containing furniture’, but the documents overlapping) subject-matters maintained by subject-matter
we have refer only to ‘chairs’; or we need to find a given experts, according to a strategy outlined in [12]. To ensure
package suspected of containing crack cocaine, but the audio consistency, these ontologies should be created as extensions
recordings we have at our disposal relating to this package of more generic higher level ontologies, subject to common
refer only to ‘bobo’ or ‘botray’ or ‘boubou’. If we are rules for example concerning the treatment of definitions, and
restricted to string search, our queries would not return the they should be based on a small common upper-level ontology
needed results. Hence, we need a framework which expands (ULO), whose domain and content neutral. For example, it
string search by capturing type and subtype information, and will include relations such as is-a (for subtype), member-of,
also incorporates synonym information. These needs are part-of, and so on. As initial ULO we choose the Basic Formal
targeted along two dimensions; first, through the fact that all Ontology (BFO) [13], which has been implemented in more
SE ontologies will be organized around a central backbone
than 100 similar projects, and which serves as the basis of the D. Implementation of the SE Strategy
already mentioned UCore Semantic Layer [8]. We can now outline the steps which are involved in realizing
this strategy in the specific context of the Dataspace, where
The ULO will be associated with a small number of Mid- we already have data structured using the DRIF.
Level Ontologies (MLOs) defined by downward population
from the ULO. The MLOs will serve in turn as bridge to a First Step: Review the contents of the Dataspace, specifically
number of Low-Level Ontologies (LLO), which will specify that concepts and predicates in Segment 1, and identify a
narrow content domains. Each MLO represents cross-domain subset of topic areas where data integration is a priority for
entities, such as Person or Information, and will be constructed
analytics.
in tandem with the LLOs which it subsumes in order to ensure
the mutual consistency and interoperability of the subsumed
Second Step: Formulate a list of MLOs that would be needed
LLOs. The MLOs and LLOs must in turn be associated with
the resources of a relation ontology, providing for the to annotate the data in corresponding areas. As far as possible
representation of content-specific relations such as Owns, identify existing ontologies which may potentially be reused
WorksFor, Audits, and so on. for this purpose, and build initial versions of new ontologies
where needed.
Initial due diligence efforts in our strategy of semantic
enhancement requires us to identify an initial collection of Third Step: Identify a specific subset of the content of the
authoritative codifications at Mid- and Lower Levels – along source data-models, and identify LLOs that will capture this
roughly the lines depicted in Table 1 – and to begin the subset in a semantically coherent fashion, ensuring that each
process of formalizing them within the BFO common upper- LLO is subsumed by some MLO. Subject matter experts
level ontological framework. In some areas ontologies will should be recruited to take charge of creation and maintenance
need to be created de novo, since no adequate authoritative of the LLOs and MLOs and of their use in annotations. In this
codifications will exist. way we can create a cadre of SMEs with expertise in
annotation and in supporting semantic enhancement.
Examples of MLO cross-domains
In realizing the above we need to maximize as far as possible
• Geospatial
the reuse of ontologies which are already being used by
• Biometrics
• Person relevant communities. This is because the strategy will be
• Provenance and Trust successful only to the degree that a critical mass of potential
• Organization users are able to be convinced of its utility and thus
• Signals and Sensors incentivized to engage in advancing it further for example by
• Equipment extending it new types of data and by disseminating the
• Facility resource to new groups of analysts. Reusing already existing
Examples of LLO domains ontologies will not merely provide a core of familiar terms
which analysts can use for search purposes, it will also
Subsumed by Geospatial
• Geospatial Feature increase the degree to which we can integrate into the
• Country Dataspace data that has already been annotated in consistent
Subsumed by Biometrics fashion by external bodies.
• Fingerprint
• Iris Fourth Step: When once a stable, initial set of ontologies has
Subsumed by Person been created, we use these ontologies to annotate the data-
• Employment Data models in corresponding portions of the Dataspace. As should
• Criminal Data by now be clear, the entire strategy is an incremental one,
• Medical Data based on a principle of low hanging fruit: the idea is not to
• Ethnicity and Tribe import the above ontologiesas a whole; rather we examine
• Skill
the existing Dataspace resources and identify expressions
Subsumed by Provenance and Trust therein for which counterparts in the ontologies already exist
• Data Quality
• Access Permissions or can easily be added. In constructing the ontologies these
• Data Source expressions will be provided with a common logical
• Evidence architecture and a common set of relations defined through
the ULO top level and in terms of which logical definitions
Table 1: Sample Ontologies within the SE for terms in the ontologies can then be formulated. The result
Structure can be used as a basis for the application of general-purpose
tools, including standard OWL reasoners FaCT++, RACER,
or Pellet, which can be used to check ontologies in the SE
resource for mutual consistency.
particular subset of a source data-model is too general for data
Stage 4 of the SE process consists in associating each set of analyst purposes, then the respective LLOs can be further
equivalent data source concepts with a single common MLO developed as needed.
or LLO expression (which will be added at the appropriate
level within the SE ontology structure where not already
Person
present). Further types of integration are thereafter brought
about automatically. Whenever any Dataspace resource Ethnicity Skill
becomes linked to one of our chosen ontologies in a way that
Computer
can be used to generate corresponding annotations, it thereby Skill
becomes linked to all the other Dataspace ( a n d
e x t e r n a l ) resources that have already been annotated with Programming Network
the same SE ontologies. This creates a snowball effect, Skill Skill
whereby each new annotation increases the value of existing Works Audits
annotations [9], and provides further incentives for the use of
the SE ontologies by new groups of users. Facility
Educational Used for
E. Organization of the SE Ontologies Hospital
annotations
Facility
Fig. 3 illustrates the organization of the SE ontology space. of source
data-
Each LLO represents the reality of a particular narrowly
models
defined domain, for example in an area such as Education and
Skills.
Middle Is-a
An MLO is a container of LLOs. Since we will be developing
LLOs in step-by-step fashion to address what are at any given Lower attribute-of
time the most urgent needs of Dataspace users, there will be semantic
data which cannot as yet be annotated with the full granularity
of detail which the annotator requires. The strategy is to use Figure 3. Simplified Example of an SE Ontology Structure.
such cases to advance the further development of the ontology
resource base, again following the model tested in the
bioinformatics domain [9]. For example, an analyst may want
IV. CONCLUSION
to use the SE resources to extract and disambiguate data from
a particular document. For different reasons the analyst may Together, the DRIF and SE provide what we believe is a
not be able to use the most detailed semantics and will use a workable data-integration solution. The DRIF is a highly
more general one. LLO taxonomies will also be used by flexible framework, with few constraints and including an
analytics to produce results of different level of detail: from a RDF-style decomposed representation of structured data
fine-grained view of narrow areas within the Dataspace to which allows the collection of data resources without loss or
coarse grained pictures of larger domains. distortion in a way that achieves syntactic integration and
preserves the local semantics of primary sources and of
Because original data and data-semantics are in every case analytics software. SE provides semantic integration in a light-
preserved without loss or distortion in the Dataspace as it weight yet incrementally extendible fashion, and in a way that
exists prior to Semantic Enhancement, there is no need to can foster global integration without adding storage and
represent all details of original storage data structures in the processing weight to already storage- and processing-heavy
SE stage. This means that complex ontologies are not needed Dataspace.
– a common and shared vocabulary is sufficient for virtual
semantic integration and search/analytics, while underlying The SE approach provides a strategy to allow the Dataspace to
details are maintained by the authors of specific primary be understood as evolving cumulatively as it accommodates
artifacts. Similarly, the collection of SE ontologies does not new kinds of data. It provides a more consistent,
need to cover all of the ad hoc local semantics within the homogeneous, and well-articulated representation of
Dataspace – content that is unlikely to be used in search or is structured content that originates in multiple internally
not important for integration can be excluded from the inconsistent and heterogeneous models. And while it involves
Enhancement step, since it will still be available in the source considerable initial SME investment in ontology creation and
data-models and can be accessed when drilling down to the annotation, we believe that it will allow the management and
appropriate level. exploitation of the Dataspace to become more cost-effective
over time.
The SE approach is highly flexible. It represents a “pay-as-
you-go” approach in the sense that investments can be made In addition, the use of the selected MLOs and LLOs brings
only in specific areas according to identified need. It is also integration with other government initiatives and brings the
tunable in the sense that, if a given body of annotations for a Dataspace endeavor closer to the federally mandated net-
centric data strategy; it also makes the integrated Dataspace [4] Rosse, C. and Mejino, J. L. V. A Reference Ontology for
Bioinformatics: The Foundational Model of Anatomy. Journal of
more effectively searchable and provides an expanding body Biomedical Informatics 36, 2003, 478-500.
of content to which more powerful analytics can be applied in [5] R6 Cloudbase documentation and source code.
the future. [6] Hadoop. http://hadoop.apache.org.
[7] http://ncor.us/ucore-sl.
Acknowledgments: This work was funded by US Army [8] B. Smith, L. Vizenor and J. Schoening, “Universal Core Semantic
CERDEC I2WD. The authors thank Mr. Kesny Parent, Layer”, Ontology for the Intelligence Community, Proceedings of the
DCGS-A Branch Chief, for continued support. Third OIC Conference, George Mason University, Fairfax, VA, October
2009, CEUR Workshop Proceedings, vol. 555.
[9] D. Hill, et al., “Gene Ontology Annotations: What they mean and where
they come from”, BMC Bioinformatics, 2008; 9(Suppl 5): S2.
REFERENCES [10] B. Smith, et al., “The OBO Foundry: Coordinated Evolution of
Ontologies to Support Biomedical Data Integration”, Nature
Biotechnology, 25 (11), November 2007, 1251-1255.
[1] S. Yoakum-Stover, T. Malyuta, N. Antunes, “A Data Integration [11] B. Smith, W. Ceusters “Ontological Realism: A methodology for
Framework with Full Spectrum Fusion Capabilities”, Presented at the coordinated evolution of scientific ontologis”, Applied Ontology 5
Sensor and Information Fusion Symposium, Las Vegas, NV, Aug 3-7, (2010) 139-188, http://x.co/adRJ.
2009.
[12] W. Ceusters, B. Smith, J. M. Fielding, “LinkSuiteTM: formally robust
[2] A. Hansen, D. Salmen, T. Malyuta, and N. Antunes. “An Evolving ontology-based data and information integration,” in Database
Integrated Dataspace on the Cloud.” Presented at the Sensor and Integration in Life Sciences, Berlin, Springer, 2004.
Information Fusion Symposium, Las Vegas, NV, July 26-29, 2010. http://ontology.buffalo.edu/bio/LinkSuite.pdf
[3] S. Yoakum-Stover, T. Malyuta, “Unified Integration Architecture for [13] Basic Formal Ontology. http://www.ifomis.org/bfo/.
Intelligence Data”, Proceedings of DAMA International Europe
Conference, London, UK, 2008.