=Paper= {{Paper |id=None |storemode=property |title=Integration of Intelligence Data through Semantic Enhancement |pdfUrl=https://ceur-ws.org/Vol-808/STIDS2011_CR_T1_SalmenEtAl.pdf |volume=Vol-808 |dblpUrl=https://dblp.org/rec/conf/stids/SalmenMHCS11 }} ==Integration of Intelligence Data through Semantic Enhancement== https://ceur-ws.org/Vol-808/STIDS2011_CR_T1_SalmenEtAl.pdf
      Integration of Intelligence Data through Semantic
                         Enhancement
    Salmen David                 Malyuta Tatiana                 Hansen Alan               Cronen Shaun                Smith Barry
 dsalmen@data-tactics.com      tmalyuta@data-tactics.com    alan.hansen1@us.army.mil   shaun.cronen@us.army.mil     phismith@buffalo.edu
   Data Tactics Corp.           Data Tactics Corp.,           Intelligence and           Intelligence and          National Center for
                              City University of New        Information Warfare        Information Warfare        Ontological Research,
                                      York                       Directorate                Directorate           University at Buffalo

Abstract—We describe a strategy for integration of data that is         As a first step towards meeting these requirements we
based on the idea of semantic enhancement. The strategy                 introduced in 2009 the Data Representation and Integration
promises a number of benefits: it can be applied incrementally; it      Framework (DRIF) [1, 2], which presents minimal barriers to
creates minimal barriers to the incorporation of new data into          the incorporation of new data into a data resource, thus
the semantically enhanced system; it preserves the existing data
                                                                        requiring no heavy pre-processing and no data or data-model
(including any existing data-semantics) in their original form
(thus all provenance information is retained, and no heavy pre-         conditioning. DRIF embraces the full spectrum of data
processing is required); and it embraces the full spectrum of data      sources, types, models, and modalities, including text, images,
sources, types, models, and modalities (including text, images,         audio, and signals, while supporting a variety of integration
audio, and signals). The result of applying this strategy to a given    and analytic processes and tools. Details are presented below.
body of data is an evolving Dataspace that allows the application
of a variety of integration and analytic processes to diverse data      The Dataspace store of intelligence data which is the subject
contents. We conceive semantic enhancement (SE) as a light-             of this communication is the result of applying the DRIF to the
weight and flexible process that leverages the richness of the          task of integrating very large heterogeneous primary data
structured contents of the Dataspace without adding storage and
                                                                        artifacts. As the Dataspace has evolved through time, so it has
processing burdens to what, in the intelligence domain, will be an
already storage- and processing-heavy starting point. SE works          incorporated progressively ever larger quantities of data, and
not by changing the data to which it is applied, but rather by          also more specific local implementations and data structures
adding an extra semantic layer to this data. We sketch how the          used by data analysts, some of which bring their own data
semantic enhancement approach can be applied consistently and           semantics. For the purposes that the Dataspace is intended to
in cumulative fashion to new data and data-models that enter the        serve, it is vital that no restrictions are imposed either on the
Dataspace.                                                              types of source-artifacts and the associated models and media
                                                                        within the Dataspace, or on the processes by which the
Keywords: integration, intelligence data, ontology, semantic            Dataspace is populated (whether by loading structured data
technology.
                                                                        from a database, by extraction from a text document through
                                                                        some Natural Language Processing application, by automatic
                        I.   INTRODUCTION                               analysis of signals, or through inference by a human analyst).
The success of the war fighter and homeland defender in the
Net-Centric Warfare environment is largely defined by the               The design of the Dataspace is such that it can incorporate
ability to quickly acquire and efficiently and accurately process       hundreds of millions of unstructured documents and similarly
intelligence information from numerous heterogeneous sources            large quantities of images, signals data, and other structured
of different structure and modality. Traditional data integration       and unstructured primary data artifacts. Each of these artifacts,
approaches fail in the face of the scale, diversity, and                when it enters the Dataspace, is represented through a set of
heterogeneity of intelligence data sources and data-models              metadata, including labels specifying image type, MIME type,
because they fail to address one or more of the following               and so forth, as well as provenance information. Further
requirements:                                                           processing may, for example, associate pixels in an image
• Integration must proceed without heavy pre-processing                 with the name of a person, or a range of characters in an
• Integration must proceed regardless of the data-models                unstructured text document with the name of a location, or
     used (or not used) in the data sources to be integrated,           extract a cell from a database table. The DRIF provides a
• Integration must proceed regardless of the data modality,             common framework in which the results of all of these
     and without loss or distortion of data, of its associated          processes are represented in a unified way, details of which
     data semantics, and of data-provenance information,                are provided below. As a result, primary data can be utilized
• Integration must involve the ability to incorporate                   immediately upon entering the Dataspace for a variety of
     multiple points of view on the data to be integrated,              different kinds of search and more sophisticated processing
     including different views of the data, for example on the          based thereon. DRIF is not, however, a magic bullet; many
     part of different analysts using different analytical tools.       issues of data integration at the syntactic level will remain,
arising for example as a result of data formats which do not         forms, which we call primary and derived, respectively. The
match, where we will need to normalize the format into an            Dataspace is divided into corresponding segments (see Figure
augmented model that will serve as the target of annotations.        1) in a way that supports a comprehensive approach to
This will involve considerable effort to ensure that the needed      integration that allows accommodation of the multiple views
actions are performed promptly and consistently whenever             of the primary and derived data and of the associated data-
new data comes in. Here, however, we focus exclusively on            semantics and metadata which arise for example as a result of
those issues which arise at the stage of what we can loosely         the workings of multiple different sorts of analytical tools.
call the ‘representational’ aspects of data integration.
                                                                     A. Approach to Integration
Some primary artifacts within the Dataspace already                  Our approach to integrating intelligence data starts with source
incorporate useable semantic content – for instance a                artifacts consisting of primary data across a variety of
structured database which incorporates meaningful column             representation modalities. This primary data is weakly
headers, or a message with a structured payload incorporating        integrated in the sense that indexes are provided to support
meaningful tags. But such content is ad hoc. It is tied to           simple (string-based) data search across all primary artifacts.
specific local implementations and typically falls short of what
is needed to secure semantic interoperability of the                 Some primary data comes with its own native structure, and
implementations involved because of the absence of a                 further structure will typically be added thorough analytical
common formally coherent approach to semantics and of a              processing. The second integration step addresses the need for
common governance process.                                           the unified storage of this structured data to support more
                                                                     complex structured search across both primary and derived
Moreover, full semantic integration is in any case prevented         artifacts.
by the needs of openness of the Dataspace to ever new sorts of
primary data and analytically derived data. It is to compensate      Importantly, we here embrace the diversity of domain-specific
for this problem that we have developed our strategy for             data-models employed throughout the Intelligence Community
semantic enhancement. We start out from the assumption that          while at the same time reaping benefits from an approach that
semantic data enrichment can be achieved only incrementally,         is data-model agnostic. This is because the unified
through the step-by-step creation of ontology modules that are       representation provided by the DRIF allows analytic
designed in coordinated fashion to work well both with each          processing of data in highly diverse primary artifacts
other and with specific bodies of Dataspace content. The             associated with different native data-models to be used as
vision is a lightweight, flexible approach comprising an extra       targets of cross-artifact analytics. For example, and most
ontology layer that leverages the contents of the Dataspace          simply, it is possible to perform unrestricted string search
without adding storage and processing weight to what is an           across structured artifacts of highly different sorts. Examples
already storage- and processing-heavy resource. We discuss           of more sophisticated analytics include computer-aided data-
the details of semantic enhancement in section IV. First,            model harmonization, for example by allowing significant
however, we introduce the DRIF and the Dataspace to which            overlap of sets of values of attributes from different databases
the SE strategy will be applied.                                     to be flagged by the analytic process as a potential indication
                                                                     that the attributes have the same meaning, thereby making it
        II.   DATA REPRESENTATION AND INTEGRATION                    possible for the relevant portions of the two databases to be
                       FRAMEWORK                                     enriched through fusion.
Our starting point is a body of U.S. Department of Defense           B. Dataspace Organization
(DoD) intelligence data within what we are here calling the
                                                                     The organization of the Dataspace is schematically illustrated
Dataspace. The implementation in the specific context upon
                                                                     in Fig. 1.
which we focus here is engineered around cloud computing
paradigms and is primarily based upon open-source cloud
                                                                     Segment 0 is a store of primary artifacts, including documents,
software stack components. This cloud computing foundation
                                                                     images, signals, and analysts’ work products vetted for re-use
leverages advantages of linear scaling and parallel distributed
                                                                     as input for further processing. The physical implementation
computation when faced with the reality of ever increasing
                                                                     of Segment 0 may be such that all data is stored internally; or
data volumes and integration processing. All the work
                                                                     it may be distributed, so that source artifact data may for
described is either deployed or in the final stages of testing
                                                                     example be either contained in the cloud store or stored
prior to deployment.
                                                                     externally to the Dataspace and referenced in the cloud store.
                                                                     Primary data vary widely by nature; they may have different
The Dataspace is built using the Data Representation and
                                                                     structures (for example of a relational database), or they may
Integration Framework (DRIF), which has been designed to
                                                                     be unstructured (for example, free text, audio or video files),
represent large quantities of data in a form that is useful to the
                                                                     and they may be of different modalities (for example they may
end user both for direct inspection and for the application of
                                                                     be cells of a relational database, audio sequences, assertions of
various kinds of analytics. Representations of source data
                                                                     an analyst).
artifacts and their contents within the Dataspace are of two
                                                                                 •   Segment 3 describes the data-models which support the
      Segment 3 – Structured descriptions of models associated with                  two sets of views just mentioned as well as synoptic
      primary and derived artifacts (e.g. attributes of a relational table           views (ultimately including SE-based views) of the type
                with associated functional dependencies)                             which can foster harmonization.
                                                                                 D. Where Models and Primary Data Come Together
         Segment 1 – Structured                Segment 2 – Structured
         descriptions of primary             descriptions of data (results       We believe that the principal contribution of the Dataspace
                  artifacts                   of processing of primary           endeavor is to resolve certain problems of storage and thus of
          (e.g. registration data;             data, e.g. through NLP,           representation, enrichment, and evolution of large bodies of
       relations to other artifacts)           image analysis, or other          data. The goal is to provide room for both primary data and
                                                   data extraction)
                                                                                 the multiple results of processing these data by different
                                                                                 analysts or analytic methods. To achieve this we introduced in
        Segment 0 – Primary artifacts (stored in CloudBase, HDFS                 [3] a strategy for description of data that is designed to enable
          and elsewhere) (including documents, images, signals,                  true data integration across a constantly evolving and highly
        databases; as well as analysts’ products, some with built in             heterogeneous resource comprehending extremely large
                               data models)                                      volumes of data. As already recognized at the very beginning
                                                                                 of contemporary high-level research in biomedical ontology
                                                                                 [4], this end can be achieved only if data are exposed in a way
 Figure 1. Organization of the Dataspace. Solid line: registration processes;    that is independent of their original intended use. This must
 curved solid lines: processes that ingest artifacts into the Dataspace,         involve some means to represent original data-models at a
 including feeding back into the Dataspace analysts’ products – results of the
 Dataspace processing ; dashed lines: derivation processes.
                                                                                 level of abstraction that is higher than that of primary data. We
                                                                                 accordingly propose an abstract data-model based on five core
Segment 1 includes primary artifact registration data as well as                 elements: sign, concept, term, predicate, and statement, which
specifications of relations between artifacts (for example,                      we believe is sufficient to represent any data-model in these
nesting of an image within a document, or attachment of one                      terms.
document to another). Segment 1 will include also data
pertaining to the way each derived artifact of Segments 2 and                    Sign: A sign gi is a string that is the abstracted proxy within
3 is derived from primary artifact(s) in Segment 0.                              the dataspace for one or more chunks of data used in some
                                                                                 primary artifact with the intention of referring to some
Segment 2 stores the structured data that is either already                      individual entity (e.g. person, location, organization, object,
present in primary artifacts or derived therefrom through                        event). Examples include: a sign of the type proper name that
analytic processing resting on data-models represented in                        is associated with an expression (for example ‘he’ or ‘Dr.
Segment 3.                                                                       Watkins’ occurring in a document; a label annotating an area
                                                                                 in a pixel array as forming an image of some building; a label
Segment 3 stores the descriptions of the data-models used in                     annotating a fragment of an audio stream or other signal as
Segment 2. These data-models may include database schemas,                       recording some explosion event. Each sign is associated with
message formats, or XML schemas. The data-models                                 one or more physical extents within those primary artifacts
themselves are primary artifacts and are thus stored in                          with which it is associated, which we call mentions (the latter
Segment 0 and registered in Segment 1.                                           are what are elsewhere called tokens). The collection G = {gi}
                                                                                 comprehends all signs extracted from primary data artifacts
The Dataspace is evolving continuously not only because of                       and changes with the incorporation of new artifacts.
new primary data ingested from the outside, but also because
new artifacts are being created, for example, through analysts’                  Concept: A concept ci is (for the purposes of this exposition) a
reports based on processing of existing data. These artifacts                    string that is used in the Dataspace to represent some general
themselves have a status of new primary artifacts.                               category or grouping. The purpose of concept strings is to
                                                                                 represent and allow reuse of classifications native to primary
C. Segments as Abstractions Over the Artifacts                                   artifacts. Concepts are taken from data-models registered in
Each of Segments 1-3 is an abstraction over the corpus of                        Segment 1. Examples of concepts are: the classes of an
primary data artifacts (Segment 0) and supports analytics of a                   ontology such as UCore SL, the tag set in an XML Schema
particular type:                                                                 Document (XSD), and the attribute or table names in a
• Segment 1 is a high-level view of the entire artifact                          relational database. The collection C = {ci} comprehends all
     corpus including the relations between the artifacts, but                   concepts within the Dataspace and changes as new data-
     with no reference to their internal contents.                               models are incorporated.
• Segment 2 is a collection of detailed views of the internal
     contents of the artifacts at the level of individual data                   Term: A term, tij, is an ordered pair of strings , where gi
     items.                                                                      ∈ G and cj ∈ C. Each term results from a process of contextual
                                                                                 disambiguation of a sign, a process which associates a sign
                                                                                 with a concept, as in <123-45-6789, SSN>. The collection T =
{tij} comprehends all terms identified by analytic processing of     current phase of evolution of DRIF, the phase of Semantic
primary artifacts.                                                   Enhancement (SE). SE, as we conceive it, is a light-weight
                                                                     and flexible solution that leverages the richness of the native
Predicate: A predicate (by which we mean here always:                source data and of any local semantics associated with these
binary relational predicate) pi is a string that is used to          data without adding storage and processing weight. The SE
connect terms in accordance with domain and range                    strategy is compliant with and complements the DRIF.
constraints. Predicates are used in the formation of statements
                                                                                                        Artifacts
(as described below). Examples of predicates are: hasSSN,
hasLocation, hasBirthDate. Predicates are derived from data-                Database A                       Database B
models registered in Segment 1, for example from table                         ID      PersonName          PersonID       Name        Address
                                                                                       …                                  …
column headings or from XML tags. The collection P = {pi}                   732        Bill                821            William     DC
comprehends all predicates within the Dataspace and changes
as new data-models are added.                                                   Document X

                                                                             ….Scott performed the database backup…
Statement: A statement si is an ordered triple consisting of a
subject, a predicate, and an object. The collection S = {si} of
statements is recursively defined. At the lowest level,
                                                                                                  Structured content
statements are ordered triples consisting of a term, a predicate,
                                                                          Sign                    Concept            Predicate
and a second term. In higher-level statements, subjects and
objects may be lower-level statements. Examples: <[Bruno,                 key           label     key       label         key           label
                                                                         1           732          1       ID          1             hasName
PersonName] hasSSN [123-45-6789, SSN]>                                   2           Bill         2       Scientist   2             hasAddress
                                                                         3           821          3       PersonID    3             sameAs
The five primitives of the DRIF (sign, concept, predicate,               4           William      4       Name        4             knows
term, and statement) define a data reference model which, by             5           DC           5       Address
effectively decoupling data from data-models, can represent              6           Scott        6       DBA
any sort of data-model at the level that is useful for                   Term
integration.                                                                      key            sign_Key    concept_Key
                                                                        1 [732, ID]             1            1
Fig. 2 schematically illustrates the representation of structured       2 [Bill, Scientist]     2            2
data in accordance with the DRIF for three sample primary               3 [821, PersonID]       3            3
                                                                        4 [William, Name]       4            4
artifacts, two of them relational databases, the third an               5 [DC, Address]         5            5
unstructured document. The example also shows how data-                 6 [Scott, DBA]          6            6
semantics come to be added to the Dataspace in ad hoc
                                                                          Statement
fashion – here, because an analyst decides to to introduce a
new Concept DBA (meaning: database administrator).                       key        term_Key_Subject      predicate_Key      term_Key_Object
                                                                         1          3 [821, PersonID]     1 hasName         4 [William, Name]
Additional Statements establishing relationships between                 2          3 [821, PersonID]     2 hasAddress      5 [DC, Address]
Terms using Predicates SameAs and Knows are also included                3          3 [821, PersonID]     3 sameAs          1 [732, ID]]
in the Figure.                                                           4          3 [821, PersonID]     4 knows           4 [Scott, DBA]

The reader familiar with the Resource Description Framework
(RDF/RDFS) may wonder what is different here. RDF                        Figure 2. Simplified example of structured content derived from 3
                                                                                                 primary artifacts.
employs a similar level of abstraction, but it is a language,
while what we are offering here is a specific, albeit still highly
abstract, data-model. This data-model could of course be             A. Goals of Semantic Enhancement
specified very easily using the RDF language; but it could be        SE is a strategy that is currently being implemented to
specified also using relational database or some other storage       improve our handling of the enormous heterogeneity of
technology. Our choice of data-model was motivated further           Dataspace content. It is centered on building a flexible and
by the fact that our implementation and security requirements        extensible framework of hierarchically organized, controlled
dictated the use of a specific type of cloud storage solution [5,    structured vocabularies – called ‘ontologies’ – covering
6] that is both highly scalable and offers highly granular           different areas of relevance to intelligence analysis. The
security access controls.                                            framework will be constructed in part by reusing already
                                                                     existing resources, in part through collaboration with other
                III.   SEMANTIC ENHANCEMENT                          defense and military organizations in the creation of new
                                                                     ontology modules. The ontologies will be used in an
The DRIF focuses on the representational aspects of the              incremental process of annotation (or ‘tagging’) of those
Dataspace and on the basic types of data integration that such       concepts and predicates already identified in data-models
representation provides. In what follows we describe the             within the Dataspace along the lines described in our
discussion of Segment 3 above. The latter amount to what we         subtype (or is_a) hierarchy; and second through the
referred to above as ‘ad hoc semantics’. Because the salient        progressive incorporation in all nodes of the SE ontologies of
data-models derive from so many heterogeneous sources, they         links to relevant synonyms derived through the annotations
use a multiplicity of partially overlapping and partially           which will link ontology nodes to the rich collection of
conflicting vocabularies, which it is the task of SE to reconcile   corresponding concepts and predicates in other areas of the
by associating co-referring concepts and predicates (strings)       Dataspace.
employed within distinct data-models in the Dataspace to
                                                                    C. The Strategy for Semantic Enhancement
single nodes within the external SE ontologies.
                                                                    Our strategy is designed to achieve its goals not by changing
To function in the needed way, annotations must be                  the Dataspace, but rather by adding an extra semantic layer
cumulative, in the sense that our strategy will ensure that tags    thereto. The strategy is thus similar to that underlying the
created by different annotators will be consistent with each        Universal Core (UCore), which arose out of the National
other. The value of annotations must also be preserved when         Information Sharing Strategy supported by multiple U.S.
the SE ontologies change, for example through refinements           Federal Government Departments, by the intelligence
created to reflect advances in knowledge, and to this end the       community, and by a number of other national and
ontologies must be subject to strict versioning policies.           international organizations [7, 8]. Here, a small controlled
                                                                    vocabulary was provided for multi-community use to associate
Finally, the SE framework must be implemented in such a way         simple summary tags to message payloads for purposes of data
that it can serve not merely as a tool of harmonization of the      search and integration.
data-models internal to the Dataspace but also in a way that
allows integration with other, external data resources wherever     Reflecting the extreme diversity of intelligence data, multiple
common ontologies are used for annotation.                          subject-matter expert communities will be contributing to the
                                                                    SE. For the strategy to work and provide useful and efficient
To address these constraints is by no means a simple matter.        integration, these multiple distributed teams must use the SE
When data value codifications do not match – for example            approach in a consistent fashion. Previous efforts to create a
when we have 1,2,3 in one data source, R, G, B in another data      broad-based, multi-community ontological approach to data
source, and RED, GREEN, BLUE in our Color ontology, then            integration in defense and intelligence domains have failed
annotation for each source to hierarchy values can be very          because the incompatible, and often over-simplistic, views of
labor intensive and require significant SME effort.                 reality incorporated into legacy databases and data-models led
                                                                    to incompatible development of ontologies in ways that
B. Sample Benefits of Semantic Enhancement                          precluded interoperability. Many advocates of semantic
We can see the sorts of benefits that SE will provide already at    approaches to data integration have still failed to appreciate
the level of search, where problems arise because of the            the tremendous challenges, both technical and human, created
multiple different ways of describing data within the               by the entrenched predisposition on the part of ontology
Dataspace. Problems that need to be confronted include:             developers to create ontologies each on the basis of their own
                                                                    potentially idiosyncratic data representations.
1. The need to find data items identified by means of terms
which are narrower or broader in meaning than the terms             The solution which we advocate is modeled on the successful
analysts will standardly use when searching;                        semantic annotation approach pioneered in the field of
2. The need to find data items in documents that are                bioinformatics by the Gene Ontology [9]. This approach is
formulated using a language or technical jargon with which          now being pursued systematically within the framework of the
analysts are unfamiliar.                                            OBO Foundry [10, 11], which starts out from the idea that the
                                                                    most effective way to ensure mutual consistency of ontologies
To provide some very simple examples: we know that a given          created by multiple independent groups over time and to
package ‘has been shipped with a red label’, but the                ensure that these ontologies are maintained in such a way as to
documents that we have pertaining to this package use only          keep pace with advances in knowledge is to organize
the word ‘vermillion’; or we need to find references to a           ontologies as a collection of modules with discrete (non-
package identified as ‘containing furniture’, but the documents     overlapping) subject-matters maintained by subject-matter
we have refer only to ‘chairs’; or we need to find a given          experts, according to a strategy outlined in [12]. To ensure
package suspected of containing crack cocaine, but the audio        consistency, these ontologies should be created as extensions
recordings we have at our disposal relating to this package         of more generic higher level ontologies, subject to common
refer only to ‘bobo’ or ‘botray’ or ‘boubou’. If we are             rules for example concerning the treatment of definitions, and
restricted to string search, our queries would not return the       they should be based on a small common upper-level ontology
needed results. Hence, we need a framework which expands            (ULO), whose domain and content neutral. For example, it
string search by capturing type and subtype information, and        will include relations such as is-a (for subtype), member-of,
also incorporates synonym information. These needs are              part-of, and so on. As initial ULO we choose the Basic Formal
targeted along two dimensions; first, through the fact that all     Ontology (BFO) [13], which has been implemented in more
SE ontologies will be organized around a central backbone
than 100 similar projects, and which serves as the basis of the    D. Implementation of the SE Strategy
already mentioned UCore Semantic Layer [8].                        We can now outline the steps which are involved in realizing
                                                                   this strategy in the specific context of the Dataspace, where
The ULO will be associated with a small number of Mid-             we already have data structured using the DRIF.
Level Ontologies (MLOs) defined by downward population
from the ULO. The MLOs will serve in turn as bridge to a           First Step: Review the contents of the Dataspace, specifically
number of Low-Level Ontologies (LLO), which will specify           that concepts and predicates in Segment 1, and identify a
narrow content domains. Each MLO represents cross-domain           subset of topic areas where data integration is a priority for
entities, such as Person or Information, and will be constructed
                                                                   analytics.
in tandem with the LLOs which it subsumes in order to ensure
the mutual consistency and interoperability of the subsumed
                                                                   Second Step: Formulate a list of MLOs that would be needed
LLOs. The MLOs and LLOs must in turn be associated with
the resources of a relation ontology, providing for the            to annotate the data in corresponding areas. As far as possible
representation of content-specific relations such as Owns,         identify existing ontologies which may potentially be reused
WorksFor, Audits, and so on.                                       for this purpose, and build initial versions of new ontologies
                                                                   where needed.
Initial due diligence efforts in our strategy of semantic
enhancement requires us to identify an initial collection of       Third Step: Identify a specific subset of the content of the
authoritative codifications at Mid- and Lower Levels – along       source data-models, and identify LLOs that will capture this
roughly the lines depicted in Table 1 – and to begin the           subset in a semantically coherent fashion, ensuring that each
process of formalizing them within the BFO common upper-           LLO is subsumed by some MLO. Subject matter experts
level ontological framework. In some areas ontologies will         should be recruited to take charge of creation and maintenance
need to be created de novo, since no adequate authoritative        of the LLOs and MLOs and of their use in annotations. In this
codifications will exist.                                          way we can create a cadre of SMEs with expertise in
                                                                   annotation and in supporting semantic enhancement.
                Examples of MLO cross-domains
                                                                   In realizing the above we need to maximize as far as possible
         •     Geospatial
                                                                   the reuse of ontologies which are already being used by
         •     Biometrics
         •     Person                                              relevant communities. This is because the strategy will be
         •     Provenance and Trust                                successful only to the degree that a critical mass of potential
         •     Organization                                        users are able to be convinced of its utility and thus
         •     Signals and Sensors                                 incentivized to engage in advancing it further for example by
         •     Equipment                                           extending it new types of data and by disseminating the
         •     Facility                                            resource to new groups of analysts. Reusing already existing
                 Examples of LLO domains                           ontologies will not merely provide a core of familiar terms
                                                                   which analysts can use for search purposes, it will also
         Subsumed by Geospatial
         • Geospatial Feature                                      increase the degree to which we can integrate into the
         • Country                                                 Dataspace data that has already been annotated in consistent
         Subsumed by Biometrics                                    fashion by external bodies.
         • Fingerprint
         • Iris                                                    Fourth Step: When once a stable, initial set of ontologies has
         Subsumed by Person                                        been created, we use these ontologies to annotate the data-
         • Employment Data                                         models in corresponding portions of the Dataspace. As should
         • Criminal Data                                           by now be clear, the entire strategy is an incremental one,
         • Medical Data                                            based on a principle of low hanging fruit: the idea is not to
         • Ethnicity and Tribe                                     import the above ontologiesas a whole; rather we examine
         • Skill
                                                                   the existing Dataspace resources and identify expressions
         Subsumed by Provenance and Trust                          therein for which counterparts in the ontologies already exist
         • Data Quality
         • Access Permissions                                      or can easily be added. In constructing the ontologies these
         • Data Source                                             expressions will be provided with a common logical
         • Evidence                                                architecture and a common set of relations defined through
                                                                   the ULO top level and in terms of which logical definitions
             Table 1: Sample Ontologies within the SE              for terms in the ontologies can then be formulated. The result
                            Structure                              can be used as a basis for the application of general-purpose
                                                                   tools, including standard OWL reasoners FaCT++, RACER,
                                                                   or Pellet, which can be used to check ontologies in the SE
                                                                   resource for mutual consistency.
                                                                    particular subset of a source data-model is too general for data
Stage 4 of the SE process consists in associating each set of       analyst purposes, then the respective LLOs can be further
equivalent data source concepts with a single common MLO            developed as needed.
or LLO expression (which will be added at the appropriate
level within the SE ontology structure where not already
                                                                                       Person
present). Further types of integration are thereafter brought
about automatically. Whenever any Dataspace resource                       Ethnicity                            Skill
becomes linked to one of our chosen ontologies in a way that
                                                                                                                              Computer
can be used to generate corresponding annotations, it thereby                                                                 Skill
becomes linked to all the other Dataspace ( a n d
e x t e r n a l ) resources that have already been annotated with                                               Programming     Network
the same SE ontologies. This creates a snowball effect,                                                         Skill           Skill
whereby each new annotation increases the value of existing                               Works        Audits
annotations [9], and provides further incentives for the use of
the SE ontologies by new groups of users.                                                 Facility

                                                                                                     Educational                   Used for
E. Organization of the SE Ontologies                                       Hospital
                                                                                                                                   annotations
                                                                                                     Facility
Fig. 3 illustrates the organization of the SE ontology space.                                                                      of source
                                                                                                                                   data-
Each LLO represents the reality of a particular narrowly
                                                                                                                                   models
defined domain, for example in an area such as Education and
Skills.

                                                                                        Middle                           Is-a
An MLO is a container of LLOs. Since we will be developing
LLOs in step-by-step fashion to address what are at any given                           Lower                            attribute-of
time the most urgent needs of Dataspace users, there will be                                                             semantic
data which cannot as yet be annotated with the full granularity
of detail which the annotator requires. The strategy is to use             Figure 3. Simplified Example of an SE Ontology Structure.
such cases to advance the further development of the ontology
resource base, again following the model tested in the
bioinformatics domain [9]. For example, an analyst may want
                                                                                                 IV.     CONCLUSION
to use the SE resources to extract and disambiguate data from
a particular document. For different reasons the analyst may        Together, the DRIF and SE provide what we believe is a
not be able to use the most detailed semantics and will use a       workable data-integration solution. The DRIF is a highly
more general one. LLO taxonomies will also be used by               flexible framework, with few constraints and including an
analytics to produce results of different level of detail: from a   RDF-style decomposed representation of structured data
fine-grained view of narrow areas within the Dataspace to           which allows the collection of data resources without loss or
coarse grained pictures of larger domains.                          distortion in a way that achieves syntactic integration and
                                                                    preserves the local semantics of primary sources and of
Because original data and data-semantics are in every case          analytics software. SE provides semantic integration in a light-
preserved without loss or distortion in the Dataspace as it         weight yet incrementally extendible fashion, and in a way that
exists prior to Semantic Enhancement, there is no need to           can foster global integration without adding storage and
represent all details of original storage data structures in the    processing weight to already storage- and processing-heavy
SE stage. This means that complex ontologies are not needed         Dataspace.
– a common and shared vocabulary is sufficient for virtual
semantic integration and search/analytics, while underlying         The SE approach provides a strategy to allow the Dataspace to
details are maintained by the authors of specific primary           be understood as evolving cumulatively as it accommodates
artifacts. Similarly, the collection of SE ontologies does not      new kinds of data. It provides a more consistent,
need to cover all of the ad hoc local semantics within the          homogeneous, and well-articulated representation of
Dataspace – content that is unlikely to be used in search or is     structured content that originates in multiple internally
not important for integration can be excluded from the              inconsistent and heterogeneous models. And while it involves
Enhancement step, since it will still be available in the source    considerable initial SME investment in ontology creation and
data-models and can be accessed when drilling down to the           annotation, we believe that it will allow the management and
appropriate level.                                                  exploitation of the Dataspace to become more cost-effective
                                                                    over time.
The SE approach is highly flexible. It represents a “pay-as-
you-go” approach in the sense that investments can be made          In addition, the use of the selected MLOs and LLOs brings
only in specific areas according to identified need. It is also     integration with other government initiatives and brings the
tunable in the sense that, if a given body of annotations for a     Dataspace endeavor closer to the federally mandated net-
centric data strategy; it also makes the integrated Dataspace               [4]  Rosse, C. and Mejino, J. L. V. A Reference Ontology for
                                                                                 Bioinformatics: The Foundational Model of Anatomy. Journal of
more effectively searchable and provides an expanding body                       Biomedical Informatics 36, 2003, 478-500.
of content to which more powerful analytics can be applied in               [5] R6 Cloudbase documentation and source code.
the future.                                                                 [6] Hadoop. http://hadoop.apache.org.
                                                                            [7] http://ncor.us/ucore-sl.
Acknowledgments: This work was funded by US Army                            [8] B. Smith, L. Vizenor and J. Schoening, “Universal Core Semantic
CERDEC I2WD. The authors thank Mr. Kesny Parent,                                 Layer”, Ontology for the Intelligence Community, Proceedings of the
DCGS-A Branch Chief, for continued support.                                      Third OIC Conference, George Mason University, Fairfax, VA, October
                                                                                 2009, CEUR Workshop Proceedings, vol. 555.
                                                                            [9] D. Hill, et al., “Gene Ontology Annotations: What they mean and where
                                                                                 they come from”, BMC Bioinformatics, 2008; 9(Suppl 5): S2.
                             REFERENCES                                     [10] B. Smith, et al., “The OBO Foundry: Coordinated Evolution of
                                                                                 Ontologies to Support Biomedical Data Integration”, Nature
                                                                                 Biotechnology, 25 (11), November 2007, 1251-1255.
[1]   S. Yoakum-Stover, T. Malyuta, N. Antunes, “A Data Integration         [11] B. Smith, W. Ceusters “Ontological Realism: A methodology for
      Framework with Full Spectrum Fusion Capabilities”, Presented at the        coordinated evolution of scientific ontologis”, Applied Ontology 5
      Sensor and Information Fusion Symposium, Las Vegas, NV, Aug 3-7,           (2010) 139-188, http://x.co/adRJ.
      2009.
                                                                            [12] W. Ceusters, B. Smith, J. M. Fielding, “LinkSuiteTM: formally robust
[2]   A. Hansen, D. Salmen, T. Malyuta, and N. Antunes. “An Evolving             ontology-based data and information integration,” in Database
      Integrated Dataspace on the Cloud.” Presented at the Sensor and            Integration     in    Life    Sciences,    Berlin,  Springer,  2004.
      Information Fusion Symposium, Las Vegas, NV, July 26-29, 2010.             http://ontology.buffalo.edu/bio/LinkSuite.pdf
[3]   S. Yoakum-Stover, T. Malyuta, “Unified Integration Architecture for   [13] Basic Formal Ontology. http://www.ifomis.org/bfo/.
      Intelligence Data”, Proceedings of DAMA International Europe
      Conference, London, UK, 2008.