<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Translation of Various Bioinformatics Source Formats for High-Level Querying</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>John McCloud</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Subhasish Mazumdar</string-name>
          <email>mazumdarg@cs.nmt.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>New Mexico Institute of Mining and Technology Socorro</institution>
          ,
          <addr-line>New Mexico</addr-line>
          ,
          <country country="US">U.S.A</country>
        </aff>
      </contrib-group>
      <fpage>39</fpage>
      <lpage>53</lpage>
      <abstract>
        <p>Representations of genomic data is currently captured for use in the eld of bioinformatics. Each of these representations have their own format, caveats, and even speci c \sub-languages" used to represent the information. This makes it necessary to use di erent specialized programs to make sense of | and then query | the data. Unfortunately, this means a great deal of lost time necessary for learning both the tool and the format itself in hopes of extracting details that can then later be used in answering queries. There is, therefore, a great need for a general purpose tool that can address all of these di erent formats at once: a tool that will abstract such low-level, format-speci c representations of raw data to a higher, operating level with terms that are immediately recognizable to experts within the bioinformatics and biology domains. In addition, it will help biologists engaged in research pose questions that were not anticipated by the authors of the special purpose software built for the data les.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>they have in any format, but there is no tool that allows them to look at anything they wish in
the data. Instead, they are forced to use whatever de nitions and terms those specialized pieces of
software de ne.</p>
      <p>We have developed a solution to this issue by manually mapping terms in the ontology for
these various, low-level data sources. Alongside an e ective, simple target at an operational level
for querying, contextual explanation for the ontology-level classi cations are presented, giving a
more cohesive understanding for adding new information to an ever-growing map of relationships
in data.</p>
      <p>The approach does not have only one, completely unambiguous ontology from which all
knowledge representation stems. Instead, it allows for multiple domain-level ontologies to coexist at once
in a shared, operational level (from which queries are made).</p>
      <p>This paper is structured as follows. In the next two sections, we present a review of related work
and a short primer on biology. Subsequently, in section 4, we delve into design issues. A case study
dealing with les using the SAM format is presented in section 5, followed by respective sections of
short analysis and concluding remarks.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Previous Work</title>
      <p>Much of the work that has been done in linking data sources usually starts out as a way to combine
low-level ontologies. While such research is extremely valuable in the work presented here, we remind
the reader that we are not merging well-de ned ontologies, but instead linking the raw or semi-raw
data that a researcher in the eld may collect | and linking data formats | with larger ontologies.
Our research attempts to implement ideas from the previous work below.
2.1</p>
      <p>C-OWL
Contextual OWL took the original language of OWL and made a system that assumed a single,
global ontology for all term interpretation would not provide suitable representation for particular
local cases. By this extension, the research realized that context within a local ontology must be
maintained. In order to represent ideas from knowledge at the local and higher levels, contextual
mappings are created with relationships to the various levels of abstractions (from local to higher
levels) linked by bridges [1].</p>
      <p>These bridges are similar to what is represented here in this research, but they are much more
abstract, often times explicitly mentioning the levels of abstraction (where a given term within the
ontology resides), and the relationship to another term in the bridge. (This allows for such mappings
as T ermA is identical to T ermB from low-level ontology A and high-level ontology B.) The bridges
and their mappings allow for denoting contextual levels of speci city of terms, but do not describe
how to actually derive terms from low-level sources.</p>
      <p>This is because both source-level and target-level information they deal with are assumed to
be OWL ontologies themselves. In our research, we assume all data from the low-level source is
not within any such formalization to begin with, but must instead be built from RAW data and
linked to a higher-level ontology. Much as in OWL, though, the source-level de nitions maintain a
useful context, and terms within that context must be dealt with di erently than terms with other
contexts (from other sources).
2.2
Perhaps the most inspiring work for this research comes from the KRAFT project, which recognized
the necessity of a shared ontology amongst many low-level ontologies and manual linking from
the low-level de nition up to the single, shared one [2]. Their agent-model design also includes a
distributed solution, whereby one user agent can receive knowledge from another user agent (and
his de nitions) across a network.</p>
      <p>By creating a new language that acts as a link between the shared ontology and the low-level
sources, this Constraint Interchange Format (CIF) made it possible for user agents to make queries
against data not de ned in the shared ontology. This turned the query into a Constraint Satisfaction
Problem. A user could then ask a question, and these constraints would be developed for the various
low-level sources, using their particular, individual languages. This kept a user from needing to know
the format and de nitions of anything other than the shared ontology.</p>
      <p>The issue with the KRAFT system is that it allows mappings to be arbitrary and possibly
completely wrong. This, however, is an acceptable drawback of the system, since we must trust
those knowledge experts who deal with their data regularly to de ne it and expose caveats, fringe
cases, and easy-to-miss constraints within the de ned resources.</p>
      <p>Additionally, it is not clear if the mappings allow for any links larger than one step. That
is, multiple transformations from low-level heterogeneous source and the shared ontology. Certain
multiple steps may be necessary to produce an e ective, overall mapping. Furthermore, combining
these intermediate de nitions may be an important task that should be available to provide richer
explanation from low-level details.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Primer on Biology and Sequencing</title>
      <p>For readers unfamiliar with the concepts and terms we will use shortly, this section provides a
quick primer on the central dogma of biology. This begins with DNA, an extremely long sequence
of nucleotide bases, which makes up almost every living thing (genome is another term used to
refer to the total genetic data of an organism, and is more general); and all living organisms are
a collection of cells containing a certain number of recognizable chromosomes, wound up in tight
bundles around proteins; each chromosome contains one or two DNA molecule(s).</p>
      <p>DNA is a code that contains all the genetic information necessary for every type of cell in our
bodies; from cells that make up our eyes to cells that form our skin. Each cell is specially tuned to
perform a certain way, depending on which portions of the DNA are manifested in it and how the
organism should respond to the environment it lives in.</p>
      <p>Portions of the DNA have regions called genes, and when these genes are expressed, it can result
in the creation of tiny, biological actors within the cell, called proteins (and if expressing the region
does result in proteins, we say that it is a coding region), which perform all types of functions. Each
cell usually only contains a certain subset of proteins, and this is dependent on the type of cell
in which the proteins are built in. When an organism exhibits a certain type of protein, we say it
expresses the gene that codes for the protein; this is referred to as gene-expression.</p>
      <p>Usually, cells produce a reasonably normal amount of proteins in each cell (as per the cell's
type). However, when the pressure from the environment results in some shift in the need of the
organism, a cell's DNA can adapt to create many of a particular type of protein. This is usually
to protect the organism or help it cope with whatever environmental stresses it needs to respond
to. Such shifts in gene expression are interesting to examine in biology, because it can assist in
discovering new regions of the genome that were not previously known. Often, such discoveries are
made by comparing an observed DNA coding region against a reference genome; an agreed upon
reference for an organism's total genetic material.</p>
      <p>For the purposes of this paper, it also helps to understand sequencing and sequence alignment.
Sequencing is the process of discovering the actual order of the nucleotide bases that constitute the
organism's DNA. Once sequenced, the various sections of the DNA are matched against a reference
to create a standard alignment of the data.</p>
      <p>As a result of sequencing, one may nd a series of reads (portions/fragments of genetic data)
that are very similar, appear in many locations of the genome, or sometimes just shifted slightly
by location. When the locations and lengths are overlaid one on top of one another, the ones
with the greatest overlap { these overlapping pileups | leads to stronger con dence in accurate
representations of that fragment of the genetic code; they are referred to as consensus sequences.
These consensus sequences are important, because one might conclude that a gene exists there,
with gene expression strength depending on how many reads were overlapping along this same area
[8].
4</p>
    </sec>
    <sec id="sec-4">
      <title>Design</title>
      <p>In building the ontology design, these core goals have been identi ed:
{ An expert in the domain should require only the knowledge of her domain and not the
additional de nitions in the le formats for querying and query resolution. We must separate the
knowledge of the low-level, speci c data formats from the domain expert's \standard"
knowledge of the domain.
{ We must create intuitive methods for domain experts to implement additional mappings as new
data formats present themselves. A tool that can deal with multiple data formats should be
easily extensible.
{ Queries at the topmost level must be independent of data formats. Seamless execution on
di erent les (and therefore, di erent formats), will allow the data to be cross-referenced, shared,
and pooled together.
4.1</p>
      <p>The Ontology's Structure
In describing the architecture of this ontology design, it helps to think of a staircase. At the top
of the staircase are the top-level de nitions, while at the bottom are the source- le de nitions.
The source- le de nitions are transformed at each ascending step, becoming a little closer to the
shared-level de nitions at the top. De nitions at the highest, top step of the staircase are what
experts in their domain might use to describe a question.</p>
      <p>The de nitions at the lowest, source-level step are what each particular le format de nes. These
are terms that are designed by individuals crafting le formats. Such de nitions can be arbitrary,
and do not generally share the language of the high-level. Instead, they de ne terms that are terse,
complex, or otherwise lacking in su cient identi cation as to be used by a general expert of the
domain. For instance, DNA Polymerase, in such a low-level, could be de ned as DNA P1 or DP1,
but this is not how biologists refer to it; this is how someone who built the le format has chosen
to represent it.</p>
      <p>The goal is to allow queries using the concepts of the high-level to pull information from the
lower levels, even though their de nitions may not immediately align. Their de nitions are mapped
from the source-level to the shared-level (and vice-versa), transforming a de nition in one format,
to intermediate de nitions, and then nally to the agreed upon de nitions of terms used by an
entire eld or domain.</p>
      <p>Take this staircase analogy and imagine that the steps are now layers. There are several layers,
and at each layer L a given de nition D exists. In order to go from one layer to another, some
transformation de ned by a function is performed on that de nition. Sets of such functions are
contained within bridges B.</p>
      <p>Users, such as biologists, may ask queries such as: Which alignment is repeated the most
in the HFD1.sam data, and does it represent a new gene? This query aims at nding
possibly unknown genes that may have been exposed by environmental stress in the experimental,
sequenced organism.</p>
      <p>Next, let us break down the query into the high-level terms (and their attributes) that need to
exist at the shared-level.</p>
      <p>{ Chromosome (Number )
{ Gene (Name, Start Location, End Location, Expression)
{ Sequence (Start Location, End Location, Length, Bases)
{ Alignment (Unique, Total Matches, Mapped, Mismatching Positions )</p>
      <p>Since these terms (and attributes) are used in queries, a domain expert must build bridges for
them from the source-level up to the shared-level (Fig. 1).</p>
      <p>Start Location</p>
      <p>Sequence
End Location Length Bases</p>
      <p>Unique</p>
      <p>Alignment
Total Matches Mapped</p>
      <p>Mismatching Positions
Chromosome</p>
      <p>Number</p>
      <sec id="sec-4-1">
        <title>Shared Layer</title>
        <p>Name End Location
Gene</p>
        <p>Start Location</p>
        <p>Tags
IH MD YR NH RG</p>
        <p>Intermediate Layer 1
Attributes
ID Name Alias Parent Target Gap
Expression</p>
        <p>Bridge
Bridge
RNAME</p>
        <p>FLAG</p>
        <p>POS
SEQ</p>
        <p>TAG
SAM Format</p>
      </sec>
      <sec id="sec-4-2">
        <title>Source Layer</title>
        <p>Start</p>
        <p>End</p>
        <p>Attributes
GFF3 Format
Fig. 1: Depiction of layers and bridges within the architecture. Each layer contains its own particular set of
de nitions, while bridges contain the sets of functions for transforming de nitions at a given level to
adjacent ones. The example here has two source formats: SAM and GFF3.</p>
        <p>Our ontology design requires at least two layers and one bridge: the shared LS and low-level
(\ground") LG layers, and the single bridge BLSLG containing two sets of functions: one set of
functions for the transformation of data de ned in the low-level layer to the shared layer fij and
another for the inverse functions (mapping shared-layer de nitions back down to the low-level)
fij 1. An arbitrary number of additional, intermediate layers and their associated bridges, however,
is possible (and necessary in some cases).</p>
        <p>
          Generalized representations for these concepts are given here:
{ A function f that lives within a bridge, transforming some de nition i at a lower layer into
de nition j at a higher layer.
{ The inverse of the transformation in Equation (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ), showing the transformation of some de nition
j into de nition i (where j exists at a higher layer and i exists at a lower layer).
{ Each pair of layers only needs one bridge, and both functions in Equations (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) are
included in the same bridge, but in two di erent sets (one set for transformations from the
higher-to-lower layer with the other set dealing with lower-to-higher layer transformations).
These sets are denoted in (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) and (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) (with generic layers A and B).
        </p>
        <p>fij : Di ! Dj :
fji : Dj ! Di :</p>
        <sec id="sec-4-2-1">
          <title>BLALB</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>BLBLA :</title>
          <p>BLALB = ffij g; ffjig :</p>
          <p>These de nitions, however, do not elucidate the full de nition of bridges and the functions within;
bridges can be non-surjective and non-injective. From the source-level, a mapping to the domain
level may exist, and if it exists, the inverse function(s) de ned by the bridge can be performed to
return to the source-level representation.</p>
          <p>That is, given fnm, a function to transform de nition n to m it's inverse would be fnm1 if it
exists.</p>
          <p>Such transformative functions can also be combinations of other, already de ned functions. This
means that a transformation from one layer to another can be dependent on terms de ned at the
same level as well as those below it. This allows one to create intermediate de nitions and functions
at various levels that can be called on at any time to be used in adding further transformation.
4.2</p>
          <p>Foundational De nitions
Upper-level ontologies de ne abstract de nitions (which we refer to as primitives ) for the elements
within domain-speci c ontologies, establish relationships amongst these abstract de nitions, and so
on. Perhaps it is easiest to think of the upper ontology | indeed, the entire ontology | from an
object-oriented point-of-view, where the top level is as generic as possible, and with the de nitions
becoming more concrete as one descends the hierarchy.</p>
          <p>
            These primitives have been identi ed in our design:
{ Component: The pieces an entity is composed of; what makes up an entity. Components
\belong to" an entity.
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )
(
            <xref ref-type="bibr" rid="ref2">2</xref>
            )
(
            <xref ref-type="bibr" rid="ref3">3</xref>
            )
(
            <xref ref-type="bibr" rid="ref4">4</xref>
            )
{ Constraint: Conditions that allow/do not allow a certain function or process to occur.
{ Entity: Any physical building block or standalone actor that performs actions or functions.
          </p>
          <p>Entities \have one-to-many" functions. Entities \have one-to-many" components.
{ Function: Actions which are performed by entities. Acts \with an" entity. May be \part of" a
process. Can \depend on one-to-many" constraints.
{ Process: Possibly non-linear sequence of interacting functions. Processes \has one-to-many"
functions. Processes may \depend on one-to-many" constraints.
{ Property: Descriptive parts of a primitive type (any of component, entity, etc). It is possible
for everything to \have one-to-many" properties.
{ Rule: Axioms for determination of validity in interaction or outcome thereof. It is possible for
everything to \have one-to-many" rules that govern them.</p>
          <p>Using such terms, we can take concepts from a given domain, set its interactions, rules, and so
on in order to relate de nitions within the ontology to one another.</p>
          <p>Upper ontologies contain basic, atomic de nitions for concepts in the world. The shared
ontology (here, the collections of domain-speci c ontologies) contains various implementations of these
atomic, abstract concepts. For example, an abstract de nition may be an entity. A particular
implementation in the nancial domain may be a bond, stock, or future. This concrete de nition may
di er depending on the domain-speci c ontology in which it is used, since a \future" in physics is
not the same concept as in nance.</p>
          <p>If \future" in both physics and nance is de ned as an entity, it means that both de nitions
have the same primitive type that possesses rules and context for interaction with one another. The
terms have both been given meaning from that base de nition, but can also extend their abilities
by de ning something more speci c to their domain-related purpose.</p>
          <p>Such relationships similar to \has a", \is a", and others are de ned for understanding connections
amongst de nitions at this level and other levels. For example, consider that an entity is an object
that is distinctly di erent from a function; that functions are performed by entities. Therefore, an
entity would \have a" function.</p>
          <p>For example, RNA Polymerase is a term that biologists know; it is an entity, with rules and
functions associated with it. A process might be DNA Replication, with a series of functions in a
particular order, and rules for how they must interact. All would be de ned at the shared-layer,
using the actual biological terms.</p>
          <p>To clarify, take the example of RNA Polymerase, while again considering the structure in Fig. 1.
At the shared-level, we develop a de nition for this entity and de ne its various relationships as
given in Fig. 2. Even from such a small example, one is able to see how these various primitive
de nitions for the model are linked together.
4.3</p>
          <p>The Shared-Level Ontologies
The shared-level ontologies are a series of domain-speci c views. At this level, multiple domain
ontologies coexist as a union of all the domains that are relevant (and would be loaded into the
system). Within each speci c domain, basic terms are de ned, and the shared-level deals with
overlaps, such as identical terms and possible de nitions, that may arise as a result of this union.</p>
          <p>In our design, the shared-level ontology is also the operational level, meaning that queries are
formed from here, potentially involving every domain in the union. This is to allow for the same
terms across domains to be used in a single query. However, this also means that a system of
ENTITY - RNA Polymerase
FUNCTIONS: Reads DNA, Reads RNA, Produces RNA
COMPONENTS: Amino Acids
ASSOC. PROCESSES: Transcription, RNA Replication</p>
          <p>DESCRIPTION: General RNA Polymerase Enyzme
Is Composed of
Has Function
ENTITY - Amino Acid
COMPONENTS: RNA
PROPERTIES: Specific RNA 3-Tuple
ASSOC. PROCESSES: Translation
DESCRIPTION: The building blocks of proteins</p>
          <p>FUNCTION - Reads DNA
INPUT: DNA
OUTPUT: None
ACTS W/ENTITIES: DNA Polymerase, RNA Polymerase
OCCURS WITHIN PROCESSES: Transcription, DNA Replication
DEPENDS ON CONSTRAINTS: Promoter
DESCRIPTION: Reading of DNA nucleotide strands</p>
          <p>Involved in Process
PROCESS - Transcription
INPUT: DNA
OUTPUT: RNA
ASSOC. RULES: RNA Compliments, DNA Compliments</p>
          <p>DESCRIPTION: Conversion of DNA to RNA
RULE - RNA Compliments
G Pairs with C
A Pairs with U</p>
          <p>Has Rule</p>
          <p>Has Rule</p>
          <p>RULE - DNA Compliments
G Pairs with C
A Pairs with T
Fig. 2: Example representation of RNA Polymerase Entity and its relationship to other types of example
primitives.
metadata tagging may be necessary for terms to be di erentiated when there are repeated terms
across domains [3]. This allows a user to make queries using terms that would overlap on the same
domain.
4.4</p>
          <p>The Low-Level Source Materials
The low-level sources are the raw sources of data that one may capture from the eld.</p>
          <p>Here, low-level sources represent the nal word in implementation speci cs for given de nitions.
One format may de ne a term or even hide a necessary de nition within a larger, format-speci c
de nition, while another format de nes it completely di erently. However, an individual performing
queries should not | indeed, does not | need to concern oneself with such implementation speci cs.
This is the familiar concept of encapsulation.</p>
          <p>Consider multiple source-level mappings to the shared-level have been made. An individual
with great familiarity of a given format has taken the raw data, with its speci c representations
of knowledge, and created a hierarchy of de nitions and mappings for us. Once this mapping is
built, the most precise de nition exists at the source-level, but we can still call on the speci c
implementation from the top of the hierarchy, hiding the source-speci cs transformation details.</p>
          <p>Consequently, a user interfaces only with the shared-level queries using shared-level terms such
as Packets and Timestamp instead of source-level terms like RPACK77 and timestamps as FIELD3.
Even when the correspondences between terms are simple, they are hidden from the user and it is
the system that translates shared-level terms into their format-speci c implementations at the low
level data source.
4.5</p>
          <p>Bridging and Linking
Data transformation from the shared (or operational ) level of the ontology is taken care of through
a series of manually mapped bridges. These bridges exist as points between levels or layers of the
overall ontology; each contains two sets of functions: transforming de nitions from a lower layer to
a higher layer, and vice-versa.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Case Study: Mappings with SAM Files</title>
      <p>To demonstrate the usefulness of this ontology design, we consider the example of building some
mappings with a low-level source format. Many di erent tools have been developed in the eld of
bioinformatics for organizing and representing the data, all with their own speci c formats; there
may even be versions of formats in response to the evolving needs of scientists. One such format is
the Sequence Alignment Mapping (SAM) [4], [5].</p>
      <p>The SAM format was originally built to be generic enough to respond to various di erent
alignment formats that were developed. In so doing, it became very exible, capable of holding a
great deal of information in such a small size. However, in order to accommodate translation from
various other formats, additional information had to be developed (by way of tagging and other
pieces of the format).</p>
      <p>The following outlines a short case of mapping between the SAM le low-level format, and the
shared-level ontology for querying.
5.1</p>
      <p>Data
Research is currently being conducted at the New Mexico Institute of Mining and Technology that
involves the examination of fat transport within rattus norvegicus (common, brown rat) on two
di erent and distinct diets. The raw data has been sequenced, and now exists in a SAM format [6].
(For the rest of this example, we will be using data from the rats fed the high-fat diet.)
5.2</p>
      <p>The SAM Format
The SAM format is similar to that of a CSV le. That is, there are a series of rows (called alignment
lines) with di erent, delimited elds. The two exceptions to this rule are found in optional header
section (which we do not fully detail in our study here) and a variable number of elements for the
\TAG" eld de ned for each alignment line. (Note that columns in SAM les are not given a named
mapping as is done in CSV. The eld/column names are implicitly derived by the SAM format.)</p>
      <p>The header section (which is optional) de nes information that corresponds to the entirety of
reads in the le. Such general information as the SAM format version number, the speci cs of the
alignment lines at various locations, the order in which data will appear, and so on are de ned
within this section.</p>
      <p>Immediately following the header section (should it exist at all) are the rows of alignment lines
(Fig. 3). Each of these lines are broken up into eleven distinct elds | all with their own caveats
| and an additional, variable-length twelfth eld (which contains various tags).
SNPSTER3_0511:3:21:542:508#0/1 0 gi|62750345|ref|NC_005100.2|NC_005100 289632 0 36M *
0 0 CTCCTCCTCCCCTTCTTTATGGGTATTCCCCTCCCC HHHHHHHHHHHGGHHHHHHHHHHHHHHHHHHHHGGG MD:Z:36
RG:Z:s_3_sequence.txt.gz IH:i:2 NH:i:2 YR:i:1
SNPSTER3_0511:3:37:1137:468#0/1 0 gi|62750345|ref|NC_005100.2|NC_005100 305783 0 36M *
0 0 CTCTCATGGCTTAAAGTCCTGCCCGGGACCACCCTG HHHHHHHHHHHHHGHHGHHHHHHEHGHHHHHHIHH@ MD:Z:36
RG:Z:s_3_sequence.txt.gz IH:i:4 NH:i:4 YR:i:1
SNPSTER3_0511:3:23:1266:1799#0/1 0 gi|62750345|ref|NC_005100.2|NC_005100 329635 0 36M *
0 0 ATTNAATATAGTTCTCGAACTCCTAGCCAGAGCAAT GGG&amp;GHHHFHFGGHHHHHHIHHIHHFIIHHGIGHIG MD:Z:3X32
RG:Z:s_3_sequence.txt.gz IH:i:2 NH:i:2 YR:i:1
SNPSTER3_0511:3:108:1205:762#0/1 0 gi|62750345|ref|NC_005100.2|NC_005100 368226 0 20M2D16M
* 0 0 GTTTTGGGGATTTAGCTCAGGTAGAGCGCTTGCCTA HGHHHHHHHHHHHHHHHHHHHGHIHHHHHIFHHHHH MD:Z:38
RG:Z:s_3_sequence.txt.gz IH:i:5 NH:i:5 YR:i:1
SNPSTER3_0511:3:31:1390:225#0/1 16 gi|62750345|ref|NC_005100.2|NC_005100 389280 0 36M *
0 0 GGGGTCAACCTTCAGTCAGCCCCTCGTAGCGAGGGA EGEGEGGEEGGGGGGDEEAFGGCEAGGGGGGGAGGG MD:Z:36
RG:Z:s_3_sequence.txt.gz IH:i:1 NH:i:1
SNPSTER3_0511:3:103:43:1576#0/1 0 gi|62750345|ref|NC_005100.2|NC_005100 394139 0 36M *
0 0 TACGAATGTGATCAATGTGGTAAAGCCTTTGCATCT HHHHHHHHHHHHHHHHHHHHGHHHHGHHHHHHHHHH MD:Z:36
RG:Z:s_3_sequence.txt.gz IH:i:4 NH:i:4 YR:i:1</p>
      <p>le. Fields are delimitted by whitespace. The nal TAG
eld itself has several sub elds of tags.
5.3</p>
      <p>The Alignment Section and Lines
The bulk of the SAM</p>
      <p>le is taken up by the alignment section, lled with various alignment lines.</p>
      <p>Each alignment line represents a particular portion of sequence that was read when sequencing was
done on the observational data. These particular sequences are referred to as reads. In an alignment
line, these reads correspond to some type of match that has been aligned to a reference genome.</p>
      <p>There are at least eleven distinct elds, delimited by whitespace, in each alignment line. The
twelfth eld, TAG, is optional and slightly di erent, since there can be multiple tags within that eld
or even none at all (whose multiple tags are also delimited by whitespace). An example alignment
line's elds (Fig. 4) are detailed below.</p>
      <p>{ FLAG: 0
{ POS: 1350933
{ MAPQ: 0
{ CIGAR: 36M
{ RNEXT: *
{ PNEXT: 0
{ TLEN 0
{ QNAME: SNPSTER3 0511:3:85:824:196#0/1
{ RNAME: gij62750345jrefjNC 005100.2jNC 005100
{ SEQ: CACACATACGAGGTGTACTTTGTGTGCAGAATGTGT
{ QUAL: E?9BBCC8?C=?;C844555FFF?F:&gt;==6&gt;CCFCE
{ TAG: MD:Z:36 RG:Z:s 3 sequence.txt.gz IH:i:3 NH:i:3 YR:i:1</p>
      <p>These identi ers correspond to</p>
      <p>elds (1-indexed, i.e., starting with index 1) 1 through 12
respectively. These are the low-level source representations of data in the SAM format.
5.4</p>
      <p>Mapping from the Low-Level Source
This is where one de nes the mappings for those shared, operational level terms such as
chromosome and others. We rst examine the alignment lines in the SAM le (many lines are given in
Fig. 3, while a single sample line to be used in examples is given in Fig. 4) to better understand
their structure.</p>
      <p>While building the bridges, it is important to keep in mind that we are trying to match a
common, shared-level de nition. Let us try this with Chromosome.</p>
      <p>At rst glance, it may be di cult to identify chromosome information within the alignment
lines. Fortunately, a domain expert is aware of such caveats and can identify that the chromosomal
numbering is given within the RNAME eld (3rd element in Fig. 4).</p>
      <p>SNPSTER3_0511:3:85:824:196#0/1 0 gi|62750345|ref|NC_005100.2|NC_005100 1350933 0
36M * 0 0 CACACATACGAGGTGTACTTTGTGTGCAGAATGTGT E?9BBCC8?C=?;C844555FFF?F:&gt;==6&gt;CCFCE
MD:Z:36 RG:Z:s_3_sequence.txt.gz IH:i:3 NH:i:3 YR:i:1</p>
      <p>The last 2 digits within RNAME (00 in Fig. 4) denote the chromosome number (0-indexed;
meaning 00 is chromosome 1) for this alignment line. Slicing o the last two digits of the RNAME
eld is then de ned as the function for this level's bridge for Chromosome.Number.</p>
      <p>
        This transformation provides a mapping from the low-level source to the shared-level.
fRNAME Chromosome:Number 2 BLGLS : RN AM E ! Chromosome:Number :
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
      </p>
      <p>If one were to code the function by-hand, one might de ne it as (in pseudo-code):
FUNC RNAME-TO-CHROMNUM(RNAME): return RNAME[-2:]</p>
      <p>This is one of the functions that the bridge would contain. Recall that bridges hold the functions
necessary for transformations. As a high-level de nition is transformed as part of a query (or a
lowlevel de nition is transformed), the bridges know which function to call for which de nition. As a
query descends the layers, each de nition is transformed accordingly via these functions.</p>
      <p>This satis es a complete mapping for the chromosome number in the SAM le format, version
1.0. This mapping goes into a collection for the SAM format with this speci c version, which is kept
within a le and pointed to by the system's framework. This helps to maintain a proper organization
for the bridges and the formats they are linked to.</p>
      <p>We examine alignment uniqueness next, since this is a particularly interesting \free-form" case
of the SAM elds. In order to nd uniqueness information, we have to consult the twelfth eld,
TAG. The information from the TAGs given in our example alignment line are:
MD:Z:36</p>
      <p>RG:Z:s_3_sequence.txt.gz</p>
      <p>IH:i:3</p>
      <p>NH:i:3</p>
      <p>YR:i:1</p>
      <p>Each tag is given a name, a type, and a value. Anyone familiar with the SAM format can look
at these tags and tell right away whether or not an alignment line is unique. Such information is
located within tag \IH."</p>
      <p>The \IH" ag denotes the number of alignments in the entire SAM le that have the same
template name. The type for our \IH" ag is \i" meaning that the value will be an (i)nteger. In
this case, it has value 3. If this template (QNAME) were unique, the value for IH would be 1.</p>
      <p>To map this value, the low-level de nition DG is set as whatever the le de nes. However, it
cannot simply su ce to set the de nition as TAG, since the TAG eld de nes many tags. Instead,
one must set a more appropriate de nition on a per-tag basis.</p>
      <p>To do this, we create an intermediate level up from the low-level source. (This is to illustrate
that an arbitrary amount of intermediate levels can be built to suit the needs of the low-level source,
combining de nitions from various levels if necessary.) The intermediate level splits up the TAG
eld into its component parts (the di erent tags).
{ Tags: IH, MD, RG, NH, YR</p>
      <p>From the low-level source to this intermediate layer, the functions de ned in the bridge are to
separate the Tag eld up by cutting it into pieces, and setting those pieces (the individual tags),
into a new set of de nitions:</p>
      <p>This transformation provides a mapping from the low-level source de nition to the intermediate
de nition of \Tags."
fT AG T ags 2 BLGL1 :</p>
      <p>
        T AG ! T ags:IH; T ags:M D; T ags:RG; T ags:N H; T ags:Y R :
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
The internals of this function might look like the following (in pseudo-code) :
FUNC TAG-EXPANSION(TAG):
local dictionary tag-map
for Tag in TAG:
      </p>
      <p>
        tag-map[Tag[:2]] = Tag[Tag.findLast(':'):]
return tag-map
To go the other direction, the inverse function su ces:
fT A1G T ags 2 BL1LG : T ags:IH; T ags:M D; T ags:RG; T ags:N H; T ags:Y R ! T AG :
(
        <xref ref-type="bibr" rid="ref7">7</xref>
        )
      </p>
      <p>Each tag in the example line (IH, MD, RG, NH, and YR) is mapped to the \Tags" de nition
in this intermediate layer. This provides a quick way for one to reference and use it in other
intermediate layers as needs be, having singled-out the information from the larger sample line itself.
Now the tags are split up, and from each of them, one can answer questions such as \uniqueness".</p>
      <p>
        Since the IH tag denotes uniqueness amongst all alignments in a SAM le, we can instead set
the IH tag from our intermediate-level \Tags" de nition to be \unique", mapping the value from
the intermediate layer to the shared layer. Only a boolean value of True or False is necessary. Also
note that this particular tag can be used to set the Total Matches for the alignment, but in that
case, the precise integer value becomes important (as opposed to being 1 or otherwise).
fT ags:IH Alignment:Unique 2 BL1L2 : T ags:IH ! Alignment:U nique :
(
        <xref ref-type="bibr" rid="ref8">8</xref>
        )
      </p>
      <p>The function for this transformation might be de ned (in pseudo-code) as:
FUNC IH-IS-UNIQUE(Tags.IH):
if Tags.IH == 1: return True
else: return False</p>
      <p>Note, however, that it may not be possible to reconstruct the IH tag merely from the \Unique"
de nition. Instead, the only way to get back the full, original value of this IH tag is from the \Total
Alignment Matches" portion of the Alignment De nition. That is, there may not exist a suitable
inverse function for uniqueness in the case where uniqueness is false, and instead, the function to
go back to the full IH tag must be:
fT a1gs:IH Alignment:Unique 2 BL2L1 : Alignment:T otalAlignmentM atches ! T ags:IH
(9)</p>
      <p>It should be clear that other mappings for the format | and indeed all di erent formats | can
be constructed similarly. Some may be more simple, while others are far more complex.
Now that the low-level source information has bridges set up to allow for de nition transformation,
we can make queries on the data. The overall query is broken up into sub-queries, answering
questions in parallel on the source les used with the aid of the bridges, transforming the SAM
format into something more recognizable by all domain experts familiar with bioinformatics, but
not necessarily the SAM format itself. Some examples are given below.1
{ Which sequence is repeated the most in the HFD1.sam data, and does it represent a new gene?
USES FORMAT SAM, GFF3
(SELECT MOST REPEATED Sequence FROM FILE HFD1.sam) AS RepeatedSeq
(SELECT Gene FROM FILE RATGENOMEv3_4.sam) AS KnownGenes
SELECT Gene.Expression
FROM FILE HFD1.sam
WHERE NOT KnownGenes
AND WHERE RepeatedSeq</p>
      <p>This will return all the possible pileups from the most repeated sequence in the HFD1.sam data
that su ciently overlap to denote a \new" gene. By forming sub-queries to count the pileup start
locations, examining their overlaps, and nding runs of su ciently piled up sequences, the system
can have certain de nitions for what constitutes | and then matches to | a possible gene region
such as this.</p>
      <p>{ Which non-coding regions of the genome, near the coding region for the leptin gene, are
expressive?
USES FORMAT SAM, GFF3
(SELECT Chromosome.Number, Sequence.StartBase.Number, Sequence.EndBase.Number
FROM FILE RATGENOMEv3_4.gff
1 Queries are presented as small snippets of SQL-like code, but we want to stress that biologists will not
have to learn SQL to make use of the system. Their high-level queries (in a language that will involve
biological concepts | still a work in progress) will be translated into similar code.
WHERE Gene.Name=lep) AS LeptinGene
SELECT Gene.Expression
FROM FILE HFD1.sam
WHERE Chromosome.Number=LeptinGene.Chromosome.Number
AND Sequence.Start NEAR LeptinGene.Base.Start</p>
      <p>Here, one is examining non-coding regions of rats that have been given a high-fat diet, exploring
any up-regulated portions of the genome around leptin, a protein known to be involved in fat
transport. This is to aid a researcher in a hunch regarding non-protein-coding regions of the genome
being relevant in dietary aberrations beyond the simple analysis of coding regions alone.</p>
      <p>Combining a sample with a particular reference is common in bioinformatics and is necessary for
asking certain types of questions (here, the reference genomic information comes from a le in the
GFF3 format [7]). With this model, a biologist can | in a single query | take two heterogeneous
source formats, combine them together, and make complex queries that result in meaningful answers
for questions. Questions that may prove useful in understanding diseases. Practically, any question
whose information exists within the source formats is answerable. And this is only an examination
of that particular eld. This model can be applied to any eld with source formats of any type and
number.</p>
      <p>Such an ability to make queries as this across many di erent, possibly overlapping source formats
at once is impossible with current tools available. There is no single language or system by which
such queries can be made with the richness of object-oriented design and the exibility of any source
format the mind can dream up, alongside coexisting, multiple domain-level ontologies, each with
their own de nitions.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Analysis</title>
      <p>The example of SAM and GFF les is a fairly complex one, and while this paper is not extensive
in its exhibition of how all transformations of that format are to be performed, they can be built
as prescribed by this proposed design methodology</p>
      <p>This proposed design still \su ers" from the same concern as existed in the KRAFT system,
mainly that arbitrary | even wrong | mappings can be de ned. This is a fundamental risk of
crowd-sourced solutions. Since we imagine bridge de nitions are just les that can be sent to others,
we believe that a type of self-checking will go on amongst the community of researchers that make
use of this design. More accurate, perhaps even faster, bridging de nitions for the di erent formats
will be accepted as the best, used most often, and signed o on, while those mappings that map
data incorrectly will wither and die.</p>
      <p>Of course, the question of how to ensure that bridge de nitions do not produce poor results
before using them is another task that we are currently looking into. The ability to verify that a
mapping produces good, accurate results is something that we would certainly like to have. This
would both cut down on conscious attempts to distort results as well as notify users about honest
mistakes in transformation of de nitions.</p>
      <p>Facilitating the creation of bridging will require a great deal of work, since not everyone knows
how to make a function to transform data into various de nitions for di erent link levels. It will,
however, be critical in producing something that experts want to use. For those that are comfortable
with building functions, the ability to just program them in should also be available.
Ontologies o er a way to organize and transfer knowledge, but they su er when overlap across
domains occur or formats di er. In such cases as arise in merging/linking ontologies, some form
of automation would be helpful in resolving issues of clarity and yield a single, well-formed and
unambiguous ontology. Unfortunately, there are many problems that arise in pursuing such an
automated approach.</p>
      <p>Instead, we have presented a solution that relies on using the work of experts who build an
ontology of their eld of expertise, e.g., biology, with that of experts who can build transformations
on data representations on a per-format basis. The result bene ts practitioners of the eld, who
have working knowledge of terms within their elds, but neither possess specialized knowledge of a
given data format, nor are technically trained to exploit it.</p>
      <p>By de ning shared-level terms for a given domain and building bridges from low-level sources
up to those domains, queries can be made on all variety of source formats. These bridges can then
be modi ed and shared by the community of researchers. Further, it is possible to query across
an ensemble of data in di erent formats allowing comparison, contrast, and union of the diverse
information content that no format-speci c software system would enable.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We would like to thank Dr. Rebecca Reiss for her assistance in helping us formulate queries and
realizing how those queries could be understood in a piece-wise manner from SAM les. We also
thank anonymous reviewers whose comments have helped improve this paper.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bouquet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Paolo</surname>
          </string-name>
          , et al.
          <string-name>
            <surname>C-OWL: Contextualizing Ontologies. The Semantic</surname>
            <given-names>Web-ISWC</given-names>
          </string-name>
          <year>2003</year>
          . Springer Berlin Heidelberg,
          <year>2003</year>
          . pp.
          <fpage>164</fpage>
          -
          <lpage>179</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Preece</surname>
          </string-name>
          et al.
          <article-title>The kraft architecture for knowledge fusion and transformation</article-title>
          .
          <source>In Proceedings of the 19th SGES International Conference on Knowledge-Based Systems and Applied Arti cial Intelligence (ES99)</source>
          . Springer,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Zhichen</surname>
          </string-name>
          , et al.
          <article-title>Towards the Semantic Web: Collaborative Tag Suggestions</article-title>
          . Collaborative Web Tagging Workshop at WWW2006, Edinburgh, Scotland.
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Heng</surname>
          </string-name>
          , et al.
          <source>The Sequence Alignment/Map Format and SAMtools. Bioinformatics</source>
          <volume>25</volume>
          .16 (
          <year>2009</year>
          ):
          <fpage>2078</fpage>
          -
          <lpage>2079</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Sequence Alignment/Map Format Speci cation.
          <source>SAMTools</source>
          . The SAM/BAM Format Speci cation Working Group, 29 May
          <year>2013</year>
          .
          <source>Web. 17 July</source>
          <year>2013</year>
          . http://samtools.sourceforge.net/SAMv1.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Fat-Rat SAM Data. NMT Biology Bioinformatics Portal</surname>
          </string-name>
          . New Mexico Institute of Mining and Technology.
          <source>Web. 03 Aug</source>
          .
          <year>2013</year>
          . http://bioinformatics.nmt.edu/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>The</given-names>
            <surname>Sequence Ontology - Resources - GFF3. The Sequence</surname>
          </string-name>
          <article-title>Ontology</article-title>
          .
          <source>Web. 18 Sept</source>
          .
          <year>2013</year>
          , http://www.sequenceontology.org/g 3.shtml
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mortazavi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ali</surname>
          </string-name>
          , et al.
          <article-title>Mapping and Quantifying Mammalian Transcriptomes by RNA-Seq</article-title>
          .
          <source>Nature methods 5</source>
          .7 (
          <year>2008</year>
          ):
          <fpage>621</fpage>
          -
          <lpage>628</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>