<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HADatAc: A Framework for Scienti c Data Integration using Ontologies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Paulo Pinheiro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henrique Santos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhicheng Liang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yue Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sabbir M. Rashid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deborah L. McGuinness</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcello P. Bax</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Rensselaer Polytechnic Institute</institution>
          ,
          <addr-line>Troy, NY, 12180</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidade Federal de Minas Gerais</institution>
          ,
          <addr-line>Belo Horizonte, MG, 31270-901</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidade de Fortaleza</institution>
          ,
          <addr-line>Fortaleza, CE, 60811-905</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>To investigate the cause and progression of a phenomenon, such as chronic disease, it is essential to collect a wide variety of data that together explains the complex interplay of di erent factors, e.g., genetic, lifestyle, environmental and social. Sharing information between studies is therefore of paramount importance. However, data that needs to be analyzed must be appropriately integrated, conceptually aligned, and harmonized. This implies that data collection must be done either in a su ciently similar or a su ciently transparent way in order to support meaningful synthesis from di erent studies. We will demonstrate4,5 how the Human-Aware Data Acquisition (HADatAc) framework integrates and harmonizes data from multiple scienti c studies and thus how to use it in interdisciplinary science investigations.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The Human-Aware Data Acquisition (HADatAc) Framework is a schema-free,
evolutionary, scalable and provenance-aware infrastructure for managing data
and metadata content from multiple scienti c studies. Three key goals of
HADatAc are: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) to extract relevant data value from instrument-generated les
and to move these values into queryable content repositories, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) to extract
relevant metadata from scientist-generated documents and to move these values
into queryable content repositories, and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) to semantically annotate these
values in a way that the entire content is logically linked and harmonized (i.e.,
uni ed representation) according to evolving collections of well-established
scienti c ontologies. HADatAc's core ontologies, that are fully integrated, aligned,
and used in multiple scienti c domains, include: W3C's Provenance Ontology
(PROV), encoding provenance knowledge, Virtual Solar-Terrestrial Observatory
(VSTO) [2], encoding knowledge about instruments and platforms,
HumanAware Science Ontology (HAScO)6 [4], encoding knowledge about studies, study
types, data elicitation from humans, and data simulation from computer
models, and the Semantic science Integrated Ontology (SIO) [1], encoding knowledge
about science-related entities and their characteristics.
4HADatAc's live demo is available at http://bit.ly/HADatAc
5Demo video: http://bit.ly/HADatAc-iswc2018
6http://hadatac.org/ont/hasco
      </p>
      <p>Without any expansion, the framework and the collection of ontologies listed
above are domain agnostic and ready for usage in scienti c domains. Existing
HADatAc deployments are built using these core ontologies, along with many
other specialized ontologies, for the domain of interest. One speci c domain
ontology is often used to import specialized ontologies into a single document
that HADatAc uses as a default namespace for a given domain of interest (e.g.
the integrated exposure and health ontology CHEAR [3] that we use in our
HADatAc backend for the NIEHS Child Health Exposure Analysis Resource
implementation).
2</p>
    </sec>
    <sec id="sec-2">
      <title>HADatAc Characteristics</title>
      <p>Schema-free claim. HADatAc uses semantic technologies, ontologies, graph
databases and non-relational databases to manage metadata and data from
relevant scienti c studies, e.g. CDC's National Health and Nutrition Survey
(NHANES)7 in the demonstration system. HADatAc is schema-free since the
content from these study les is stored without a prede ned and xed structure.
This allows HADatAc to include objects, e.g., subjects, samples, locations, into
its repositories as they are presented, including attributes from objects regarded
as relevant for the studies.</p>
      <p>Evolutionary claim. Ontologies and the underlying graph, including the
loaded data, are managed by one Apache SOLR repository and by one
Blazegraph RDF graph database. The SOLR repository manages data values and a
collection of URIs for each data value. URIs, in the collection of URIs of each
stored data value, are links between the data value and semantic annotations
in the Blazegraph. The Blazegraph repository is used to manage the HADatAc
knowledge graph. HADatAc's knowledge graph evolves by either importing
existing ontologies or by de ning concepts and relations that are not available or
not appropriate for reuse from existing ontologies.</p>
      <p>Scalability claim. Scienti c data management platforms must handle
increasing data volumes. HADatAc's backend SOLR repository provides the
required scalability for very large data repositories. As the metadata volume is
signi cantly smaller than the data volume, it is easily managed by the
Blazegraph triple-store. SOLR is also used to compute aggregate faceted values into
the scienti c study indicators, that would be costly to do with Blazegraph.</p>
      <p>Provenance-Aware claim. HADatAc has been developed to work with
a broad range of data sources, and the infrastructure captures and preserves
provenance on how data values where acquired by a broad notion of instruments.
HADatAc classi es data sources according to the instruments (and detectors)
used to acquire the data. HADatAc understands if study data is the result of
any of the data acquisition strategies: (a) empirical measurement, which is done
using physical instruments like sensors; (b) data and knowledge elicitation from
humans, which is done using questionnaires as instruments; (c) computer data
generation (simulation), which is done using simulation models as instruments.
7https://www.cdc.gov/nchs/nhanes/index.htm</p>
    </sec>
    <sec id="sec-3">
      <title>HADatAc Architecture</title>
      <p>Metadata
file</p>
      <p>File
Management</p>
      <p>Data files</p>
      <p>Study
Management</p>
      <p>Instrument</p>
      <p>Management
SOLR
(data)</p>
      <p>Blazegraph
(metadata)</p>
      <p>HAScO
POJO</p>
      <p>Library
Metadata Data
Ingestion Ingestion</p>
      <p>Content Ingestion
Core Component</p>
      <p>Search
Download
Alignment</p>
      <p>API</p>
      <p>HADatAc is a framework and also a web application. Most of its web user
interface is implemented as part of the six satellite subsystems. The API Subsystem
is a special subsystem composed of a collection of RESTful services with
programmatic access to HADatAc's content. The Core Component has the elements
required to support the satellite subsystems including: the SOLR and Blazegraph
content repositories, a Java API encoding the concepts of the Human-Aware
Science Ontology (HAScO) as POJO Classes, and the subsystems responsible for
extracting, annotating, and storing study content from data and metadata les
into SOLR and Blazegraph. The HAScO POJO classes are used to build and
maintain HADatAc's knowledge graph.</p>
      <p>Content is added into HADatAc either through the parsing of uploaded
les through the File Management Subsystem or on-line through user
interaction with the Study Management Subsystem and the Instrument Management
Subsystem. Content is directly presented to users through the Search Subsystem
and downloaded through the Object Alignment Subsystem and API Subsystem.
4</p>
    </sec>
    <sec id="sec-4">
      <title>HADatAc's Demonstration</title>
      <p>
        Our demonstration includes the following six steps: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) registration of CDC's
NHANES as a new study in HADatAc including the generation of NHANES
subjects as RDF instances, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) registration of a semantic data dictionary (SDD)
for NHANES that identi es how the content of NHANES data les are extracted,
integrated, and harmonized, (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) upload and processing of NHANES data les
that store NHANES data and metadata into databases, (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) automatic processing
of uploaded data les, (5) display of harmonized content in a semantic faceted
search, as shown in Figure 2, and nally (6) download of datasets generated
from the current selection of a search in the data faceted search tool. We have
used a combination of techniques to identify and retrieve variables to be used in
the demo, mainly the NHANES Variable Search8 and, to generate datasets for
ingestion into HADatAc, the RNHANES R package9, which allows extraction of
select named variables across NHANES datasets. Our claim is that HADatAc can
play the roles of those multiple techniques for variable identi cation, retrieval,
and dataset generation for diverse domains.
Acknowledgements This work was partially funded by the National Institute
of Environmental Health Sciences (NIEHS) Award 0255-0236-4609 /
1U2CES026555-01 and CAPES Foundation Award 88881.120772 / 2016-01. It has been
co-deployed with Mount Sinai School of Medicine and the Gates Foundation
Healthy Birth, Growth, and Development Knowledge Integration program with
collaborators from Yale's Center for Ecosystems in Architecture.
8https://wwwn.cdc.gov/nchs/nhanes/search/default.aspx
9https://CRAN.R-project.org/package=RNHANES
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Dumontier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baran</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callahan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chepelev</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz-Toledo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Del Rio</surname>
            ,
            <given-names>N.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duck</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furlong</surname>
            ,
            <given-names>L.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keath</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , et al.:
          <article-title>The Semanticscience Integrated Ontology (SIO) for Biomedical Research and Knowledge Discovery</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <volume>14</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cinquini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>West</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benedict</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Middleton</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Ontology-supported Scienti c Data Frameworks: The Virtual Solarterrestrial Observatory Experience</article-title>
          .
          <source>Computers &amp; Geosciences</source>
          <volume>35</volume>
          (
          <issue>4</issue>
          ),
          <volume>724</volume>
          {
          <fpage>738</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>McCusker</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rashid</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chastain</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinheiro</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stingone</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          :
          <article-title>Broad, Interdisciplinary Science in Tela: An Exposure and Child Health Ontology</article-title>
          .
          <source>In: Proceedings of the 2017 ACM on Web Science Conference</source>
          . pp.
          <volume>349</volume>
          {
          <fpage>357</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Pinheiro</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bax</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rashid</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCusker</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          :
          <article-title>Annotating Diverse Scienti c Data with HAScO</article-title>
          .
          <source>In: Proceedings of the Seminar on Ontology Research in Brazil 2018 (ONTOBRAS</source>
          <year>2018</year>
          ).
          <source>Sa~o Paulo</source>
          ,
          <string-name>
            <given-names>SP</given-names>
            ,
            <surname>Brazil</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>