<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Hybrid Natural Language Approach to Manage Semantic Interoperability for Public Health Analytics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maxime Lavigne</string-name>
          <email>maxime.lavigne@mail.mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arash Shaban-Nejad</string-name>
          <email>arash.shaban-nejad@mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anya Okhmatovskaia</string-name>
          <email>anya.okhmatovskaia@mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luke Mondor</string-name>
          <email>luke.mondor@mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David L. Buckeridge</string-name>
          <email>david.buckeridge@mcgill.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>McGill Clinical &amp; Health Informatics, Department of Epidemiology and Biostatistics, McGill University</institution>
          ,
          <addr-line>Montreal, Quebec, H3A 1A3</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper discusses the integration of an ontology with a natural language query engine to calculate and interpret epidemiological indicators for population health assessment. In this paper, we discuss the application of this approach to one type of possible query, which retrieves health determinants, causally associated with diabetes mellitus.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology</kwd>
        <kwd>Natural language interface</kwd>
        <kwd>Causal inference</kwd>
        <kwd>Epidemiology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The Population Health Record (PopHR) platform [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] aims to improve
population health decision-making. It calculates and presents measures or indicators
of health determinants and health outcomes in a manner that, unlike most
current web portals, is intuitive to access and provides up-to-date indicators that
are contextualized by public health knowledge. In this paper, we describe our
approach to querying the PopHR knowledge base using an natural language
interface (NLI).
      </p>
      <p>Early in its development, it became apparent that even though we were
restraining the language of recognized queries, the breadth of pre- and
postconditions made implementation di cult. We therefore partitioned the space of
possible queries and called these, query types. By partitioning intents of user
inputs into collectively exhaustive and mutually exclusive query types, we were
able to overcome the di culty of designing a single data processing pathway for
all queries. Linking a query's concepts with our domain ontology is simpli ed and
it allows us, for example, to disambiguate concepts, which could have di erent
interpretation in di erent query types.Partitioning restricts the software contract
of our system when processing a query. Finally, this approach allows us to make
assumptions about the domains of concepts, such as statistics and geography,
which are relevant in our context.</p>
      <p>In this paper, we use a representative query as an example: What
determinants increase the risk of diabetes? The following sections introduce the relevant
parts of our domain ontology and then describe the strategy used to answer the
query.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Ontological Representation</title>
      <p>
        PopHR uses its own domain ontology representing knowledge relevant to
population health, including a taxonomy of human diseases, various groups of health
determinants and public health interventions, measures of disease occurrence
and other epidemiological concepts. In addition to the hierarchy of concepts,
the ontology encodes associative relations to allow for meaningful inference. One
speci c type of associative relation represents a causal link between two entities
(i.e., cause and e ect). For example, body mass index (BMI) has a positive
effect on an individual's disposition towards developing type 2 diabetes mellitus
(see Figure 1). More generally, this relationship is an example of a probabilistic
causal link from a health determinant to a process of developing a disease. We
can also describe a causal relation between a health determinant and a process
that modi es another health determinant.
In PopHR, the natural language interface is the preferred method for querying
information. All queries must respect a proper subset of the English language
that is formally de ned to be context free. The subset is built around question
answering and was conceived with the intent to provide all the expressivity
needed. This design decision implied that we needed an intuitive, consistent user
experience. For the system to succeed at providing proper guidance, it needed
to suit both the needs of the inexperienced users and experts.
We used our formally de ned grammar in conjunction with the ANTLR
framework[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for langage processing. The rst step of the process is to break the input
down into lexemes. The token stream produced by this step from the example
input is:
      </p>
      <p>Once the individual components of the question are separated, an LL(*)
parser uses production rules to generate a syntactic tree. If the creation of such
a tree is impossible, then we know that the input text was not part of our
language and proper guidance will be given on how to correct the issue. This
syntactic tree (Figure 2) is the artifact that will be used by the rest of the
system.</p>
      <p>A formal representation of the question is a necessary but not a su cient
step to understand the intent of the user. Although it is trivial for a human,
performing this step programmatically requires the ability to match the query
to some known patterns. This role is played by the oracle: all known patterns
are manually entered in the system and take the general form: What (To Be)
ID? is a description query. With this mapping, we are making the assumption
that a question that starts with the question word What and uses a derivate of
the verb To Be that has a nal concept ID is asking the system for a
description of this concept. Applying the Oracle to our example would classify it as
a CausalityEnumerationQueryWithConcept. We can intuitively concur that we
did want an enumeration of all determinants that have some causal relationship
to diabetes mellitus.
3.2</p>
      <p>Semantics
At this point in the process, we have gathered information regarding the domain
and general intent of the query. Nevertheless, we still have no information on
which concepts are used and what they mean. It is at this point that we query the
ontology for concept such as determinants, increase, risk, and diabetes. Fetching
these concepts by their textual representation, searching labels, synonyms and
other annotations, we obtain the following:</p>
      <p>
        It is noteworthy that misspelings, di erence in case and such are handled
outside of the ontology. We then check for special markings that de ne processing
triggers to activate. In our example, `has positive e ect on' requires transitively
walking upstream to identify additional causal factors. At this point, all of the
information needed to understand what the user requested has been gathered.
We would then reformulate the question into a format that can be answered by a
description logic (DL) reasoner, such as Fact++[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. From our example, we need
SubClassOf+ of `Health Determinants' that are described by: \ `has positive
e ect on' some (`is disposition to' only (results in some `diabetes mellitus') "
      </p>
      <p>The results, will be a list of health determinants that directly in uence the
risk of the event in a positive way. From our processing trigger associated with
`has positive e ect on', we know that the answer should also include any health
determinants that positively a ects health determinant having a direct in uence
on the risk event. We know, for example, that BMI is one of those direct factors
and that it is a measurable property. Therefore, we look for other health
determinants that have a positive e ect on a disposition to increase the level of BMI.
If the result is a measurable property we would repeat the same step.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Discussion</title>
      <p>Developing a system that is accessible via a natural language interface is
challenging. To address this challenge, we make use of all the contextual information
we can learn about the intent of the query, we restrict ourselves to a proper
context-free subset of the English language, and we use a domain ontology. The
resulting system gives useful and correct answers to practical questions. We
are looking forward validating our solution in user testing. It will enable us to
broaden our scope from a prototype state to that of day to day use.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>The Canadian Foundation for Innovation (CFI) and the Canadian Institutes of
Health Research (CIHR) provide funding for this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Buckeridge</surname>
            <given-names>DL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Izadi</surname>
            <given-names>MT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaban-Nejad</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mondor</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jauvin</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dube</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jang</surname>
            <given-names>Y</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tamblyn</surname>
            <given-names>R.</given-names>
          </string-name>
          <article-title>An infrastructure for real-time population health assessment and monitoring</article-title>
          .
          <source>IBM Journal of Research and Development</source>
          ,
          <year>2012</year>
          ,
          <volume>56</volume>
          (
          <issue>5</issue>
          ):
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Izadi</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaban-Nejad</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okhmatovskaia</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mondor</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckeridge</surname>
            <given-names>DL</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Population Health Record: An Informatics Infrastructure for Management, Integration, and Analysis of Large Scale Population Health Data</article-title>
          .
          <source>In Proc. of AAAI (HIAI</source>
          <year>2013</year>
          ), Belleveue, Washington, USA July 14-18.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Parr</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fisher</surname>
            <given-names>KS</given-names>
          </string-name>
          ,
          <article-title>LL(*): The Foundation of the ANTLR Parser Generator</article-title>
          ,
          <article-title>PLDI 2011</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Tsarkov</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Horrocks</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>FaCT++ description logic reasoner: system description</article-title>
          .
          <source>In proc. of IJCAR 2006</source>
          , Springer, pp.
          <fpage>292</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>