<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Holistic Approach to Scientific Reasoning Based on Hybrid Knowledge Representations and Research Objects</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andres Garcia Expert System Madrid</string-name>
          <email>jmgomez@expertsystem.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Spain rdenaux@expertsystem.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jose Manuel Gomez-Perez Expert System Madrid</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Raul Palma PSNC Poznan</institution>
          ,
          <country country="PL">Poland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Ronald Denaux Expert System Madrid</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>MOTIVATION AND GOALS</title>
      <p>
        Under the light of current developments in AI it appears the time
is ripe for a shared partnership with machines, whereby humans
can benefit from augmented reasoning and information
management capabilities provided that machines are endowed with the
necessary intelligence to assist with such tasks. This seems to be
particularly the case of the scientific domain, where some envision
the development of an AI that can make major scientific
discoveries and that eventually becomes worthy of a Nobel Prize [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This
vision may still be far from realization, but it is not completely new
nevertheless.
      </p>
      <p>
        NLP technologies based on well-formed, logically sound
structured knowledge representations (knowledge graphs, ontologies)
leverage expressive and actionable descriptions of the domain of
interest through logical deduction and inference, and can provide
logical explanations of reasoning outcomes. Closely related to this
family of approaches, project Halo [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] aimed to develop a
Digital Aristotle able to answer novel questions in scientific domains
with expertise equivalent to Advanced Placement competence level.
Halo enabled subject matter experts (SMEs) to model complex
scientific knowledge from textbooks and related questions, based on
an underlying logical formalism and a knowledge modeling
workbench to assist SMEs in the task. The resulting system achieved
an unprecedented question answering performance level for
SMEentered knowledge, but it also had a number of severe drawbacks,
including brittleness (coverage, precision or granularity gaps),
scalability issues, and the need for a considerable force of well trained
human labor to manually encode large amounts of scientific
knowledge.
      </p>
      <p>
        On the other hand, the last decade has witnessed a shift towards
statistical methods due to the increasing availability of raw data
and cheap computing power. These have proved to be powerful and
convenient in many linguistic tasks, such as part-of-speech tagging
or dependency parsing. However, they are also limited, e.g. humans
seek causal explanations, which are hard to provide based on
statistical induction rather than logical deduction. Recent results in the
ifeld of distributional semantics [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] have shown promising ways
to learn features from text that can complement the knowledge
already captured explicitly in structured representations.
Embeddings provide a compact and portable representation of words and
their meaning that stems directly from a document corpus. In this
scenario, a notion of semantic portability [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] emerges that refers to
the capability to capture as an information artifact (a vector) the
semantics of a linguistic unit (a word) from its occurrences in the
corpus and how such artifact enables that meaning to be merged
with other forms of knowledge representation.
      </p>
      <p>
        Furthermore, scientific knowledge is heterogeneous and can
present itself in many forms. During its analysis phase, Halo
produced an inventory of the diferent types of knowledge identified.
Such knowledge types include among others: factual knowledge,
procedural, classification, mathematical, diagrammatic, tabular and
experimental. It is therefore clear that successfully reading and
understanding scientific knowledge (either by humans or machines)
requires addressing the diferent knowledge types in a holistic way,
which remains a challenging task. We argue that addressing such
challenge requires generalizing the notion of semantic portability
from a text understanding scenario to a broader one where other
modalities, such as diagrams, processes, experiments and related
artifacts like scientific workflows and their execution provenance,
are also involved. This can be achieved by learning individual
models for each modality in the form of concept embeddings following
a distributional semantics [
        <xref ref-type="bibr" rid="ref12 ref8">8, 12</xref>
        ] and learning the corresponding
transformations between each vector space. The result will be a
shared, hybrid formalism that encompasses the diferent modalities
involved in scientific knowledge. Using embeddings to represent
not only words but arbitrary features has been recently popularized
by Chen and Manning in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        At this point, the question remains where to obtain the
crossmodal data required to learn such models and the necessary
transformations between them. We argue that the growing collections
of research objects from diferent scientific disciplines available
in repositories like ROHub.org [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] will play a key role in this
regard. Conceptually speaking, a research object [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a container
of scientific knowledge, a semantically rich aggregation of all the
materials involved in a scientific investigation, such as papers and
bibliography, numerical data, hypotheses, methods, experiments,
workflows encoding such experiments and the provenance of their
executions. A research object thus becomes the carrier of the
scientific knowledge associated to a specific investigation. They also
bring together all the necessary information to preserve scientific
Jose Manuel Gomez-Perez, Ronald Denaux, Andres Garcia, and Raul Palma
work against potential decay [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and can be shared, reused and
cited in scholarly communications. As scholars move away from
paper towards digital content, research objects have a key role to
play in the way scientific results are communicated and validated
by the communities, given the need for mechanisms that support
the production of self-contained publications involving not only
text but also data, methods and software implementations.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], we show how research objects are key pieces of a
humanmachine scientific partnership. Building on that, we aim at
furthering the role of research objects in such partnership, leveraging
research object corpora of cross-modal scientific knowledge to
develop hybrid models for scientific reasoning and question
answering. During the workshop, we aim at sharing and discussing
these ideas, explore related lines of work and establish areas of
common interest and collaboration with the participants. Key topics
and research questions we wish to address include: approaches for
hybrid reasoning, question answering and explanation, methods
to build portable knowledge representations of multimodal data,
how to combine the knowledge extracted from each modality in
the research objects to recompose a coherent, more complete view
of the scientific facts documented by them, and how each modality
interplay with each other in doing so.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>ABOUT THE AUTHORS</title>
      <p>
        This research is conducted by a team of researchers at Expert
System’s COGITO Lab and the Poznan Supercomputing and
Networking Center. Through the years, we have developed a body of work
in the intersection of several areas of AI that converge in the ideas
discussed in this document, including NLP, Knowledge Discovery,
Representation and Reasoning and new ways of scholarly
communication and preservation of scientific knowledge (as research
objects). This work aims at enabling machines to understand text
and other modalities in which knowledge can be expressed in a
way similar to how humans read, bridging the gap between both
through semantically rich knowledge representations and
humanmachine interfaces. In doing so, we believe that such vision is best
served through a combination of structured knowledge and
probabilistic approaches. The main author of this document participated
in project Halo as a member of the DarkMatter team, focused on
process knowledge acquisition from textbooks and question
answering by domain experts [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. He is also one of the founders
and key personnel behind ROHub.org, the reference platform for
research object management. ROHub currently hosts almost 2,500
research objects and 180 scientists in a variety of experimental and
observational scientific disciplines like Biology, Astrophysics and
Earth Science.
      </p>
    </sec>
    <sec id="sec-3">
      <title>ACKNOWLEDGMENTS</title>
      <p>This research is funded by the EU H2020 and national research
projects EVER-EST (674907), xLiMe-ES (20160805) and DANTE
(700367).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S</given-names>
            <surname>Bechhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I</given-names>
            <surname>Buchan</surname>
          </string-name>
          , D De Roure,
          <string-name>
            <given-names>P</given-names>
            <surname>Missier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Ainsworth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Bhagat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P</given-names>
            <surname>Couch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Cruickshank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Delderfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I</given-names>
            <surname>Dunlop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Gamble</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Michaelides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Owen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Sufi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C</given-names>
            <surname>Goble</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Why linked data is not enough for scientists</article-title>
          .
          <source>Future Generation Computer Systems</source>
          <volume>29</volume>
          ,
          <issue>2</issue>
          (
          <year>2013</year>
          ),
          <fpage>599</fpage>
          -
          <lpage>611</lpage>
          . https://doi.org/10. 1016/j.future.
          <year>2011</year>
          .
          <volume>08</volume>
          .004 Special section: Recent advances in e-Science.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Danqi</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>A Fast and Accurate Dependency Parser using Neural Networks</article-title>
          .
          <source>In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          .
          <article-title>Association for Computational Linguistics</article-title>
          , Doha, Qatar,
          <fpage>740</fpage>
          -
          <lpage>750</lpage>
          . http://www.aclweb.org/anthology/D14-1082
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Ronald</given-names>
            <surname>Denaux and Jose M Gomez-Perez</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Towards a Vecsigrafo: Portable Semantics in Knowledge-based Text Analytics</article-title>
          .
          <source>In Proceedings of the 2017 workshop on Hybrid Statistical Semantic Understanding and Emerging Semantics (HSSUES)</source>
          .
          <source>CEUR Workshop Proceedings, Held in Conjunction with the 16th International Semantic Web Conference</source>
          , Vienna, Austria.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Jose</given-names>
            <surname>Manuel</surname>
          </string-name>
          Gomez-Perez, Michael Erdmann,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Greaves</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Oscar</given-names>
            <surname>Corcho</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A Formalism and Method for Representing and Reasoning with Process Models Authored by Subject Matter Experts</article-title>
          .
          <source>IEEE Trans. on Knowl. and Data Eng</source>
          .
          <volume>25</volume>
          ,
          <issue>9</issue>
          (Sept.
          <year>2013</year>
          ),
          <fpage>1933</fpage>
          -
          <lpage>1945</lpage>
          . https://doi.org/10.1109/TKDE.
          <year>2012</year>
          .127
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Jose</given-names>
            <surname>Manuel</surname>
          </string-name>
          Gomez-Perez, Michael Erdmann, Mark Greaves, Oscar Corcho, and V. Richard Benjamins.
          <year>2010</year>
          .
          <article-title>A framework and computer system for knowledgelevel acquisition, representation, and reasoning with process knowledge</article-title>
          .
          <volume>68</volume>
          (
          <issue>10</issue>
          <year>2010</year>
          ),
          <fpage>641</fpage>
          -
          <lpage>668</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Jose</surname>
            <given-names>M</given-names>
          </string-name>
          <string-name>
            <surname>Gomez-Perez</surname>
          </string-name>
          ,
          <article-title>Andres Garcia-Silva, and</article-title>
          <string-name>
            <given-names>Raul</given-names>
            <surname>Palma</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Towards a Human-Machine Scientific Partnership Based on Semantically Rich Research Objects</article-title>
          . In eScience.
          <source>IEEE Computer Society</source>
          , 1-
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>David</given-names>
            <surname>Gunning</surname>
          </string-name>
          , Vinay K Chaudhri,
          <string-name>
            <surname>Peter E Clark</surname>
          </string-name>
          , Ken Barker,
          <string-name>
            <surname>Shaw-Yi</surname>
            <given-names>Chaw</given-names>
          </string-name>
          , Mark Greaves, Benjamin Grosof, Alice Leung,
          <string-name>
            <surname>David D McDonald</surname>
            ,
            <given-names>Sunil</given-names>
          </string-name>
          <string-name>
            <surname>Mishra</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>Others. 2010. Project</given-names>
            <surname>Halo Update-Progress Toward</surname>
          </string-name>
          Digital Aristotle.
          <source>AI</source>
          Magazine
          <volume>31</volume>
          ,
          <issue>3</issue>
          (
          <year>2010</year>
          ),
          <fpage>33</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Zellig</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Harris</surname>
          </string-name>
          .
          <year>1981</year>
          . Distributional Structure. Springer Netherlands, Dordrecht,
          <fpage>3</fpage>
          -
          <lpage>22</lpage>
          . https://doi.org/10.1007/
          <fpage>978</fpage>
          -94-009-8467-
          <issue>7</issue>
          _
          <fpage>1</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Hiroaki</given-names>
            <surname>Kitano</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Artificial Intelligence to Win the Nobel Prize and Beyond: Creating the Engine for Scientific Discovery</article-title>
          .
          <source>AI</source>
          Magazine
          <volume>37</volume>
          ,
          <issue>1</issue>
          (
          <year>2016</year>
          ),
          <fpage>39</fpage>
          -
          <lpage>49</lpage>
          . http://www.aaai.org/ojs/index.php/aimagazine/article/view/2642
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Tomác</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed Representations of Words and Phrases and their Compositionality.</article-title>
          . In NIPS. https://doi.org/10.1162/jmlr.
          <year>2003</year>
          .
          <volume>3</volume>
          .4-
          <fpage>5</fpage>
          .951 arXiv:
          <fpage>1310</fpage>
          .
          <fpage>4546</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Raul</surname>
            <given-names>Palma</given-names>
          </string-name>
          , Piotr Hołubowicz, Oscar Corcho,
          <string-name>
            <surname>Jose M Gomez-Perez</surname>
            , and
            <given-names>Cezary</given-names>
          </string-name>
          <string-name>
            <surname>Mazurek</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>ROHub-A Digital Library of Research Objects Supporting Scientists Towards Reproducible Science</article-title>
          .
          <source>In Semantic Web Evaluation Challenge</source>
          . Springer,
          <fpage>77</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Magnus</given-names>
            <surname>Sahlgren</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>The distributional hypothesis</article-title>
          .
          <source>Italian Journal of Linguistics 20</source>
          ,
          <issue>1</issue>
          (
          <year>2008</year>
          ),
          <fpage>33</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>JM</given-names>
            <surname>Gomez-Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K</given-names>
            <surname>Belhajjame</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G</given-names>
            <surname>Klyne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E</given-names>
            <surname>García-Cuesta</surname>
          </string-name>
          ,
          <string-name>
            <surname>A Garrido</surname>
          </string-name>
          , KM Hettne,
          <string-name>
            <given-names>M</given-names>
            <surname>Roos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>De Roure</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C</given-names>
            <surname>Goble</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Why workflows break - Understanding and combating decay in Taverna workflows</article-title>
          ..
          <source>In eScience. IEEE Computer Society</source>
          , 1-
          <fpage>9</fpage>
          . http://dblp.uni-trier.de/db/conf/eScience/eScience2012. html#ZhaoGBKGGHRRG12
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>