<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Collaborative Development of a Process Chemistry Domain Ontology, PROCO</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wes Schafer</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincent Antonucci</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yongqun Oliver He</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna Dunn</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zach E.X. Dance</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Nespor</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lama Saeeda</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Analytical Research &amp; Development, Merck &amp; Co., Inc.</institution>
          ,
          <addr-line>Rahway, NJ</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IT Eng., Dev. &amp; Integration, MSD</institution>
          ,
          <addr-line>Prague</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Research and Development Sciences IT, Merck &amp; Co., Inc.</institution>
          ,
          <addr-line>Rahway, NJ</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Michigan Medical School</institution>
          ,
          <addr-line>Ann Arbor, MI</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Process chemists embracing data mining, machine learning and artificial intelligence rapidly discover that the lack of structured data hampers their efforts. Although raw and processed instrument data have been addressed with public ontologies such as Allotrope and somewhat by commercial enterprise content management solutions, the chemical context of that data has been largely neglected. Recognizing this foundational gap and the impact it would have on developing new and better scientific data capture systems like electronic laboratory notebooks, Merck chemists reached out to peers in other companies and academia to develop PROCO, a domain ontology around process chemistry that studies the development and optimization of the production processes for chemical compounds. The scope was set using specific use cases and is being rounded out modeling public databases such as ORD. Development was based on “up-scaling” the chemists semantic skill sets and the domain knowledge of the ontologists. Simplified public ontology tools such as WebProtege provided a collaborative online space for ontology developers and subject matter experts with features like commenting, suggesting and approving changes. To increase internal adoption and applicability of PROCO, the ontology was loaded into Merck's master ontology for discovery, pre-clinical and early development space (MDO) via its CENtree ontology management system. Specific applications use MDO as the master source to develop their 'application ontologies'. As an example, ELN application ontology covers experimental metadata and feeds it into the Perkin-Elmer Signals notebook to define values of drop-down lists in the notebook. This enables standardized data capture which consequently makes the data interoperable and reusable for analytics / data science. Using ontologies as a metadata input for data capture enables data to be 'born FAIR' (findable, accessible, interoperable, and reusable) which is significantly more efficient and less expensive than FAIRifying the data at later stages. This industrial and academic collaboration proved to be an effective means of achieving better structured data with limited enterprise resources. The final PROCO ontology has been submitted to the OBO Foundry to broaden the development pool and usage of the ontology.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;PROCO</kwd>
        <kwd>Process Chemistry Ontology</kwd>
        <kwd>process chemistry</kwd>
        <kwd>ontology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body />
  <back>
    <ref-list />
  </back>
</article>