<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>ChemProp: A Dataset with Annotations for Instructional Language in Chemical Patents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sopan Khosla</string-name>
          <email>sopankh@amazon.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carolyn Rose</string-name>
          <email>cprose@cs.cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AWS AI Labs</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carnegie Mellon University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Chemical Patents, Information Extraction, Program Synthesis</institution>
          ,
          <addr-line>Coreference Resolution, Relation Extraction</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>The Third AAAI Workshop on Scientific Document Understanding</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>patents. In this paper, we propose a new set of annotations for the ChEMU Chemical Reaction Corpus. Our annotations (ChemProp) non-trivially incorporate the signals from ChEMU 2020 and 2021 schema to extract the instructional structure from chemical patents with details about inputs, outputs, and reaction attributes for each event in the reaction snippet. We propose a semi-automatic algorithm to create ChemProp and benchmark state-of-the-art models proposed for ChEMU 2020 and ChEMU 2021 on it. We hope that ChemProp can play an important part in modeling the instructional language present in chemical</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        chemical processes and synthesis, which are often
deChemical research relies heavily on the knowledge of puts) of the underlying chemical patent. For example,
theentire instructional structure (e.g. a chronological
sequence of inputs, reaction-steps, conditions, and
outeven though ChEMU 2020 schema tries to relate
reacabout new compounds [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Despite the significant value
scribed in chemical patents or research literature, withtion events with associated compounds or conditions,
patents also serving as a critical source of informationit only operates on named-entities, and therefore does
not cover important lexical items (noun-phrases) that
relies on costly manual processes 2[]. High influx of such
of the information present in these documents, extracd-escribe relevant reaction conditions and participants
tion and organization of this information still heavilyusing co-referring generic expressions, for example,the
mixture, the organic layer, or the filtrate .
present in these text3s,[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
documents in chemistry has introduced the need for auto- In this work, we propose an algorithm that augments
matic systems that can extract the structured knowledCgeLEF ChEMU 2020 annotations with ChEMU 2021
annotations to create a more complete annotation
frame
      </p>
      <p>CLEF ChEMU shared-task series released the ChEMU work for converting natural language chemical patents
Chemical Reaction Corpus that contains reaction snip-into structured recipes. Instructional language is a
usepets extracted from chemical patents.</p>
      <p>
        For ChEMUful structure that comprises of step-by-step instructions
duces five domain-specific relations (including bridging
2020 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the authors annotate information about relationt-hat need to be performed to complete a task.
Howships between reaction events (steps) and named-entitieesver, most of the prior art in the instructional language
involved in that step. ChEMU 20214[] on the other hand paradigm focuses on cooking recipes. We propose a new
focuses specifically on extracting chemical relations be- datasetC,hemProp1, that merges the ChEMU 2020 and
tween a pair of entity-mentions. The framework intro2-021 annotations to create labels for the instructional
language present in chemical patents. For each
reacin the discourse. Finally, ChEMU 2022 [5] reused the
and coreference) that link diferent noun-phrases present tion snippet, we annotate constituting events (reaction/
work-up steps), their relative chronological order, and
enexpression-level tasks from 2020 and 2021, and also in- tities that are associated with each of these events. More
troduced other document-level information extractionspecifically, for each reaction step in that snippet, we
      </p>
      <p>None of these shared tasks however fully capture (entities) that depict the (i) INPUT, (ii) OUTPUT, and (iii)
tasks.
nEvelop-O
LGOBE
(C. Rose)
htp:/ceur-ws.org
ISN1613-073</p>
      <p>Attribution 4.0 International (CC BY 4.0).</p>
      <p>CEUR</p>
      <p>Workshop Proceedings (CEUR-WS.org)
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License
annotate the trigger event verb, and the noun phrases
reaction-attributes (RXN_ATTR) of that reaction step.</p>
      <sec id="sec-2-1">
        <title>We leverage the raw reaction snippets from the ChEMU</title>
      </sec>
      <sec id="sec-2-2">
        <title>Chemical Reaction Corpus as our data and annotate it</title>
        <p>by (i) automatically combining the annotations of CLEF</p>
      </sec>
      <sec id="sec-2-3">
        <title>ChEMU shared tasks 2020 and 2021, and (ii) manually</title>
        <p>incorporating events/entities that are missed by the two</p>
      </sec>
      <sec id="sec-2-4">
        <title>1https://github.com/sopankhosla/chemprop</title>
        <p>annotation schemes. 2. Furthermore, CC20 does not capture relationships</p>
        <p>Furthermore, we show the significance of these aug- between reaction steps (events) and noun-phrase
mentations by evaluating the performance of the best mentions that denote combinations/ mixtures
performing models on ChEMU 2020 and 2021 shared- (e.g., the reaction mixture) or coreferent
exprestasks on ChemProp. Our experiments show that models sions (e.g., the product).
trained on ChemProp training data only achieve 0.69
Micro-F1 points on the test data, thus highlighting thCeLEF ChEMU 2021 Annotation Schema (CC21).
room for improvement. We also show that ChemProp Next year, He et al[.4] proposed an additional layer of
contains novel entities and relationships that are noatnnotation to the patents corpus, which focuses on the
present in ChEMU shared-tasks thus making it beneficial identification of anaphoric references. The new corpus
as a standalone benchmark for instructional language contains annotations for both COREFERENCE and
bridgmodeling from chemical patents. ing relations (Figure2). The authors define four
domainspecific sub-types for bridging: TRANSFORMED,
RE2. Prior Art ACTION_ASSOCIATED, WORK_UP, CONTAINED. As
a standalone schema, CC21 sufers from the following
In this section, we brielfy describe the ChEMU Chemical issues:
Reaction Corpus and the two state-of-the-art annotation 1. CC21 does not contain explicit information about
schemes proposed during ChEMU shared-tasks ’20 &amp; ’21. reaction steps. Therefore, it is less useful, in
isolation, for information extraction from chemical
patents.
2. Furthermore, CC21 difers from CC20 on its
definition ofmentions and therefore makes the
combination of two annotations non-trivial. In the
next section, we describe the algorithm to handle
these ambiguities.</p>
        <sec id="sec-2-4-1">
          <title>CLEF ChEMU 2020 Annotation Schema (CC20).</title>
          <p>
            He et al.[
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] annotated a corpus of 1,500 patent
snippets sampled from 170 patents from the European Patent
Ofice and the United States Patent and Trademark Ofice.
          </p>
          <p>Their annotation schema aims at extraction of chemical
reactions (i.e. REACTION_STEP, WORKUP) from patent
snippets. It identifies trigger words that describe reaction
steps and relates them to named-entities linked to the
step (i.e. chemical compounds, time, temperature, and3. ChemProp: Annotation
yields; Figure 1). Despite being a comprehensive
annotation schema, CC20 sufers from two major drawbacks: CC20 contains relationships between events and
named1. CC20 does not annotate reaction steps that deontities, whereas CC21 connects noun phrases (including
not relate to any named-entity in the discoursenamed-entities) based on their anaphoric relationships.
snippet. E.g., as shown in Figure1, CC20 does Together, CC20 and CC21 provide somewhat
complenot annotate the evencotncentrated (in line 7). mentary information about each reaction snippet in the</p>
          <p>ChEMU corpus. In this section, we describe the steps we3.1.1. Pre-processing
take to merge these somewhat heterogeneous annotation
schema to create our new dataset ChemProp. As a pre-processing step, we setup data-structures that
help with the conversion algorithm. We create
First, we present our algorithm that automatically merges
signals from CC20 and CC21 and extracts lexical spans
in the discourse, that most closely represent the inputs,
outputs, &amp; reaction-attributes for each reaction step.</p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>1. A many-to-one map from CC20 named-entities</title>
        <p>to CC21 named-entity annotations to tackle
the small annotation diferences between the
two schemas. The map is many-to-one because
CC21 annotates entire solutions, whereas CC20
keeps the individual compounds/elements, e.g.,
[a solution ofethanol 20 and water 20 ] 21 .</p>
      </sec>
      <sec id="sec-2-6">
        <title>We also store the inverse one-to-many map from CC21 to CC20. So for the given example, we store</title>
        <p>ethanol20
wate r20</p>
        <p>↔ a solution of ethanol and wat e2r1
↔ a solution of ethanol and wat e2r1 .</p>
      </sec>
      <sec id="sec-2-7">
        <title>2. The mapping from ARG1 of CC21 reaction</title>
        <p>oriented relations (i.e.,
REACTION_ASSOCIATED (R_ASSOC), WORK_UP, and
TRANSFORMED (TRANS)) to their corresponding ARG2
mentions. For example, for the relation
TRANSFORMED betweenThe reaction mixture and The
reaction in Figure 2 (line 4, 5), we store
The reaction mixture21</p>
        <p>Rxn:TRANS
−−−−−−−−−→ The reactio n21 .</p>
        <p>1. If ARG2 21 is also present in the CC20
annotation, we use thisARG2 21 as a pivot to combine
the two annotations. We extract the CC20
relation (Rxn 20 ) it is linked with.
2. If ARG2 21 is not present in CC20, it means that</p>
        <p>ARG2 21 is a non named-entity noun-phrase
and therefore needs to be resolved further. To
ground such cases, we rely on the fact that the
starting compounds in each reaction snippet
starts are named-entities. Therefore, we can
assume that in order to reach currenAtRG1 21 , all
otherARG2 21 s that also appeared asARG1 21
(noun-phrases) have been resolved in earlier
iterations of this algorithm. Therefore, the
latestRxn 20 in the patent snippet between the
CC20 reaction event associated witAhRG2 21
(i.e. ARG2 21 →− Rxn 20 ) and ARG1 21 is
returned.
3. The mapping from CC20 events (ARG1) to their From this list of relations, we consider thRexn 20 that
corresponding argument mentions (ARG2). For is closest toARG1 21 , but occurs before it, to be the
example, for the eventstirred in Figure 1 (line 4), lexical event trigger that outputs AR G211 :
we store
stirre d20
stirre d20</p>
        <p>Event:TEMP
←−−−−−−−−→ r.t.20
Event:TIME
←−−−−−−−−→ 2 h 20 .</p>
        <p>Cand_Rxn 20
Rxn 20
ARG1 21
= [Rxn 20 , Rxn220 , ..., Rxn 20 , ...]</p>
        <p>1
= 
_</p>
        <p>(ARG1, Cand_Rxn 20 )
Output:Rxn
←−−−−−−−→ Rxn 20 .</p>
      </sec>
      <sec id="sec-2-8">
        <title>4. Finally, a dictionary to store the COREFERENCE</title>
        <p>relationships between diferent CC21 mentions. Consider the case whereARG1 21 is The combined
orWe create a separate map for coreference as they ganic phases (Figure 2; line 6). It is related to four
mendenote that both mentions in the pair are equivt-ions (ARG2 21 s) {The reaction mixture, a pad of celite,
alent i.e., point to the same underlying entity inwater, ethyl acetate} by the relation WORK_UPR( xn 21 ).
the discourse: Three of these mentions{a pad of celite, water, ethyl
acetate} are also present in the mapping created in
preARG1 21 ←C−o−r−→ef ARG2 21 . processing step 1, whereas{The reaction mixture} is not.</p>
        <p>As discussed earlier, forARG2 21 s present in CC20, we
The above mappings store the relationships annotated infirst create a list ofRxn 20 s (Cand_Rxn 20 ) they
corCC20 and CC21 that are relevant to creating ChemPropr.espond to{filtered, diluted, extracted} (Figure 1). For The
reaction mixture, the correspondingRxn 20 , based on
3.1.2. Algorithm its resolution in the previous step,Tishe reaction mixture
←→ stirred.</p>
        <p>We allow for three relationships between reaction events The latest CC20 relation betweenstirred and
and entities in our new ChemProp benchmark – IN- ARG2 21 , extracted, is then inserted intCoand_Rxn 20 .
PUT, OUTPUT, RXN_ATTR. CC20 already annotates Combining the four relations we ge{textracted, filtered,
RXN_ATTRs (reaction-attributes) and named-entity
IN</p>
        <p>diluted, extracted}. From these,extracted occurs closest
PUTS/ OUTPUTS. Therefore, to complete the schema, to the current ARG1 while occurring before it.
Therewe devise an algorithm that can find a mapping between fore, The combined organic phases is considered to be the
each non named-entity noun-phrase in CC21 to a reac- OUTPUT of extracted (The combined organic phases ←→
tion event in CC20 (AR G21 ←→ Rxn 20 ). extracted; Figure 3).</p>
        <p>For each CC21 entity A(RG1 21 ) that is related to
other CC21 entitiesA(RG2 21 ) (that occur before it Exceptional Cases. Although, most of the events can
in the discourse) via one of the three reaction-orientedbe fully annotated in ChemProp format using the above
CC21 relations (Rxn 21 ) i.e., REACTION_ASSOCIATED,
WORK_UP, and TRANSFORMED, two possibilities need steps, there are certain exceptions that arise due to the
mismatch between the motivation of CC20 and CC21
to be considered. For each ARG 221 in ARG2 21 : schemas.
about the chemical patent snippets that will be
systemat3.2. Manual Quality Assurance ically missed by the systems trained on CC20.
Next, we manually go through the development and test We observe thatCC21_best (gold entities) achieves
data to fix the exceptions (described in the previous sec- an F1 score of 0.86 on ChemProp test set. This suggests
tion). that the model is able to somewhat reliably figure out</p>
        <p>As discussed earlier, the CC20 annotation does notwhich named-entities/ noun phrases are related to which
annotate reaction steps that do not relate to any namerdea-ction event in the snippetC.C21_best (end-to-end), in
entity in the snippet. However, in our case such eventsaddition to relation classification, also extracts mentions
are equally relevant and need to be extracted to get farom raw patent snippets, and therefore expectedly
per</p>
        <sec id="sec-2-8-1">
          <title>System</title>
          <p>Trained on ChemProp
CC21_best (gold-entities)
CC21_best (end-to-end)
0.74
0.62
forms much worse thanCC21_best (gold entities) with
an overall F1 score of 0.69.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Conclusion</title>
      <p>In this work, we propose a new corpus ChemProp that
non-trivially combines properties from ChEMU 2020 and
2021 annotation schema to extract the instructional
structure from chemical patents. We provide a semi-automatic
algorithm to create ChemProp. Evaluating
state-of-theart models on the our new dataset suggests that there
is still room for improvement in extracting relevant
instructional triggers from patent text. We believe that
ChemProp can act as an important benchmark for
instructional language modeling.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Senger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bartek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Papadatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gaulton</surname>
          </string-name>
          ,
          <article-title>Managing expectations: assessment of chemistry databases generated by automated extraction of chemical structures from patents</article-title>
          ,
          <source>Journal of cheminformatics 7</source>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Muresan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Petrov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Southan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Kjellberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kogej</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tyrchan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Varkonyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>Making every sar point count: the development of chemistry connect for the large-scale integration of structure and bioactivity data</article-title>
          ,
          <source>Drug Discovery Today</source>
          <volume>16</volume>
          (
          <year>2011</year>
          )
          <fpage>1019</fpage>
          -
          <lpage>1030</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoessel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yoshikawa</surname>
          </string-name>
          , et al.,
          <source>Overview of chemu</source>
          <year>2020</year>
          <article-title>: named entity recognition and event extraction of chemical reactions from patents</article-title>
          ,
          <source>in: International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>237</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yoshikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cavedon</surname>
          </string-name>
          , et al.,
          <year>Chemu 2021</year>
          :
          <article-title>reaction reference resolution and anaphora resolution in chemical patents</article-title>
          ,
          <source>in: ECIR (2)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>