<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Extended Overview of ChEMU 2021: Reaction Reference Resolution and Anaphora Resolution in Chemical Patents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuan Li</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Biaoyan Fang</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiayuan He</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hiyori Yoshikawa</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saber A. Akhondi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Druckenbrodt</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Camilo Thorne</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zubair Afzal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zenan Zhai</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Timothy Baldwin</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karin Verspoor</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Elsevier BV</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Elsevier Information Systems GmbH</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Fujitsu Limited</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>RMIT University</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>The University of Melbourne</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>2</volume>
      <fpage>1</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>In this paper, we provide an overview of the Cheminformatics Elsevier Melbourne University (ChEMU) evaluation lab 2021, part of the Conference and Labs of the Evaluation Forum 2021 (CLEF 2021). The ChEMU evaluation lab focuses on information extraction over chemical reactions from patent texts. As the second instance of our ChEMU lab series, we build upon the ChEMU corpus developed for ChEMU 2020, extending it for two distinct tasks related to reference resolution in chemical patents. Task 1 Chemical Reaction Reference Resolution - focuses on paragraph-level references and aims to identify the chemical reactions or general conditions specified in one reaction description referred to by another. Task 2 - Anaphora Resolution - focuses on expression-level references and aims to identify the reference relationships between expressions in chemical reaction descriptions. Herein, we describe the resources created for these tasks and the evaluation methodology adopted. We also provide a brief summary of the results obtained in this lab, finding that one submission achieves substantially better results than our baseline models.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Reaction reference resolution</kwd>
        <kwd>Anaphora resolution</kwd>
        <kwd>Chemical patents</kwd>
        <kwd>Text mining</kwd>
        <kwd>Information Extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The discovery of new chemical compounds is perceived as a key driver of the chemical industry
and many other industrial sectors, and information relevant for this discovery is found in
chemical synthesis descriptions in natural language texts. In particular, patents serve as a critical
source of information about new chemical compounds. Compared with journal publications,
patents provide more timely and comprehensive information about new chemical compounds [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1,
2, 3</xref>
        ], since they are usually the first venues where new chemical compounds are disclosed.
      </p>
      <p>
        Despite the significant commercial and research value of the information in patents, manual
extraction of such information is costly, considering the large volume of patents available [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
Thus, developing automatic natural language processing (NLP) systems for chemical patents,
which convert text corpora into structured knowledge about chemical compounds, has become
a focus of recent research [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ].
      </p>
      <p>
        The ChEMU campaign focuses on information extraction tasks over chemical reactions
in patents1. ChEMU 2020 [
        <xref ref-type="bibr" rid="ref6 ref8 ref9">6, 8, 9</xref>
        ] provided two information extraction tasks, named entity
recognition (NER) and event extraction, and attracted 37 teams around the world to participate.
In the ChEMU 2021 lab, we provide two new information extraction tasks: chemical reaction
reference resolution and anaphora resolution, focusing on reference resolution in chemical
patents. Compared with previous shared tasks dealing with anaphora resolution, e.g., the
CRAFT-CR task [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], our proposed tasks extend the scope of reference resolution by considering
reference relationships on both paragraph-level and expression-level (see Fig. 1). Specifically, our
ifrst task aims at the identification of reference relationships between reaction descriptions. Our
second task aims at the identification of reference relationships between chemical expressions,
including both coreference and bridging. Moreover, we focus on chemical patents while the
CRAFT-CR task focused on journal articles.
      </p>
      <p>Unfortunately, we didn’t receive any submissions to Task 1, chemical reaction reference
resolution. The complexity of this task in particular combined with relatively short time periods
for people to develop their systems may have made it dificult for people to participate. We plan
to re-run it in 2022, to give the opportunity for more people to participate since the data and
task definitions will have been around for a longer period of time. As a result, the remainder of
this paper will focus on the second task, anaphora resolution.</p>
      <p>The rest of the paper is structured as follows. We first discuss related work and shared tasks
in Section 2 and introduce the corpus we created for use in the lab in Section 3. Then we give
an overview of the task in Section 4 and detail the valuation framework of ChEMU in Section 5
including the evaluation methods and baseline models. We present the evaluation results in
Section 6 and finally conclude this paper in Section 7.</p>
      <sec id="sec-1-1">
        <title>1Our main website is http://chemu.eng.unimelb.edu.au</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Shared Tasks</title>
      <p>
        Several shared tasks have addressed reference resolution in scientific literature. BioNLP2011
hosted a subtask on protein coreference [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. CRAFT 2019 hosted a subtask on coreference
resolution (CRAFT-CR) in biomedical articles [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. However, these shared tasks difer from ours
in several respects.
      </p>
      <p>
        First, previous shared tasks considered diferent domains of scientific literature. For example,
the dataset used in BioNLP2011 is derived from the GENIA corpus [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], which primarily focuses
on the biological domain, viz. gene/proteins and their regulations. The dataset used in
CRAFTCR shared task is based on biomedical journal articles in PubMed [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. Our ChEMU shared
task, in contrast, focuses on the domain of chemical patents. This diference entails the critical
importance for this shared task: information extraction methodologies for general scientific
literature or the biomedical domain will not be efective for chemical patents [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. It is widely
acknowledged that patents are written quite diferently as compared with general scientific
literature, resulting in substantially diferent linguistic properties. For example, patent authors
may trade some clarity in wording for more protection of their intellectual property.
      </p>
      <p>Secondly, our reference resolution tasks include both paragraph-level and entity-level
reference phenomena. Our first task aims at identification of reference relationships between reaction
descriptions, i.e. paragraph-level. This task is challenging because a reaction description may
refer to an extremely remote reaction and thus requires processing of very long documents. Our
second task aims at anaphora resolution, similarly to previous entity-level coreference tasks.
However, a key diference is that we extend the scope of this task by including both coreference
and bridging phenomena. That is, we not only aim at finding expressions referring to the same
entity, but also expressions that are semantically related or associated.</p>
    </sec>
    <sec id="sec-3">
      <title>3. The ChEMU Chemical Reaction Corpus</title>
      <p>
        In this section, we explain how the dataset is created for the anaphora resolution task. The
complete annotation guidelines are made available at [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>3.1. Corpus Selection</title>
        <p>
          We build on the ChEMU corpus [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] developed for the ChEMU 2020 shared task [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. The ChEMU
corpus contains patents from the European Patent Ofice and the United States Patent and
Trademark Ofice, available in English in a digital format. It is based on the Reaxys ® database,2
containing reaction entries for patent documents manually created by experts in chemistry. It
consists of ‘snippets’ extracted from chemical patents, where each snippet corresponds to a
reaction description. It is common that several snippets are extracted from the same chemical
patent.
        </p>
        <p>2Reaxys® Copyright ©2021 Elsevier Life Sciences IP Limited except certain content provided by third parties.
Reaxys is a trademark of Elsevier Life Sciences IP Limited, used under license. https://www.reaxys.com</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Mention Type</title>
        <p>
          We aim to capture anaphora in chemical patents, with a focus on identifying chemical compounds
during the reaction process. Consistent with other anaphora corpora [
          <xref ref-type="bibr" rid="ref13 ref19 ref20">19, 13, 20</xref>
          ], only mentions
that are involved in referring relationships (as defined in Section 3.3) and related to chemical
compounds are annotated. The mention types that are considered for anaphora annotation are
listed below. It should be noted that verbs (e.g. mix, purify, distil) and descriptions that refer to
events (e.g. the same process, step 5) are not annotated in this corpus.
3.2.1. Chemical Names
Chemical names are a critical component of chemical patents. We capture as atomic mentions
the formal name of chemical compounds, e.g.
N-[4-(benzoxazol-2-yl)-methoxyphenyl]-S-methylN’-phenyl-isothiourea or 2-Chloro-4-hydroxy-phenylboronic acid. Chemical names often include
nested chemical components, but for the purposes of our corpus, we consider chemical names
to be atomic and do not separately annotate internal mentions. Hence
4-(benzoxazol-2-yl)methoxyphenyl and acid in the examples above will not be annotated as mentions, as they are
part of larger chemical names.
3.2.2. Identifiers
In chemical patents, identifiers or labels may also be used to represent chemical compounds,
in the form of uniquely-identifying sequences of numbers and letters such as 5i. These can be
abbreviations of longer expressions incorporating that identifier that occur earlier in the text,
such as chemical compound 5i, or may refer back to an exact chemical name with that identifier.
Thus, the identifier is annotated as an atomic mention as well.
3.2.3. Phrases and Noun Types
Apart from chemical names and identifiers, chemical compounds are commonly presented as
noun phrases (NPs). An NP consists of a noun or pronoun, and premodifiers; NPs are the most
common type of compound expressions in chemical patents. Here we detail NPs that are related
to compounds:
1. Pronouns: In chemical patents, pronouns (e.g. they or it) usually refer to a previously
mentioned chemical compounds.
2. Definite and indefinite NPs: Commonly used to refer to chemical compounds, e.g. the
solvent, the title compound, the mixture, and a white solid, a crude product.
        </p>
        <p>Furthermore, there are a few types of NPs that need specific handling in chemical patents:
1. Quantified NPs: Chemical compounds are usually described with a quantity. NPs with
quantities are considered as atomic mentions if the quantities are provided, e.g. 398.4 mg
of the compound 1.
2. NPs with prepositions: Chemical NPs connected with prepositions (e.g. in, with, of ) can be
considered as a single mention. For example, the phrase
2,4-dichloro-6-(6-triuoromethylpyridin2-yl)-1,3,5-triazine (5.0 g, 16.9 mmol) in tetrahydrofuran (100 mL) is a single mention,
as it describes a solvent that contains
2,4-dichloro-6-(6-triuoromethylpyridin-2-yl)-1,3,5triazine (5.0 g, 16.9 mmol) and tetrahydrofuran (100 mL).</p>
        <p>NPs describing chemical equipment containing a compound may also be relevant to anaphora
resolution. This generally occurs when the equipment that contains the compound undergoes a
process that also afects the compound. Thus, equipment expressions such as the flask and the
autoclave can also be mentions if they are used to implicitly refer to a contained compound.</p>
        <p>
          Unlike many annotation schemes, our annotation allows discontinuous mentions. For
example, the underlined spans of the fragment 114 mg of
4-((4aS,7aS)-6-benzyloctahydro-1-pyrrolo[3,4b]pyridine-1-yl)-7H-pyrrolo[2,3-d]pyrimidine was obtained with a yield of about 99.1% are treated
as a single discontinuous mention. This introduces further complexity into the task and helps
to capture more comprehensive anaphora phenomena.
3.2.4. Relationship to ChEMU 2020 entities
Since this dataset is built on the ChEMU 2020 corpus [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], annotation of related chemical
compounds is available by leveraging existing entity annotations introduced for the ChEMU
2020 named entity recognition (NER) task. However, there are some diferences in the definitions
of entities for the two tasks.
        </p>
        <p>In the original ChEMU 2020 corpus, entity annotations identify chemical compounds (i.e.
REACTION_PRODUCT, STARTING_MATERIAL, REAGENT_CATALYST, SOLVENT, and OTHER
COMPOUND), reaction conditions (i.e. TIME, TEMPERATURE), quantity information (i.e.
YIELD_PERCENT, YIELD_OTHER), and example labels (i.e. EXAMPLE_LABEL). There is
overlap with our definition of mention for the labels relating to chemical compounds. However,
in our annotation, chemical names are annotated along with additional quantity information,
as we consider this information to be an integral part of the chemical compound description.
Furthermore, the original entity annotations do not include generic expressions that corefer
with chemical compounds such as the mixture, the organic layer, or the filtrate , and neither do
they include equipment descriptions.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Relation Types</title>
        <p>Anaphora resolution subsumes both coreference and bridging. In the context of chemical
patents, we define four sub-types of bridging, incorporating generic and chemical knowledge.</p>
        <p>
          A referring mention which cannot be interpreted on its own, or an indirect mention, is
called an anaphor, and the mention which it refers back to is called the antecedent. In relation
annotation, we preserve the direction of the anaphoric relation, from the anaphor to the
antecedent. Following similar assumptions in recent work, we restrict annotations to cases
where the antecedent appears earlier in the text than the anaphor.
3.3.1. Coreference
Coreference is defined as expressions/mentions that refer to the same entity [
          <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
          ]. In
chemistry, identifying whether two mentions refer to the same entity needs to consider various
chemical properties (e.g. temperature or pH). As such, for two mentions to be coreferent, they
must share the same chemical properties. We consider two diferent cases of coreference:
        </p>
        <sec id="sec-3-3-1">
          <title>1. Single Antecedents: the anaphor refers to a single antecedent.</title>
          <p>2. Multiple Antecedents: the anaphor refers to multiple antecedents, e.g. start materials
refers to all the chemical compounds or materials that are used at the beginning.
It is possible for there to be ambiguity as to which mention of a given antecedent an anaphor
refers to (where the mention is identical); in these cases the closest mention is selected.
3.3.2. Bridging
As stated above, when we consider the anaphora relations, we take the chemical properties of
the mention into consideration. Coreference is insuficient to cover all instances of anaphora in
chemical patents, and bridging occurs frequently. We define four bridging types:
TRANSFORMED Links between chemical compounds that are initially based on the same
components, but which have undergone a change in condition, such as pH or temperature. Such
cases must be one-to-one relations (not one-to-many). As shown in Figure 2, the mixture in line
2 and the first-mentioned mixture in line 3 have the TRANSFORMED relation, as they have the
same chemical components but diferent chemical properties.</p>
          <p>REACTION-ASSOCIATED The relationship between a chemical compound and its
immediate source compounds is via a mixing process, where the source compounds retain their original
chemical structure. This relation is one-to-many from the anaphor to the source compounds
(antecedents). For example, the mixture in line 2 has REACTION-ASSOCIATED links to three
mentions on line 1 that are combined to form it: (1) the solution of Compound (4) (0.815 g, 1.30
mmol) in THF (4.9 ml); (2) acetic acid (9.8 ml); and (3) water (4.9 ml)).</p>
          <p>WORK-UP Chemical compounds are used to isolate or purify an associated output product,
in a one-to-many relation, from the anaphor to the compounds (antecedents) that are used for
the work-up process. As demonstrated in Figure 2, The combined organic layer in line 5 comes
from the extraction of The mixture and ethyl acetate in line 4, and they are hence annotated as
WORK-UP.</p>
          <p>CONTAINED A chemical compound is contained inside equipment. It is a one-to-many
relation from the anaphor (equipment) to the compounds (antecedents) that it contains. An
example of this is a flask and the solution of Compound (4) (0.815 g, 1.30 mmol) in THF (4.9 ml)
on line 1, where the compound is contained in the flask.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Annotation Process</title>
        <p>
          For the corpus annotation, we use the BRAT text annotation tool.3 In total 1500 snippets have
been annotated by two chemical experts, a PhD candidate and a final year bachelor student
in Chemistry. A draft of the annotation guideline was created and refined with chemical
experts, then four rounds of annotation training were completed prior to beginning oficial
annotation. In each round, the two annotators individually annotated the same 10 snippets
(diferent across each round of annotation), and their annotations were compared and combined
by an adjudicator; annotation guidelines were then refined based on discussion. After several
rounds of training, we achieved a high inner-annotator agreement of Krippendorf’s  = 0.92
[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] at the mention annotation level,4 and  = 0.84 for relations. Finally, the development and
test sets were double annotated by the two expert annotators, with any disagreements merged
by the adjudicator.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Data Partitions</title>
        <p>We randomly partitioned the whole dataset into three splits for training, development, and
test purposes, with a ratio of 0.6/0.15/0.25. The training and development sets were released to
participants for model development. Note that participants are allowed to use the combination
of training and development sets and to use their own partitions to build models. The test set
is withheld for use in the formal evaluation. The statistics of the three splits including their
number of snippets, total number of sentences, and average number of tokens per sentence, are
summarized in Table 1.</p>
        <p>To ensure the snippets included in the training, development, and test splits have similar
distributions, we compare the distribution of relation types (five types of relations in total).
Based on the numbers in Table 1, we confirm that the label distribution in the three splits are
similar, with very little variation (≤ 2%) across the three splits observed for each relation type.</p>
        <sec id="sec-3-5-1">
          <title>3https://brat.nlplab.org/</title>
          <p>4With the lowest agreement being  = 0.89 for coreference mentions.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Task definition</title>
      <p>This task requires the resolution of general anaphoric dependencies between expressions in
chemical patents. Five types of anaphoric relationships are defined:
1. Coreference: two expressions/mentions that refer to the same entity.
2. Transformed: two chemical compound entities that are initially based on the same
chemical components and have undergone possible changes through various conditions (e.g.,
pH and temperature).
3. Reaction-associated: the relationship between a chemical compound and its immediate
sources via a mixing process. The immediate sources do need to be reagents, but they
need to end up in the corresponding product. The source compounds retain their original
chemical structure.
4. Work-up: the relationship between chemical compounds that were used for isolation or
purification purposes, and their corresponding output products.
5. Contained: the association holding between chemical compounds and the related
equipment in which they are placed. The direction of the relation is from the related equipment
to the previous chemical compound.</p>
      <p>Taking the text snippet in Figure 3 as an example, several anaphoric relationships can be
extracted from it. [The mixture]4 and [the mixture]3 refer to the same “mixture” and thus,
form a coreference relationship. The two expressions [The mixture]1 and [the mixture]2 are
initially based on the same chemical components but the property of [the mixture]2 changes
after the “stir” and “cool” action. Thus, the two expressions should be linked as “Transformed”.
The expression [The mixture]1 comes from mixing the chemical compounds prior to it, e.g.,
[water (4.9 ml)]. Thus, the two expressions are linked as “Reaction-associated”. The expression
[The combined organic layer] comes from the extraction of [ethyl acetate]. Thus, they are
linked as “Work-up”. Finally, the expression [the solution] is contained by the entity [a flask] ,
and the two are linked as “Contained”.</p>
      <p>[Acetic acid (9.8 ml)] and [water (4.9 ml)] were added to [the solution] in [a flask]. [The mixture]1 was stirred
for 3 hrs at 50°C and then cooled to 0°C . 2N-sodium hydroxide aqueous solution was added to [the mixture]2 until the
pH of [the mixture]3 became 9. [The mixture]4 was extracted with [ethyl acetate] for 3 times. [The combined
organic layer] was washed with water and saturated aqueous sodium chloride.</p>
      <p>ID
AR1
AR2
AR3
AR4
AR5</p>
      <p>Relation type
Coreference
Transformed
Reaction_associated
Work-up
Contained</p>
      <p>Anaphor
[The mixture]4
[the mixture]2
[The mixture]1
[The combined organic layer]
[a flask]</p>
      <p>Antecedent</p>
    </sec>
    <sec id="sec-5">
      <title>5. Evaluation Framework</title>
      <sec id="sec-5-1">
        <title>5.1. Evaluation Methods</title>
        <p>We use BRATEval5 to evaluate all the runs that we receive. Three metrics are used to evaluate
the performance of all the submissions: Precision, Recall, and 1 score. We use two diference
matching criteria, exact matching and relaxed matching (approximate matching), as in some
practical applications it also makes sense to understand if the model can identify the approximate
region of mentions.</p>
        <p>Formally, let  = ( , , ) denote an entity where  is the type of ,  and  are the
beginning position (inclusive) and end position (exclusive) of the text span of . Then two
entities 1 and 2 are exactly matched (1 = 2), if 1 = 2, 1 = 2, and 1 = 2.
While two entities 1 and 2 are approximately matched (1 ≈ 2) if 1 = 2, 2 &lt; 1,
and 1 &lt; 2, i.e. the two spans [1, 1) and [2, 2) overlaps.</p>
        <p>Furthermore, let  = ( , , ) be a relation where  is the type of ,  the
anaphor of ,  the antecedent of . Then 1 and 2 are exactly matched (1 = 2) if
1 = 2, 1 = 2, and 1 = 2. While 1 and 2 are approximately matched
(1 ≈ 2) if 1 = 2, 1 ≈ 2, and 1 ≈ 2.</p>
        <p>In summary, we require strict type match in both exact and relaxed matching, but are lenient
in span matching.
5.1.1. Exact Matching
With the above definitions, the metrics for exact matching can be easily calculated. The true
positives (TP) are exact matching pairs found in gold relations and predicted relations. Then false
positives (FP) are the predicted relations that don’t have a match, i.e.   = # −   , where
# is the number of predicted relations. Similarly, false negatives   are the gold relations
that are not matched by any predicted relations, i.e.   = # −   where # is the
5https://bitbucket.org/nicta_biomed/brateval/src/master/
4
3
2
1
0
4
3
2
1
0
4
3
2
1
0
4
3
2
1
0
4
3
2
1
0
4
3
2
1
0
(a) All pairs matched in relaxed
setting
(b) A non-optimal bipartite
matching (  = 3)
(c) A maximum bipartite
matching (  = 4)
number of gold relations. Finally Precision  =  /(  +   ), Recall  =  /(  +   ),
and 1 = 2/(1/ + 1/).
5.1.2. Relaxed Matching
Unlike exact matching, relaxed matching is not well-defined and metrics in this setting have
more than one way to calculate, therefore we need to clearly define all the metrics.
5
Let consider an example shown in Figure 4a where nodes {}=1 are predicted relations,
5
{}=1 are gold relations, and every edge between a  node and a  node means they are
approximately matched. At first glance, one may think that   =   = 0 because every gold
relation has at least a match and so does every predicted relation. However, it is impossible to
ifnd 5 true positive pairs from this graph without using one node more than once. Therefore,
if   =   = 0, then   +   ̸= # = 5 and   +   ̸= # = 5, which is
inconsistent with the formulas in exact setting.</p>
        <p>So, instead of defining   as the number of gold relations that don’t have a match, we
just define   = # −   . Similarly   is defined as # −   . Then the problem
remained is how to calculate   . Actually, finding true positive pairs can be considered as
bipartite matching. Figure 4b shows a matching with   = 3 but is not optimal. Figure 4c
shows one possible maximum bipartite matching with   = 4. Another optimal matching is
replacing edge 0 − 0 with 0 − 1.</p>
        <p>In summary, we define   as the maximum bipartite matching for the graph constructed by
all approximately matched pairs, then   = # −   and   = # −   , finally
Precision  =  /(  +   ), Recall  =  /(  +   ), and 1 = 2/(1/ + 1/). This
has been implemented in the latest BRATEval.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Coreference Linkings</title>
        <p>We consider two types of coreference linking, i.e. (1) surface coreference linking and (2) atomic
coreference linking, due to the existence of transitive coreference relationships. By transitive
coreference relationships we mean multi-hop coreference such as a link from an expression
T1 to T3 via an intermediate expression T2, viz., “T1→T2→T3”. Surface coreference linking
will restrict attention to one-hop relationships, viz., to: “T1→T2” and “T2→T3”. Whereas
atomic coreference linking will tackle coreference between an anaphoric expression and its first
antecedent, i.e. intermediate antecedents will be collapsed. Thus, these two links will be used
for the above example, “T1→T3” and “T2→T3”. Note that we only consider transitive linking
in coreference relationships.</p>
        <p>Note that {T1→T2,T2→T3} infers {T1→T3,T2→T3}, but the reverse is not true. This leads
to a problem about how to score a prediction {T1→T3,T2→T3}, when the gold relation is
{T1→T2,T2→T3}. Both T1→T3 and T2→T3 are true, but some information is missing here.</p>
        <p>Our solution is to first expand both the prediction set and gold set where all valid relations
that can be inferred will be generated and added to the set, and then to evaluate the two sets
normally. In the above example, the gold set will be expanded to {T1→T2,T2→T3,T1→T3}, and
then the result is   = 2,   = 1. Likewise, when evaluate {T1→T4,T2→T4,T3→T4} against
{T1→T2,T2→T3,T3→T4}, the gold set will be expanded into 6 relations, while the prediction set
won’t be expanded as no new relation can be inferred. So the evaluation result will be   = 3,
  = 3. One may worry that if there is a chain of length  then its expanded set will be in
(2), when  is large, this local evaluation result will have too much influence on the overall
result. But we find in practice that coreference chains are relatively short, with 3 or 4 being the
most typical lengths, so it is unlikely to be a big issue.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Baselines</title>
        <p>
          Our baseline model adopts an end-to-end architecture for coreference resolution [
          <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
          ], as
depicted in Figure 5. Following the methods presented in [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], we use GloVe embeddings and a
character-level CNN as input to a BiLSTM to obtain contextualized word representations. Then
all possible spans are enumerated and fed to a mention classifier which detects if the input is a
mention. Based on the same mention representations, pairs of mentions are fed to a coreference
classifier and a bridging classifier, where the coreference classifier does binary classification
and the bridging one classifies pairs into 4 bridging relation types and a special class for no
relation. Training is done jointly with all losses added together.
        </p>
        <p>
          We released the code for training our baseline models to help the participants to get started
on the shared task.6 Two variants of the baseline model are evaluated on the test set, one using
the ELMO embeddings as input to the BiLSTM component, while the other used pretrained
ChELMO, based on the embeddings of [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] pre-trained on chemical patents, with the hope of
benefiting more from domain-specific pretraining.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results and Discussions</title>
      <p>A total of 19 teams registered on our submission website for the shared task. Among them, we
ifnally received 2 submissions on the test set. One team is from Carnegie Mellon University,
US (CMU) and the other one is from Hokkaido University, Japan (HUKB). More details about
their systems are provided in Section 7. In this section, we report their results along with the
performance of our two baseline systems.</p>
      <p>We report the overall performance of all runs in Table 2. The rankings of diferent systems
are fully consistent across all metrics. The CMU team achieves an 1 score of 0.7847 in exact
matching, outperforming our two baselines which get 0.7633 and 0.7453, followed by the HUKB
team who obtains 0.6907. The lead of the CMU team is even larger in relaxed matching, with an
1 score of 0.8723, about 7 points higher than our baselines. This shows the potential of the
CMU model and indicates that the performance in exact matching may be further boosted if
the boundary errors of their model could be corrected in a post-processing step.</p>
      <p>Our baselines have higher precision in the exact setting and precision in relaxed setting is
also very close to the best, which indicates that our models are more conservative and could
possibly be enhanced by making more aggressive predictions to improve recall. The use of
domain-pretrained embeddings (ChELMO vs. ELMO) does, as expected, benefit performance.</p>
      <p>Table 3 provides more details about the performance of all models for each relation type.
The CMU team outperforms others on TRANSFORMED relation by a large margin. While
our baselines performs the best on CONTAINED relation type. For the other three relation
types, the CMU model wins 1 score and recall, while our models achieve the highest precision,
which is similar to our observation on the overall results. Given that the models perform very
diferently, it would be very interesting to do more analysis when the details of all the models</p>
      <sec id="sec-6-1">
        <title>6Code available at https://github.com/biaoyanf/ChEMU-Ref</title>
        <p>are disclosed, and hopefully every team can borrow ideas from others and further improve the
performance.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Overview of Participants’ Approaches</title>
      <p>We received paper submissions from both the participating teams, i.e. the HUKB team and the
CMU team. We first describe their approaches, then summarize the same and diferent aspects
of them.
7.1. HUKB
The HUKB team used a two-step approach for the anaphora resolution task, where mentions
including both antecedent and anaphor are first detected, then mentions are classified into
diferent types and relations between them are determined. For step one, they found that
although existing parsers such as Chemical Tagger can generate useful features, they are not
enough for mention detection. Therefore they trained a BioBERT model to find candidate
mentions. This is done by treating mention detection as a NER task, where a BIOHD format is
used to convert gold spans into sequence of labels. The BIOHD format is an extension to the
well-known BIO format to support discontinuous spans which exist in this amphora resolution
task. In the second step, diferent types of relations are determined based on diferent rules.
COREFERENCE is first detected by 5 regular expression rules. The remaining relations are
detected based on the features generated by ChemicalTagger. When no more relations can be
found, a post-processing step is carried out to handle the transitivity property of COREFERENCE
relations, i.e. enumerating all valid COREFERENCE relations based on the transitivity property.
7.2. CMU
The CMU team proposed a pipelined system for anaphora resolution, where mentions are first
extracted and then relations between them are determined. The first step is done by a BERT-CRF
model which is trained using BIO tagging. In the second step, for each pair of mentions, the
sequence of sentences that contain the pair is fed to a BERT model to obtain the encoded
representation of every token, then the representation of a mention is simply the mean of all
tokens in it, and finally the representations of two mentions are concatenated and classified into
6 classes using a linear layer (5 relations + a class for no relation). To correct boundary errors
in mention detection, a rule-based post processing is done between step 1 and 2. Furthermore,
ensembling of 5 models are used in both step 1 and 2 to improve performance.</p>
      <sec id="sec-7-1">
        <title>7.3. Summary</title>
        <p>Both of them adopted a two-step approach where mentions are first detected and then relations
between them are determined. They also both relied on BERT-like models to extract
contextualized representations for mention detection. While the CMU team used a BERT-like model
in the relation extraction, the HUKB team chose a rule-based method. In addition, the CMU
team used ensembling of 5 models with majority voting in both mention detection and relation
extraction, as well as a post-processing step between them to correct potential boundary errors
in mention detection, where both techniques contribute to their superior overall performance.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusions</title>
      <p>This paper presents a general overview of the activities and outcomes of the ChEMU 2021
evaluation lab. As the second instance of our ChEMU lab series, ChEMU 2021 targets two new
tasks focusing on reference resolution in chemical patents. Our first task aims at identification
of reference relationships between chemical reaction descriptions, and our second task aims
at identification of reference relationships between expressions in chemical reactions. The
evaluation result includes diferent approaches to tackling the shared task, with one submission
clearly outperforming our baseline methods. We look forward to fruitful discussion and deeper
understanding of the methodological details of these submissions at the workshop.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>Funding for the ChEMU project is provided by an Australian Research Council Linkage Project,
project number LP160101469, and Elsevier. We acknowledge the support of our ChEMU-Ref
annotators, Dr. Sacha Novakovic and Colleen Hui Shiuan Yeow at the University of Melbourne,
and the annotation teams supporting the reaction reference task annotation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schwörer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Toomey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nau</surname>
          </string-name>
          , G. Ilchmann,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sheehan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Irmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bobach</surname>
          </string-name>
          , et al.,
          <article-title>Automatic identification of relevant chemical compounds from patents</article-title>
          ,
          <year>Database 2019</year>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bregonje</surname>
          </string-name>
          ,
          <article-title>Patents: A unique source for scientific technical information in chemistry related industry?</article-title>
          ,
          <source>World Patent Information</source>
          <volume>27</volume>
          (
          <year>2005</year>
          )
          <fpage>309</fpage>
          -
          <lpage>315</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Senger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bartek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Papadatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gaulton</surname>
          </string-name>
          ,
          <article-title>Managing expectations: Assessment of chemistry databases generated by automated extraction of chemical structures from patents</article-title>
          ,
          <source>Journal of Cheminformatics</source>
          <volume>7</volume>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cinciruk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Walsh</surname>
          </string-name>
          ,
          <article-title>Improving automated patent claim parsing: Dataset, system, and experiments</article-title>
          ,
          <source>arXiv preprint arXiv:1605.01744</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Muresan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Petrov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Southan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Kjellberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kogej</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tyrchan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Varkonyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>Making every SAR point count: The development of Chemistry Connect for the large-scale integration of structure and bioactivity data</article-title>
          ,
          <source>Drug Discovery Today</source>
          <volume>16</volume>
          (
          <year>2011</year>
          )
          <fpage>1019</fpage>
          -
          <lpage>1030</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoessel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yoshikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Albahem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cavedon</surname>
          </string-name>
          , T. Cohn,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verspoor</surname>
          </string-name>
          ,
          <year>Chemu 2020</year>
          :
          <article-title>Natural language processing methods are efective for information extraction from chemical patents</article-title>
          ,
          <source>Frontiers Res. Metrics Anal</source>
          .
          <volume>6</volume>
          (
          <year>2021</year>
          )
          <article-title>654438</article-title>
          . URL: https://doi.org/ 10.3389/frma.
          <year>2021</year>
          .
          <volume>654438</volume>
          . doi:
          <volume>10</volume>
          .3389/frma.
          <year>2021</year>
          .
          <volume>654438</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Leitner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Rabal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vazquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oyarzabal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Valencia</surname>
          </string-name>
          ,
          <string-name>
            <surname>CHEMDNER:</surname>
          </string-name>
          <article-title>The drugs and chemical names extraction challenge</article-title>
          ,
          <source>Journal of Cheminformatics</source>
          <volume>7</volume>
          (
          <year>2015</year>
          )
          <article-title>S1</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoessel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yoshikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Albahem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cavedon</surname>
          </string-name>
          , T. Cohn,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verspoor</surname>
          </string-name>
          , Overview of ChEMU 2020:
          <article-title>Named entity recognition and event extraction of chemical reactions from patents, in: Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Eleventh International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ), volume
          <volume>12260</volume>
          , Lecture Notes in Computer Science,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yoshikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoessel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          , T. Cohn,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          , et al.,
          <article-title>ChEMU: Named entity recognition and event extraction of chemical reactions from patents</article-title>
          ,
          <source>in: European Conference on Information Retrieval</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>572</fpage>
          -
          <lpage>579</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>W. A.</given-names>
            <surname>Baumgartner Jr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Ciosici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hailu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Pielke-Lombardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Regan</surname>
          </string-name>
          , L. Hunter,
          <article-title>CRAFT shared tasks 2019 overview-integrated structure, semantics, and coreference</article-title>
          ,
          <source>in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>174</fpage>
          -
          <lpage>184</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-D. Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Tsujii</surname>
          </string-name>
          ,
          <article-title>Overview of BioNLP 2011 protein coreference shared task</article-title>
          ,
          <source>in: Proceedings of BioNLP Shared Task 2011 Workshop</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ohta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tateisi</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-D. Kim</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mima</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Tsujii</surname>
          </string-name>
          ,
          <article-title>The GENIA corpus: An annotated research abstract corpus in molecular biology domain</article-title>
          ,
          <source>in: Proceedings of the Second International Conference on Human Language Technology Research</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>82</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>K. B. Cohen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lanfranchi</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bada</surname>
            ,
            <given-names>W. A. B.</given-names>
          </string-name>
          <string-name>
            <surname>Jr.</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Panteleyeva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Verspoor</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>L. E.</given-names>
          </string-name>
          <string-name>
            <surname>Hunter</surname>
          </string-name>
          ,
          <article-title>Coreference annotation and resolution in the colorado richly annotated full text (CRAFT) corpus of biomedical journal articles</article-title>
          ,
          <source>BMC Bioinform</source>
          .
          <volume>18</volume>
          (
          <year>2017</year>
          )
          <volume>372</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>372</lpage>
          :
          <fpage>14</fpage>
          . URL: https://doi.org/10.1186/s12859-017-1775-9. doi:
          <volume>10</volume>
          .1186/ s12859-017-1775-9.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shipley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sitnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Baumgartner</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. A.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. B.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verspoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Blake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Hunter</surname>
          </string-name>
          ,
          <article-title>Concept annotation in the CRAFT corpus</article-title>
          ,
          <source>BMC Bioinformatics 13</source>
          (
          <year>2012</year>
          )
          <article-title>161</article-title>
          . URL: https://www.ncbi.nlm.nih.gov/pubmed/22776079. doi:
          <volume>10</volume>
          .1186/
          <fpage>1471</fpage>
          -2105-13-161.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Mayer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Trippe</surname>
          </string-name>
          ,
          <article-title>Current challenges in patent information retrieval</article-title>
          , volume
          <volume>37</volume>
          , Springer,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yeow Hui Shiuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Novakovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hössel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mistica</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verspoor</surname>
          </string-name>
          ,
          <article-title>Chemu-ref dataset for modeling anaphora resolution in the chemical domain</article-title>
          ,
          <year>2021</year>
          . doi:
          <volume>10</volume>
          .17632/r28xxr6p92.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>K.</given-names>
            <surname>Verspoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoessel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Zhai,</surname>
          </string-name>
          <article-title>ChEMU dataset for information extraction from chemical patents</article-title>
          ,
          <year>2020</year>
          . doi:
          <volume>10</volume>
          . 17632/wy6745bjfj.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoessel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yoshikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Albahem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cavedon</surname>
          </string-name>
          , T. Cohn,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verspoor</surname>
          </string-name>
          , Overview of chemu 2020:
          <article-title>Named entity recognition and event extraction of chemical reactions from patents</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickhof</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Névéol</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality</source>
          , Multimodality, and Interaction - 11th
          <source>International Conference of the CLEF Association, CLEF</source>
          <year>2020</year>
          , Thessaloniki, Greece,
          <source>September 22-25</source>
          ,
          <year>2020</year>
          , Proceedings, volume
          <volume>12260</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2020</year>
          , pp.
          <fpage>237</fpage>
          -
          <lpage>254</lpage>
          . URL: https: //doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -58219-7_
          <fpage>18</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -58219-7\_
          <fpage>18</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pradhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moschitti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Xue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Uryupina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Conll
          <article-title>-2012 shared task: Modeling multilingual unrestricted coreference in ontonotes</article-title>
          , in: S. Pradhan,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moschitti</surname>
          </string-name>
          , N. Xue (Eds.),
          <source>Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning - Proceedings of the Shared Task: Modeling Multilingual Unrestricted Coreference in OntoNotes</source>
          , EMNLP-CoNLL
          <year>2012</year>
          , July 13,
          <year>2012</year>
          ,
          <string-name>
            <given-names>Jeju</given-names>
            <surname>Island</surname>
          </string-name>
          , Korea,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          ,
          <year>2012</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          . URL: https://www.aclweb.org/anthology/W12-4501/.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ghaddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Langlais</surname>
          </string-name>
          ,
          <string-name>
            <surname>Wikicoref:</surname>
          </string-name>
          <article-title>An english coreference-annotated corpus of wikipedia articles</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Goggi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Grobelnik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Maegaard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mazo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Odijk</surname>
          </string-name>
          , S. Piperidis (Eds.),
          <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation LREC</source>
          <year>2016</year>
          , Portorož, Slovenia, May
          <volume>23</volume>
          -28,
          <year>2016</year>
          ,
          <string-name>
            <given-names>European</given-names>
            <surname>Language Resources Association</surname>
          </string-name>
          (ELRA),
          <year>2016</year>
          . URL: http://www.lrec-conf.org/proceedings/lrec2016/summaries/192.html.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <article-title>Machine learning for entity coreference resolution: A retrospective look at two decades of research</article-title>
          , in: S.
          <string-name>
            <given-names>P.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          Markovitch (Eds.),
          <source>Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9</source>
          ,
          <year>2017</year>
          , San Francisco, California, USA, AAAI Press,
          <year>2017</year>
          , pp.
          <fpage>4877</fpage>
          -
          <lpage>4884</lpage>
          . URL: http://aaai.org/ocs/index.php/AAAI/AAAI17/ paper/view/14995.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>Entity-centric coreference resolution with model stacking, in: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th</article-title>
          <source>International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, ACL 2015, July 26-31</source>
          ,
          <year>2015</year>
          , Beijing, China, Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers</given-names>
          </string-name>
          , The Association for Computer Linguistics,
          <year>2015</year>
          , pp.
          <fpage>1405</fpage>
          -
          <lpage>1415</lpage>
          . URL: https://doi.org/10.3115/v1/p15-
          <fpage>1136</fpage>
          . doi:
          <volume>10</volume>
          .3115/v1/p15-
          <fpage>1136</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>B.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          , K. Verspoor,
          <string-name>
            <surname>ChEMU-Ref</surname>
          </string-name>
          :
          <article-title>A corpus for modeling anaphora resolution in the chemical domain</article-title>
          ,
          <source>in: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, Association for Computational Linguistics</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>K.</given-names>
            <surname>Krippendorf</surname>
          </string-name>
          ,
          <article-title>Measuring the reliability of qualitative text analysis data</article-title>
          ,
          <source>Quality and quantity 38</source>
          (
          <year>2004</year>
          )
          <fpage>787</fpage>
          -
          <lpage>800</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <article-title>End-to-end neural coreference resolution</article-title>
          , in: M.
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hwa</surname>
          </string-name>
          , S. Riedel (Eds.),
          <source>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP</source>
          <year>2017</year>
          , Copenhagen, Denmark, September 9-
          <issue>11</issue>
          ,
          <year>2017</year>
          , Association for Computational Linguistics,
          <year>2017</year>
          , pp.
          <fpage>188</fpage>
          -
          <lpage>197</lpage>
          . URL: https: //doi.org/10.18653/v1/d17-
          <fpage>1018</fpage>
          . doi:
          <volume>10</volume>
          .18653/v1/d17-
          <fpage>1018</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <article-title>Higher-order coreference resolution with coarse-to-fine inference, in: M. A</article-title>
          .
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Stent (Eds.),
          <source>Proceedings of the</source>
          <year>2018</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-</article-title>
          <string-name>
            <surname>HLT</surname>
          </string-name>
          , New Orleans, Louisiana, USA, June 1-6,
          <year>2018</year>
          , Volume
          <volume>2</volume>
          (
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>687</fpage>
          -
          <lpage>692</lpage>
          . URL: https://doi.org/10.18653/v1/n18-
          <fpage>2108</fpage>
          . doi:
          <volume>10</volume>
          .18653/v1/n18-
          <fpage>2108</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Akhondi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Druckenbrodt</surname>
          </string-name>
          , T. Cohn,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gregory</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verspoor</surname>
          </string-name>
          ,
          <article-title>Improving chemical named entity recognition in patents with contextualized word embeddings</article-title>
          ,
          <source>in: Proceedings of the 18th BioNLP Workshop</source>
          and Shared Task, Association for Computational Linguistics, Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>328</fpage>
          -
          <lpage>338</lpage>
          . URL: https: //www.aclweb.org/anthology/W19-5035. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W19</fpage>
          -5035.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>