<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Self-Supervised Learning for Visual Sum mary Identification in Scientific Publications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shintaro Yamamoto</string-name>
          <email>s.yamamoto@fuji.waseda.jp</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anne Lauscher</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simone Paolo Ponzetto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Goran Glavaš</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shigeo Morishima</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data and Web Science Group, University of Mannheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Pure and Applied Physics, Waseda University</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Waseda Research Institute for Science and Engineering</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>5</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>Providing visual summaries of scientific publications can increase information access for readers and thereby help deal with the exponential growth in the number of scientific publications. Nonetheless, eforts in providing visual publication summaries have been few and far apart, primarily focusing on the biomedical domain. This is primarily because of the limited availability of annotated gold standards, which hampers the application of robust and high-performing supervised learning techniques. To address these problems we create a new benchmark dataset for selecting figures to serve as visual summaries of publications based on their abstracts, covering several domains in computer science. Moreover, we develop a self-supervised learning approach, based on heuristic matching of inline references to figures with figure captions. Experiments in both biomedical and computer science domains show that our model is able to outperform the state of the art despite being self-supervised and therefore not relying on any annotated training data.</p>
      </abstract>
      <kwd-group>
        <kwd>scientific publication mining</kwd>
        <kwd>multimodal retrieval</kwd>
        <kwd>visual summary identification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Given the exponential growth in the number of scientific publications [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], providing concise
summaries of scientific literature becomes increasingly important. Accordingly, previous
work has focused on the automatic creation of textual summaries [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5 ref6 ref7">2, 3, 4, 5, 6, 7</xref>
        ]. However,
specifically in the case of scientific publications (and especially in some domains), information
is also conveyed in the form of figures, which allow the reader to understand the scientific
contributions better, ofering visual representations of data, experimental design, and results.
Some scientific publishing companies (e.g., Elsevier) even require authors to submit a figure
as a Graphical Abstract (GA), which is “a single, concise, pictorial and visual summary of the
main findings of the article” 1. GAs, in turn, are then used to provide multi-modal online search
results, following the observations that humans better remember and recall visual information
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
online
© 2021 Copyright for this paper by its authors.
      </p>
      <p>CEUR
Workshop
Proceedings
htp:/ceur-ws.org
ISN1613-073</p>
      <p>CEUR Workshop Proceedings (CEUR-WS.org)
1https://www.elsevier.com/authors/journal-authors/graphical-abstract</p>
      <p>
        Recently, Yang et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] introduced the concept of the central figure, referring to the figure
that is the best candidate for a GA of a paper. To build a dataset for automatically finding
central figures in scientific publications, they asked authors of papers in PubMed 2 to identify
one central figure in each of their scientific publications. Though a GA is not required for all
publications, authors were shown to be able to identify the central figure from their publications
for 87.6% of the papers. Using the obtained datasets of publications with GAs, they devised a
supervised machine learning approach for identifying central figures of scientific publications.
Such approaches for automatically identifying central figures can be employed to create GAs for
large collections of scientific documents. As a result, researchers and students can profit from the
visual support in concrete scenarios: for instance, as they obtain an impression of the discussed
research at a glance without reading the text, more eficient analysis of online search results is
possible. In this paper, we address two major limitations of Yang et al.’s seminal contribution.
First, the dataset of Yang et al. consists of PubMed data only, limiting the applicability of the
devised supervised central figure identification model to the biomedical domain. The use of
ifgures in scientific literature, however, is a common practice in a much broader set of research
ifelds and areas [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Secondly, while supervised learning is known to generally provide the
best results, it critically depends on (suficiently) large amounts of labeled data to be used for
training the models: expensive and time-consuming data annotation processes impede the
scalability of central figure identification across the plethora of research domains in which
ifgures encode valuable information. Whereas in some tasks, labeled data can be acquired more
economically with crowd-sourcing, this is not the case for the task at hand: identification of
the central figure for a publication requires annotators to be knowledgeable in the publication
domain. In other words, collecting datasets large enough to support supervised learning for
central figure identification for a wide range of many domains is impractical (if not infeasible)
due to the high annotation costs stemming from having to recruit expert annotators.
      </p>
      <p>
        To alleviate these issues, we propose (1) a novel benchmark for central figure identification
covering several subareas of computer science, and (2) a self-supervised learning approach for
which we do not need any labeled training data. For our proposed benchmark for central figure
identification, we ask two (semi-expert) annotators to rank the top three figures in a scientific
paper that would be the best candidates for a graphical abstract. The papers are collected
from four computer science subdomains: natural language processing (NLP), computer vision
(CV), artificial intelligence (AI), and machine learning (ML). Accordingly, our newly collected
dataset allows for a comparison of the performance of central figure identification models
across diverse (sub)domains. Secondly, to eliminate the reliance on labeled training data, we
introduce a self-supervised learning approach for automatic identification of central figures in
scientific publications. The core idea of our approach is outlined as follows. In most scientific
publications, a figure is mentioned in an article’s body by using a direct reference (e.g., “In Figure
3, we illustrate ⋯”). This typically means that the paragraph of the direct link sentence roughly
describes in text what the figure depicts visually, i.e., that the paragraph’s content is clearly
associated with the content of the figure. We exploit these direct links between an article’s
body and use these paragraph-figure pairs as training instances for a supervised central figure
identification model. We then train several Transformer-based [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] models, which take pairs of
body text and figure captions as input, and we train the models to judge whether the paragraph
text (from the body of the article) matches its paired figure. At inference (i.e., test) time, we
rank the article’s figures by (1) predicting the scores for each article figure by pairing them all
with the abstract and feeding them to the model and (2) ranking the figures based on the scores
output by the model reflecting their degree of match with the article’s abstract. In contrast
to sentence matching approaches [
        <xref ref-type="bibr" rid="ref12 ref13 ref14 ref15">12, 13, 14, 15</xref>
        ], which perform sentence-pair classification,
we tackle a ranking problem, scoring and ordering all figures of an article given its abstract.
Although self-supervised, our approach outperforms the existing fully supervised learning
approach for central figure identification [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] in terms of top-1 accuracy. Finally, we provide an
extensive analysis of performance diferences across diferent domains.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        While the majority of related work in scientific paper summarization has focused on
automatically creating textual summaries [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5 ref6 ref7">2, 3, 4, 5, 6, 7</xref>
        ], only a few have investigated the creation of
visual summaries, i.e., selection of images that best reflect the publication content. Kuzi and Zhai
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] proposed Keyword-based figure retrieval: they tackle the related problem of ranking figures
from multiple papers (in ACL anthology reference corpus [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]). The task that we tackle in this
work difers in that we focus on selecting the best figure for a single publication, considering
only the figures from that publication as candidates. Similarly, in [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ] the authors rank
ifgures from a single paper based on their importance.
      </p>
      <p>
        In this paper, we consider the problem of automatically identifying a central figure for a
paper, which would then be a candidate for the paper’s visual summary, referred to as Graphical
Abstract (GA) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Several works have focused on analyzing GAs, e.g., their use [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and design
pattern [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        The automatic selection of a central figure for scientific papers was first proposed by Yang
et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In their work, they built a dataset for the central figure identification from PubMed
(biomedical and life science) papers. To extend the study of central figure identification, we
propose a novel dataset consisting of computer science papers from several subdomains. Yang et
al. proposed a supervised learning approach for central figure identification, a methodology that
can hardly scale across a variety of scientific disciplines, due to the need for expert annotation
of central figures. Limited sizes of existing datasets for various tasks in scientific publication
mining [
        <xref ref-type="bibr" rid="ref22 ref23 ref7 ref9">9, 22, 7, 23</xref>
        ], additionally suggest that obtaining any kind of gold expert annotations on
scientific text is expensive and time-consuming. To remedy for this bottleneck of annotation
cost, we propose a self-supervised approach in which we make use of direct inline figure
references in the article body to heuristically pair article paragraphs with figure captions and
use those pairs as distant supervision.
      </p>
      <p>
        The similarity between an abstract and a figure caption is the most important feature for the
supervised model of [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Accordingly, we treat the task of identifying a central figure as an
abstract-to-caption matching problem [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Approaches for sentence matching can be divided
into two types: a sentence encoding-based approach and an attention-based approach. In the
sentence encoding approach, sentences are encoded separately [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], which, in contrast to the
attention-based approach [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">15, 14, 13</xref>
        ], does not capture semantic interactions between them.
We employ an attention-based approach and build the model on top of pretrained Transformer
networks [
        <xref ref-type="bibr" rid="ref24 ref25">24, 25</xref>
        ]. In contrast to current research on sentence matching as a classification task,
we treat central figure identification as a ranking problem where all figures in a paper are scored
according to their suitability to be used as a central figure.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Annotation Study</title>
      <p>
        Data Collection. According to [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the use of figures in scientific literature difers according
to the field and the topic of research. To investigate fine-grained diferences for automatic
central figure identification across research domains, we collect papers published between
2017 and 2019 for four diferent research fields in computer science, namely natural language
processing (NLP), computer vision (CV), artificial intelligence (AI) and machine learning (ML).
In order to make the dataset suficiently challenging, we keep only the publications with more
than five figures. Table 1 provides the dataset statistics (number of publications and average
number of figures per publication for each subdomain).
      </p>
      <p>Annotation Process. Our annotation task is defined as follows: given a paper abstract and
the figures extracted from the paper, identify and rank the top 3 figures according to the degree to
which they match the abstract and can therefore serve as a visual summary. In our annotation
guidelines we adopt the definition of a graphical abstract (GA) as given in the Elsevier author
guidelines (cf. footnote 1). Annotations were carried out by two coders with a university degree
in computer science, who were instructed to study the examples provided on the publisher page
and discuss them in a group to make sure they understood the notion of a graphical abstract.</p>
      <p>To facilitate the annotation process, we develop a web-based annotation tool with a graphical
user interface displaying a paper abstract and all figures extracted from the same paper, which
are randomly shufled to avoid the bias induced by the order. We first asked our annotators
to read the abstract in order to obtain an overview of the paper and then to study each figure
carefully. Next, the annotators were asked to choose and rank the top 3 GA candidates. All
instances are either doubly or singly annotated, and the inter-annotator agreement across the
doubly annotated data amounts to .43 Krippendorf’s  (ordinal), which reflects the dificulty
and the subjective nature of the task.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>
        Problem Definition. Central figure identification can be defined in two diferent ways [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
namely figure-level and the paper-level: in the figure-level setting, each individual figure is
classified as being a central figure or not, while in the paper-level setting, a central figure is
determined from all figures in a paper. In this work, we are primarily interested in retrieving
GAs as a form of summarization: hence, we opt for a document-level approach and cast it a
ranking problem in which all figures from a paper are to be scored based on their suitability to
provide a central figure for the publication.
      </p>
      <p>
        Building on the result from [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] that the similarity between an abstract and a figure caption is
the most important factor for central figure identification, we use a pair of abstract and figure
caption as input. Given the sets of figures extracted from a paper  = {  ∶ } and an abstract  ,
we learn a scoring function  (,  ) that predicts the appropriateness of the figure to act as the
central figure for the abstract (and accordingly, the paper). All figures are then ranked according
to the model’s prediction  = {  ∶   =  (  ,  )} .
      </p>
      <p>
        Model. Our model consists of two components, a Transformer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] as a language encoder and
a score predictor (Figure 1).
      </p>
      <p>
        We build upon finding from recent work in NLP that has shown the benefits of an
attentionbased approach for sentence matching [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">14, 15, 13</xref>
        ], and accordingly opt for a pre-trained BERT
[
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] model as the text encoder: specifically, we use in our experiments a SciBERT model
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], which is pretrained on scientific publications from Semantic Scholar [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. We provide
a text pair consisting from an abstract and a figure caption 3 as input to BERT, augmented
with Transformer’s special tokens: “[ C L S ] a b s t r a c t [ S E P ] c a p t i o n [ S E P ] ”. The transformed
hidden vector of the sequence start token [ C L S ] token, xCLS , is then forwarded into a linear
transformation layer that produces the final relevance score:  = xCLS W +  , with the vector
W ∈ ℝ and scalar  ∈ ℝ as regressor’s parameters ( = 768 is BERT’s hidden state size). For
BERT, the length of the input sequence is restricted to be up to a maximum of 512 tokens. We
considered increasing the input sequence length for handling longer abstracts, but we decided
against this option as it would require training instances with longer sequences and as it would
result in a non-negligible increase of the required GPU memory. To overcome this limitation
and allow for abstracts of arbitrary sizes, we divide an abstract into sentences and aggregate
scores across sentences. Given a function  to score pairs of sentences (from the abstract) and
ifgure captions (  ), and a set of sentences in an abstract  = {  ∶ } , the scoring function is
defined as  (,  ) =
      </p>
      <p>∑ (,   ).</p>
      <p>Training Instance Creation. The annotation of scientific publications requires expert
knowledge of the field of research. To avoid manual annotation of the training data, we introduce
a self-supervised approach by leveraging explicit inline references to figures (e.g., “Figure 2
depicts the results of the ablation experiments…”). In a scientific publication, an inline reference
to a figure indicates that the paragraph and the figure are related to each other. We denote the
 -th paragraph that mentions the figure   and the set of paragraphs referring to figures in a
paper as   and  = {  ∶ } , respectively. Instead of directly identifying a central figure during

training, we learn the matching of the figure  and the paragraph  . At training time, we make
positive and negative pairs of paragraphs and figures as (  ,   ), as shown in Figure 2. We treat
3We feed the abstract sentences as input only at inference time. In training, input instances couple the paragraphs
from the article’s body explicitly mentioning the figure with the figure caption.
the pair (  ,   ) as positive if  = , while  ≠  for a negative one.</p>
      <p>
        Optimization. We train the model to rank the positive pairs higher than negative ones. Due
to BERT’s input token sequence length restriction, we randomly sample one sentence from
the paragraph. The pair of a sampled sentence and a caption is fed into the model as ”[CLS]
sentence [SEP] caption [SEP]”. For the training objective, we formulate the following loss
similar to the Triplet loss [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] as  = max(  −   + , 0) , where   and   denote the predictions
of the model for the positive and negative pairs, respectively. In the experiments, we set  = 1.0 .
For a single training instance, we sample one positive and one negative pairs including the
ifgure   as (  ,   ) and (  ,   ′) ( ≠  ′), respectively. The training objective makes the score for a
positive pair lower than that for a negative one: therefore, the figure with the lower score is
ranked higher.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <sec id="sec-5-1">
        <title>5.1. Implementation Details</title>
        <p>
          We conduct our experiments using BERT’s implementation from the Hugging Face library [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ].
In all fine-tuning procedures, we use the Adam optimizer [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] with the learning rate 1 − 6 , train
in batches of size 32, and apply a dropout at the rate of 0.2 and a gradient clipping threshold of 5.
We train the model for 1 epoch. To extract text from collected PDF versions of papers, we rely
on the Science Parse library4. Explicit inline references of figure are identified via the keywords
”Figure” or ”Fig.”. To extract figure captions, we employ the image-based approach from [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ].
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Experimental Setting</title>
        <p>
          PubMed. In [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], 7, 295 biomedical and life science papers from PubMed are annotated for
central figure identification. We managed to obtain the PDFs from PubMed for 7, 113 of those
papers and divide the papers into training, validation, and test portions in the same ratio as
Yang et al. (8:1:1). Using our figure mention heuristic, we create 40 paragraph-figure pairs
from the training portion of the dataset. As only a single figure is annotated as the central figure
in the PubMed dataset, we use the top-1 and top-3 accuracy as evaluation metrics following [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
Computer Science (CS). We additionally evaluate our model on our new CS dataset (Section 3).
Unlike the PubMed dataset, in which only a single figure is annotated as central, our annotators
ranked three figures for each CS paper. Consequently, we use Mean Average Precision (MAP),
Mean Reciprocal Rank (MRR), and normalized Discounted Cumulative Gain (nDCG) as our
evaluation metrics on the CS dataset. For the optimization procedure, we collect papers from
the same subdomains as the annotated test data, from between 2015 and 2018. We divide them
into training and validation portions with ratio of 90% and 10%, respectively. Here we also
obtain around 40 paragraph-figure instances for model training. For performance evaluation,
we utilize our annotated data described in Section 3.
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Experiments</title>
        <sec id="sec-5-3-1">
          <title>Model</title>
          <p>
            Random
Pick first
Text-only
Full
Vanilla BERT
RoBERTa
SciBERT
Performance on the PubMed dataset. We first evaluate our self-supervised approach using
the PubMed dataset (Table 2). We follow [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] and make use of two baselines: random and ‘select
ifrst image’. The random baseline ranks figures randomly and the ‘select first image’ ranks
based on the order of the figures as they appear in the paper (i.e., figure 1 is ranked 1st). For
comparison, we also provide the results of two methods from [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ], a text-only model that uses
cosine similarity of TF-IDF between the abstract and the figure caption as the input feature, and
a full model that takes the figure type label (e.g., diagram, plot) and layout (e.g., section index,
ifgure order) as inputs, as well as text features. We compare these against the the performance
of our models based on three diferent pretrained Transformers, namely vanilla BERT [
            <xref ref-type="bibr" rid="ref25">25</xref>
            ],
RoBERTa [
            <xref ref-type="bibr" rid="ref31">31</xref>
            ] and SciBERT [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ].
          </p>
          <p>
            Regardless of the text encoder, our approach outperforms the baselines in terms of both top-1
and top-3 accuracy. This result indicates that our method for generating training data creation
is efective for central figure identification. Among the text encoders, SciBERT performs the
best for both metrics, arguably because it has been trained on a corpus of scientific papers, thus
minimizing problems related to domain transfer. Despite not requiring manual annotation for
training, our approach with SciBERT also outperforms the supervised approach of Yang et al.
[
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] in terms of top-1 accuracy.
          </p>
          <p>
            Performance on the CS dataset. We also evaluate the performance of the model on our CS
dataset (Table 3). We follow the same setting as for the PubMed data and use a random and
‘choose first image’ methods as baselines. Here, we compare SciBERT with vanilla BERT and
RoBERTa, so as to additionally verify the efectiveness of SciBERT in the CS domain – since
over 80% of the papers in the corpus for SciBERT pre-training are from the biomedical domain
and the ratio of CS papers account for only 18% of SciBERT’s pretraining corpus [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ].
          </p>
          <p>Our approach outperforms the random baseline in terms of MAP, MRR, and nDCG. Among
base models, SciBERT outperforms vanilla BERT and RoBERTa as in PubMed papers. Though
most of the corpus for SciBERT pre-training is from the biomedical domain, a certain number
of CS papers seen in pretraining still contributes to the downstream performance on central
ifgure identification.</p>
          <p>However, as opposed to the case of PubMed papers, the ‘pick first’ baseline here is much
stronger and hard-to-beat, even for our Transformer-based approach. Note that in the proposed
self-supervised learning approach, the order of the figures is not taken into account.
Consequently, our models do not consider the order in which the figures appear. This result indicates
that CS papers tend to use Graphical Abstract (GA) in the beginning, and empirically highlights
that our new dataset is more challenging than the PubMed-based dataset, as the ‘pick first’
baseline is hard to beat.</p>
          <p>
            Cross-domain Experiments. Image usage in scientific publications is known to be diferent
across scientific fields [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]. Accordingly, we set next to empirically evaluate the robustness of
our approach in a domain transfer setup. Due to the diferent granularity of our PubMed and
CS datasets, the latter including papers from four diferent research areas of computer science
(AI, NLP, ML, and CV), we are able to perform two sets of domain transfer experiments, namely
biomedical vs. computer science as well as across diferent CS subdomains.
          </p>
          <p>We first compare model performance by training and testing on datasets from diferent
domains – i.e., biomedical papers from PubMed vs. computer science publications from our CS
dataset – using SciBERT as a base model (Table 4).</p>
          <p>In the test with PubMed papers, training on the same domain performs better both in terms of
top-1 and top-3 accuracy. Despite the slightly lower performance, training on the CS domain also
outperforms the random baseline. On the other hand, training on diferent domains, somewhat
surprisingly, does not degrade the performance. This would imply that papers from diferent
domains exhibit similar text-figure (caption) matching properties.</p>
          <p>Next, we examine the results of domain transfer for diferent CS subdomains dataset belonging
to diferent areas of computer science, due to the fact that image usage and volume may
potentially vary among fields like, e.g., natural language processing and computer vision, with
papers from the latter containing typically more images. We train models on four diferent
areas (NLP, CV, AI, and ML) and test them on all others. Domain comparison within several
research topics in CS is summarized in Table 5.</p>
          <p>Overall, the results are rather consistent across areas and indicate that, within computer
science, the research topics of papers do not afect the model performance. Among the four
topics, the performance is the lowest on machine learning (ML) papers. We therefore manually
analyzed ML papers with poor model performance and observed that these papers tend to have
ifgures that look rather similar. This makes the identification of the central figure – even with
manual efort – dificult. The ‘pick first’ scores higher for CV papers than for NLP and ML
papers, whereas the random baseline naturally performs worse on CV papers, which contain
more figures.</p>
          <p>
            Model Analysis. To understand the model behavior, we analyze the attention in SciBERT.
We visualize the attention in the Transformer model. We find that most attention maps are
consistent with typical classes reported in [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ], such as vertical or diagonal attention patterns.
In some attention heads, the model attends to the lexical overlap between abstract and caption.
The examples of attention matrices produced by heads attending over the same or semantically
similar tokens, are shown in Figure 3. In this example, instances of tokens like ’tracking’
and ’when’ appearing in both in abstract and caption have mutually high attention weights.
Additionally, pairs of tokens with similar meaning like ’restore’ and ’recovering’ also receive
high mutual attention weights.
          </p>
          <p>
            We also compare attention patterns among the model trained with diferent topics (NLP, CV,
AI, and ML). Following [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ], we calculate the cosine similarity of attention maps. We show the
mean cosine similarity of flattened attention map for randomly selected 100 samples in Table 6.
Cosine similarity is high for all combinations; this means that the attention patterns across the
diferent CS domains are virtually identical, confirming empirically our previous assumption
that there are no relevant diferences between CS domains when it comes to text-figure matching
(see Table 5).
          </p>
          <p>
            Kovaleva et al. reported that after fine-tuning attention maps change the most in the last two
transformer layers [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ]. We therefore analyze the change in attention patterns after our
taskspecific fine-tuning. We compare the standard fine-tuning, in which we update all SciBERT’s
parameters (and which we used in all our previous experiments), and the feature-based training,
in which we freeze SciBERT’s parameters and train only the regressor’s parameters. The
comparison of attention patterns between fine-tuned and frozen SciBERT is summarized in
Table 7. On the one hand, if we freeze SciBERT’s parameters, we observe a major drop in
performance (6 MAP points). On the other hand, high cosine similarity of attention maps
between the fine-tuned and frozen SciBERT that the two transformers still exhibit similar
attention patterns. This suggests that only slight changes in the parameters of Transformer’s
attention heads have the potential to substantially change the predictions of the regressor.
(b) Cosine similarity of attention map for randomly sampled 100 sentence-caption pairs in each layer.
          </p>
        </sec>
        <sec id="sec-5-3-2">
          <title>Layer</title>
          <p>Similarity</p>
          <p>Layer
Similarity</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>While research eforts have been mostly spent on increasing information access to scientific
literature by creating textual summaries, it is known that a large amount of information is often
conveyed visually in the form of figures. In this work, we have addressed the problem of central
ifgure identification from scientific publications, the task of identifying a candidate for a visual
summary. Starting from previous work, which has introduced a dataset enabling supervised
learning for central figure identification, we identified and addressed two main issues: (1) the
only existing data set is limited to the biomedical domain, and (2) large-scale annotations for
new domains are impractical and costly. To alleviate these issues, we first presented a new
benchmark collection of scientific publications annotated for central figures in the computer
science domain covering four diferent subfields. Secondly, we proposed a self-supervised
approach to central figure identification. Our method exploits the link between portions of
text explicitly referencing figures and figure captions, thereby bypassing the need for large
manually annotated training data. We have experimentally demonstrated the efectiveness of
our approach, outperforming the supervised approach in terms of rank-1 accuracy. Finally,
a follow-up analysis of cross-domain performance diferences and models’ attention scores
revealed only slight diferences across the individual CS subdomains, but interestingly, our
ifndings also indicate that the positioning of the central figure difers between the CS and the
biomedical domain. We hope that our results fuel further research on cost-efective visual
summary creation for increased information access in light of the exponentially growing body
of scientific literature.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was supported by the Program for Leading Graduate Schools, ”Graduate Program for
Embodiment Informatics” of the Ministry of Education, Culture, Sports, Science and Technology
(MEXT) of Japan, and JST ACCEL (JPMJAC1602). Computational resource of AI Bridging
Cloud Infrastructure (ABCI) provided by National Institute of Advanced Industrial Science and
Technology (AIST) was used.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bornmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mutz</surname>
          </string-name>
          ,
          <article-title>Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references</article-title>
          ,
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>66</volume>
          (
          <year>2015</year>
          )
          <fpage>2215</fpage>
          -
          <lpage>2222</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dernoncourt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goharian</surname>
          </string-name>
          ,
          <article-title>A discourseaware attention model for abstractive summarization of long documents</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>2</volume>
          (
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <year>2018</year>
          , pp.
          <fpage>615</fpage>
          -
          <lpage>621</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goharian</surname>
          </string-name>
          ,
          <article-title>Scientific article summarization using citation-context and article's discourse structure</article-title>
          ,
          <source>in: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>390</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Mei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <article-title>Generating impact-based summaries for scientific literature</article-title>
          ,
          <source>in: Proceedings of 46th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>816</fpage>
          -
          <lpage>824</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Qazvinian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Radev</surname>
          </string-name>
          ,
          <article-title>Scientific paper summarization using citation summary networks</article-title>
          ,
          <source>in: Proceedings of the 22nd International Conference on Computational Linguistics - Volume</source>
          <volume>1</volume>
          ,
          <year>2008</year>
          , pp.
          <fpage>689</fpage>
          -
          <lpage>696</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lauscher</surname>
          </string-name>
          , G. Glavaš, K. Eckert, University of mannheim@ clscisumm-
          <fpage>17</fpage>
          :
          <article-title>Citation-based summarization of scientific articles using semantic textual similarity</article-title>
          ,
          <source>in: CEUR workshop proceedings</source>
          , volume
          <year>2002</year>
          , RWTH,
          <year>2017</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yasunaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kasai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Fabbri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Radev</surname>
          </string-name>
          ,
          <article-title>Scisummnet: A large annotated corpus and content-impact models for scientific paper summarization with citation networks</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>7386</fpage>
          -
          <lpage>7393</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. S.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Walling</surname>
          </string-name>
          ,
          <article-title>Pictorial superiority efect</article-title>
          .,
          <source>Journal of experimental psychology: Human learning and memory 2</source>
          (
          <year>1976</year>
          )
          <fpage>523</fpage>
          -
          <lpage>528</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.-S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kazakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>West</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Howe</surname>
          </string-name>
          ,
          <article-title>Identifying the central figure of a scientific paper</article-title>
          ,
          <source>in: 2019 International Conference on Document Analysis and Recognition (ICDAR)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1063</fpage>
          -
          <lpage>1070</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.-S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>West</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Howe</surname>
          </string-name>
          , Viziometrics:
          <article-title>Analyzing visual information in the scientific literature</article-title>
          ,
          <source>IEEE Transactions on Big Data</source>
          <volume>4</volume>
          (
          <year>2018</year>
          )
          <fpage>117</fpage>
          -
          <lpage>129</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , L. u. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          ,
          <year>2017</year>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Bowman</surname>
          </string-name>
          , G. Angeli,
          <string-name>
            <given-names>C.</given-names>
            <surname>Potts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>A large annotated corpus for learning natural language inference</article-title>
          ,
          <source>in: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>632</fpage>
          -
          <lpage>642</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hamza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Florian</surname>
          </string-name>
          ,
          <article-title>Bilateral multi-perspective matching for natural language sentences</article-title>
          ,
          <source>in: Proceedings of the 26th International Joint Conference on Artificial Intelligence</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>4144</fpage>
          -
          <lpage>4150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Xu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Original semantics-oriented attention and deep fusion network for sentence matching</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2652</fpage>
          -
          <lpage>2661</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Attention-fused deep matching network for natural language inference</article-title>
          ,
          <source>in: Proceedings of the 27th International Joint Conference on Artificial Intelligence, IJCAI'18</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>4033</fpage>
          -
          <lpage>4040</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kuzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <article-title>Figure retrieval from collections of research articles</article-title>
          ,
          <source>in: European Conference on Information Retrieval</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>696</fpage>
          -
          <lpage>710</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Dorr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gibson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joseph</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Powley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Radev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. F.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>The ACL anthology reference corpus: A reference dataset for bibliographic research in computational linguistics</article-title>
          ,
          <source>in: Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>1755</fpage>
          -
          <lpage>1759</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Learning to rank figures within a biomedical article</article-title>
          ,
          <source>PLOS ONE 9</source>
          (
          <year>2014</year>
          )
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <article-title>Automatic figure ranking and user interfacing for intelligent ifgure search</article-title>
          ,
          <source>PLOS ONE 5</source>
          (
          <year>2010</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Chung,</surname>
          </string-name>
          <article-title>An investigation on graphical abstracts use in scholarly articles</article-title>
          ,
          <source>International Journal of Information Management</source>
          <volume>37</volume>
          (
          <year>2017</year>
          )
          <fpage>1371</fpage>
          -
          <lpage>1379</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hullman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bach</surname>
          </string-name>
          , Picturing science:
          <article-title>Design patterns in graphical abstracts</article-title>
          , in: P. Chapman,
          <string-name>
            <given-names>G.</given-names>
            <surname>Stapleton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moktefi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Perez-Kriz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bellucci</surname>
          </string-name>
          (Eds.),
          <source>Diagrammatic Representation and Inference</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>183</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lauscher</surname>
          </string-name>
          , G. Glavaš,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <article-title>Investigating the role of argumentation in the rhetorical analysis of scientific publications with neural multi-task learning models</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>3326</fpage>
          -
          <lpage>3338</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>X.</given-names>
            <surname>Hua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Badugu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Argument mining for understanding peer reviews</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <year>2019</year>
          , pp.
          <fpage>2131</fpage>
          -
          <lpage>2137</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Cohan,
          <article-title>SciBERT: A pretrained language model for scientific text</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>3615</fpage>
          -
          <lpage>3620</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>W.</given-names>
            <surname>Ammar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Groeneveld</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bhagavatula</surname>
          </string-name>
          , I. Beltagy,
          <string-name>
            <given-names>M.</given-names>
            <surname>Crawford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Downey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dunkelberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elgohary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Feldman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kinney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kohlmeier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Murray</surname>
          </string-name>
          , H.
          <string-name>
            <surname>-H. Ooi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Power</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Skjonsberg</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Wilhelm</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Yuan</surname>
          </string-name>
          , M. van
          <string-name>
            <surname>Zuylen</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <article-title>Construction of the literature graph in semantic scholar</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>3</volume>
          (
          <string-name>
            <surname>Industry</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <year>2018</year>
          , pp.
          <fpage>84</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ailon</surname>
          </string-name>
          ,
          <article-title>Deep metric learning using triplet network</article-title>
          , in: A.
          <string-name>
            <surname>Feragen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Pelillo</surname>
          </string-name>
          , M. Loog (Eds.),
          <source>Similarity-Based Pattern Recognition</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>84</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Debut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chaumond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delangue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cistac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Louf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Funtowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Brew</surname>
          </string-name>
          ,
          <article-title>Huggingface's transformers: State-of-the-art natural language processing</article-title>
          , ArXiv abs/
          <year>1910</year>
          .03771 (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          ,
          <source>arXiv preprint arXiv:1412.6980</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>N.</given-names>
            <surname>Siegel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lourie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Power</surname>
          </string-name>
          , W. Ammar,
          <article-title>Extracting scientific figures with distantly supervised neural networks</article-title>
          ,
          <source>in: Proceedings of the 18th ACM/IEEE on Joint Conference on Digital Libraries</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>223</fpage>
          -
          <lpage>232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          , ArXiv abs/
          <year>1907</year>
          .11692 (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>O.</given-names>
            <surname>Kovaleva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Romanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rogers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rumshisky</surname>
          </string-name>
          ,
          <article-title>Revealing the dark secrets of BERT</article-title>
          , in
          <source>: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4365</fpage>
          -
          <lpage>4374</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>