<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Computa-
[</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1145/3197026.3197040</article-id>
      <title-group>
        <article-title>ACL-Fig: A Dataset for Scientific Figure Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zeba Karishma</string-name>
          <email>zebakarishma@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ShauryaRohatg</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kavya ShrinivasPuranik</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jian Wu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>C. Lee Giles</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Scientific Figures, Figure Classification, ACL Anthology, ACL-Fig</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Old Dominion University</institution>
          ,
          <addr-line>5115 Hampton Blvd, Norfolk, VA 23529</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The Pennsylvania State University</institution>
          ,
          <addr-line>Westgate Building, University Park, PA 16802</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <volume>20</volume>
      <issue>1987</issue>
      <fpage>248</fpage>
      <lpage>255</lpage>
      <abstract>
        <p>Most existing large-scale academic search engines are built to retrieve text-based information. However, there are no largescale retrieval services for scientific figures and tables. One challenge for such services is understanding scientific figures' semantics, such as their types and purposes. A key ob- stacle is the need for datasets containing annotated scientific ifgures and tables, which can then be used for classification, question-answering, and auto-captioning. Here, we develop a pipeline that extracts figures and tables from the scientific lit- erature and a deep-learning-based framework that classifies scientific figures using visual features. Using this pipeline, we built the first large-scale automatically annotated corpus, ACL-FIG consisting of 112,052 scientific figures extracted from ≈ 56K research papers in the ACL Anthology. The ACL-FIGPILOT dataset contains 1,671 manually labeled scientific figures belonging to 19 categories. The dataset is ac- cessible at Workshop Proceedings Washington, DC'23:The AAAI-23 Workshop on Scientific Document</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Figures are ubiquitous in scientific papers illustrating
experimental and analytical results. We refer to these
ifgures as scientific figures</p>
      <p>to distinguish them from
natural images, which usually contain richer colors and
gradients. Scientific figures provide a compact way to
present numerical and categorical data, often facilitating
researchers in drawing insights and conclusions.
Machine understanding of scientific figures can assist in
developing efective retrieval systems from the hundreds
of millions of scientific papers readily available on the
Web [1]. The state-of-the-art machine learning models
can parse captions and shallow semantics for specific
categories of scientific figures. [2] However, the task of
reliably classifying general scientific figures based on
their visual features remains a challenge.
contextualized scientific figure datasets. Applying the
pipeline on 55,760 papers in the ACL Anthology
(downloaded from https://aclanthology.org/ in mid-2021), we
built two datasetAsC:L-Fig and ACL-Fig-pilot. ACL-Fig
references, and metadatAa.CL-Fig-pilot (Figure1) is a
subset of unlabeleAdCL-Fig, consisting of 1671 scientific
ifgures, which were manually labeled into 19 categories.</p>
      <p>Here, we propose a pipeline to build categorized anFdigure 1: Example figures of each type in ACL-Fig-pilot.
consists of 112,052 scientific figures, their captions, inline source and configurable, enabling others to expand the
for scientific figure classification. The pipeline is
opendatasets from other scholarly datasets with pre-defined
or new labels.</p>
      <p>CEUR</p>
      <p>ceur-ws.org</p>
      <p>The ACL-Fig-pilot dataset was used as a benchmark 2. Related</p>
    </sec>
    <sec id="sec-2">
      <title>Work</title>
      <p>CEUR
htp:/ceur-ws.org
ISN1613-073</p>
      <p>Attribution 4.0 International (CC BY 4.0).</p>
      <p>CEUR</p>
      <p>Workshop ProceedingsC(EUR-WS.org)</p>
      <sec id="sec-2-1">
        <title>Scientific Figures Extraction</title>
        <sec id="sec-2-1-1">
          <title>Automatically extract</title>
          <p>ing figures from scientific papers is essential for many
downstream tasks, and many frameworks have been
developed. A multi-entity extraction framework called</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>PDFMEF incorporating a figure extraction module was attention to compound figure detection and separation. © 2022 Copyright for this paper by its authors. Use permitted under Creative Commons Licpenrseoposed [3]. Shared tasks such as ImageCLEF4[] drew</title>
        </sec>
        <sec id="sec-2-1-3">
          <title>Clark and Divvala[5] proposed a framework callePdDF</title>
          <p>Figures that extracted figures and captions in research
papers. The authors extended their work and built a more
robust framework callePdDFFigures2 [6]. DeepFigures Figure
was later proposed to incorporate deep neural networkExtraction
models [2].</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Scientific Figure Classification Scientific figure clas</title>
        <p>sification [ 7, 8] aids machines in understanding figures.</p>
        <p>Early work used a visual bag-of-words representation
with a support vector machine classifier7][. Zhou and
Tan applied hough transforms to recognize bar charts inClustering
document images. Siegel et a[l1.0] used handcrafted
features to classify charts in scientific documents. Tang et al.
[11] combined convolutional neural networks (CNNs)
and the deep belief networks, which showed improved
performance compared with feature-based classifiers .</p>
        <p>DeepFigures PDFFigures2
Figure classification Datasets There are several ex- Automatic
isting datasets for figure classification such as DocFigure annotation Labeled figures with metadata
[12], FigureSeer1[0], Revision [7], and datasets presented
by Karthikeyani and Nagarajan[13] (Table1). FigureQA
is a public dataset that is similar to ours, consisting of
over one million question-answer pairs grounded in ovFeirgure 2: Overview of the data generation pipeline.
100,000 synthesized scientific images [14] with five styles.</p>
        <p>Our dataset is diferent from FigureQA because the
figures were directly extracted from research papers. Es3pe.- Data Mining Methodology
cially, the training dataDoefepFigures are from arXiv
and PubMed, labeled with only “figure” and “table”, andThe ACL Anthology is a sizable, well-maintained PDF
does not include fine-granular labels. Our dataset cocno-rpus with clean metadata covering papers in
computatains fine-granular labels, inline context, and is compiletdional linguistics with freely available full-text. Previous
from a diferent domain. work on figure classification used a set of pre-defined
categories (e.g.,1[4], which may only cover some figure
types. We use an unsupervised method to determine
ifgure categories to overcome this limitation. After the
category label is assigned, each figure is automatically
annotated with metadata, captions, and inline referencpeisv.ot point (elbow) of the curve determines the number
The pipeline includes 3 steps: figure extraction, clusteor-f clusters.
ing, and automatic annotation (Figu2r).e Silhouette Analysis determines the number of clusters
by measuring the distance between clusters. It considers
3.1. Figure Extraction multiple factors such as variance, skewness, and high-low
diferences and is usually preferred to the Elbow method.</p>
        <p>To mitigate the potential bias of a single figure extractTohre, Silhouette plot displays how close each point in one
we extracted figures usingpdffigures2 [6] and deep- cluster is to points in the neighboring clusters, allowing
figures [2] which work in diferent ways. PDFFigures2 us to assess the cluster number visually.
ifrst identifies captions and the body text because they
are identified relatively accurately. Regions containin3g.3. Linking Figures to Metadata
ifgures can then be located by identifying rectangular
bounding boxes adjacent to captions that do not overlTahpis module associates figures to metadata, including
with the body textD.eepFigures uses the distant super-captions, inline reference, figure type, figure boundary
vised learning method to induce labels of figures fro mcoordinates, caption boundary coordinates, and figure
a large collection of scientific documents in LaTeX andtext (text appearing on figures, only available for results
XML format. The model is based on TensorBox, applyingfromPDFFigures2). The figure type is determined in
the Overfeat detection architecture to image embedditnhges clustering step above. The inline references are
obgenerated using ResNet-1012][. We utilized the publiclytained using GROBID (see below). The other metadata
available model weigh1tstrained on 4M induced figures fields were output by figure extractorsP. DFFigures2
and 1M induced tables for extraction. The model ouantd- DeepFigures extract the same metadata fields
exputs the bounding boxes of figures and tables. Unlesscept for “image text” and “regionless captions” (captions
otherwise stated, we collectively refer to figures andfotar-which no figure regions were found), which are only
bles together as “figures”. We used multi-processing taovailable for resultsPoDf FFigures2.
process PDFs. Each process extracts figures following An inline reference is a text span that contains a
referthe steps below. The system processed, on average, 20e0nce to a figure or a table. Inline references can help to
papers per minute on a Linux server with 24 cores. understand the relationship between text and the objects
it refers to. After processing a paper, GROBID outputs a
1. Retrieve a paper identifier from the job queue. TEI file (a type of XML file), containing marked-up
full2. Pull the paper from the file system. text and references. We locate inline references using
3. Extract figures and captions from the paper. regular expressions and extract the sentences containing
4. Crop the figures out of the rendered PDFs using der-eference marks.</p>
        <p>tected bounding boxes.
5. Save cropped figures in PNG format and the metadata
in JSON format. 4. Results</p>
        <p>4.1. Figure Extraction
3.2. Clustering Methods
Next, we use an unsupervised method to label extracted
ifgures automatically. We extract visual features using
VGG16 [15], pretrained on ImageNet16[]. All input fig- 14283
ures are scaled to a dimension 2o2f4 × 224 to be
compatible with the input requirement of VGG16. The featurePsDFFigures2
were extracted from the second last hidden (dense) layer,
consisting of 4096 features. Principal Component AnalyFi-gure 3: Numbers of extracted images.
sis was adopted to reduce the dimension to 1000.</p>
        <p>Next, we cluster figures represented by the
1000dimension vectors using-means clustering. We com- The numbers of figures extracted bPyDFFigures2 and
pare two heuristic methods to determine the optiDmeaelpFigures are illustrated in Figu3r,ewhich indicates
number of clusters, including the Elbow method and tahseignificant overlap between figures extracted by two
Silhouette Analysi1s7[]. The Elbow method examines software packages. However, either package extracte≈d (
theexplained variation, a measure that quantifies the dif- 5%) figures that were not extracted by the other package.
ference between the between-group variance to the toBtyalinspecting a random sample of figures extracted by
variance, as a function of the number of clusters. Theeither software package, we found thatDeepFigures
tended to miss cases in which two figures were vertically
240623</p>
        <p>9046
DeepFigures
label count</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Supervised Scientific Figure</title>
    </sec>
    <sec id="sec-4">
      <title>Classification</title>
      <p>trees
natural images
confusion matrix</p>
      <p>graph
architecture diagram</p>
      <p>Screenshots
bar charts
neural networks
NLP text_grammar_eg</p>
      <p>Line graph_chart</p>
      <p>tables
algorithms
pie chart
scatter plot</p>
      <p>maps
boxplots
word cloud
venn diagram
pareto
4.2. Automatic Figure Annotation</p>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <p>The extraction outputs 151,900 tables and 112,052 figuresB. ased on the ACL Anthology papers, we designed a
Only the figures were clustered using t h-emeans algo- pipeline and used it to build a corpus of automatically
rithm. We varie d from 2 to 20 with an increment of 1labeled scientific figures with associated metadata and
to determine the number of clusters. The results wceornetext information. This corpus, namAedCL-Fig,
conanalyzed using the Elbow method and Silhouette Anasliys-ts of≈ 250k objects, of which about 42% are figures
sis. No evident elbow was observed in the Elbow methoadnd about 58% are tables. We also buAilCtL-Fig-pilot, a
curve. The Silhouette diagram, a plot of the numbersoufbset ofACL-Fig, consisting of 1671 scientific figures
clusters versus silhouette score exhibited a clear twuirtnh- 19 manually verified labels. Our dataset includes
ing point a t= 15 , where the score reached the globafiglures extracted from real-world data and contains more
maximum. Therefore, we grouped the figures into 15classes than existing datasets, e.g., DeepFigures and
Figclusters. ureQA.</p>
      <p>To validate the clustering results, 100 figures randomlyOne limitation of our pipeline is that it used VGG16
sampled from each cluster were visually inspected. Duprr-e-trained on ImageNet. In the future, we will improve
ing the inspection, we identified three new figure types:figure representation by retraining more sophisticated
word cloud, pareto, and venn diagram. The ACL-Fig-pilot models, e.g., CoCa, 1[9], on scientific figures. Another
dataset was then built using all manually inspected lfigi-mitation was that determining the number of clusters
ures. Two annotators manually labeled and inspectreedquired visual inspection. We will consider
densitythese clusters. The consensus rate was measured usinbgased methods to fully automate the clustering module.
Cohen’s Kappa coeficient, which was −0.78 (substantial
agreement) for thAeCL-Fig-pilot dataset. For
completeness, we added 100 randomly selected tables. ThereforRe,eferences
theACL-Fig-pilot dataset contains a total of 1671 figures
and tables labeled with 19 classes. The distribution of a[l1l] M. Khabsa, C. L. Giles, The number of scholarly
classes is shown in Figur4e. documents on the public web, PLoS ONE 9 (2014)
e93949.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>