<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Detect, Retrieve, Comprehend: A Flexible Framework for Zero-Shot Document-Level Question Answering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tavish McDonald</string-name>
          <email>mcdonald53@llnl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brian Tsan</string-name>
          <email>btsan@ucmerced.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amar Saini</string-name>
          <email>saini5@llnl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juanita Ordonez</string-name>
          <email>ordonez2@llnl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luis Gutierrez</string-name>
          <email>gutierrez74@llnl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Phan Nguyen</string-name>
          <email>nguyen97@llnl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Blake Mason</string-name>
          <email>mason35@llnl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brenda Ng</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lawrence Livermore National Laboratory</institution>
          ,
          <addr-line>7000 East Avenue Livermore, California 94550</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SDU'23: The Third AAAI Workshop on Scientific Document Under-</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of California Merced</institution>
          ,
          <addr-line>5200 Lake Rd, Merced, CA 95343</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Researchers produce thousands of scholarly documents containing valuable technical knowledge. The community faces the laborious task of reading these documents to identify, extract, and synthesize information. To automate information gathering, document-level question answering (QA) ofers a flexible framework where human-posed questions can be adapted to extract diverse knowledge. Finetuning QA systems requires access to labeled data (tuples of context, question and answer). However, data curation for document QA is uniquely challenging because the context (i.e., text passage containing evidence to answer the question) needs to beretrieved from potentially long, ill-formatted documents. Existing QA datasets sidestep this challenge by providing short, well-defined contexts that are unrealistic in real-world applications. We present a three-stage document QA approach: (1) text extraction from PDF; (2) evidence retrieval from extracted texts to form well-posed contexts; (3) QA to extract knowledge from contexts to return high-quality answers - extractive, abstractive, or Boolean. Using the QASPER dataset for evaluation, ourDetect-Retrieve-Comprehend (DRC) system achieves a +7.19 improvement in Answer -1 over existing baselines due to superior context selection. Our results demonstrate thaDtRC holds tremendous promise as a lfexible framework for practical scientific document QA.</p>
      </abstract>
      <kwd-group>
        <kwd>Answering</kwd>
        <kwd>document understanding</kwd>
        <kwd>information retrieval</kwd>
        <kwd>question answering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Growth in new machine learning publications has
exploded in recent years, with much of this activity
occurring outside traditional publication venues. For
example, arXiv hosts researchers’ manuscripts detailing
the latest progress and burgeon
        <xref ref-type="bibr" rid="ref18">ing initiatives. In 2021</xref>
        alone, over 68,000 machine learning papers were
submitted to arXiv. Since 2015, submissions to this category
have increased yearly at an average rate of 52%. While it
is admirable that the accelerated pace of AI research has
produced many innovative works and manuscripts, the
sheer amount of papers makes it prohibitively dificult
      </p>
      <sec id="sec-1-1">
        <title>Increasingly, researchers turn to scientific search engines (e.g., Semantic Scholar and Zeta Alpha), powered by neu</title>
        <p>CEUR</p>
        <p>CEUR
CEUR</p>
        <p>ceur-ws.org
Question
Evidence
What is the seed lexicon?
three types.</p>
        <p>The seed lexicon consists of positive
and negative predicates. If the
predicate of an extracted event is in the
seed lexicon and does not involve
complex phenomena like negation,
we assign the corresponding
polarity score (+1 for positive events and
-1 for negative events) to the event.</p>
        <p>We expect the model to automatically
learn complex phenomena through
label propagation. Based on the
availability of scores and the types of
discourse relations, we classify the
extracted event pairs into the following
Answer
a vocabulary of positive and negative
predicates that helps determine the
polarity score of an event
to keep pace with the latest developments in the field. evidence retrieval to generate ananswer.
ral information retrieval, to find relevant literature. To tion of document-derived contents, particularly titles and
date, scientific search engines [1, 2, 3] have focused on
K Answers (K=3)
no answer
matches neither
the former nor the
latter event
positive
tive
prediacnadtesnegaPDF
pdf2image</p>
        <p>DiT
pdfminer.six
PDF File</p>
        <p>Page Images</p>
        <p>Paragraph Bounding Boxes
Question Text
Wlexhiactoni?s the seed</p>
        <p>ELECTRA_CE</p>
        <p>Top-K Paragraph Texts (K=3)</p>
        <p>CO (CONCESSION Pairs) r=0.07
The seed lexicon matches
neiCthAer(CtAhUeSEformPaeir s)noTrhe r=0.10
itdshisettsCceheolreOeaudrtNrettsevClheereeExneifStvcor,SoeerInlmaaOtnte,Nmirdoa.anntotdWchrhteteihhyrseapeisendrl-aeisti--- r=0.93
sumcoeuTrthshee sterweoldateliveoexnictostnyhpacevoensis ts
theCrAoeUfveSrpEsoe.sdiWtpiveoelaarasintsiudems.neegtahteive
twporevdeicnattsesh.avIef
ththeesapmredipoclartietioefs.an extracted</p>
        <p>UnifiedQA</p>
        <sec id="sec-1-1-1">
          <title>Detect (Text Extraction)</title>
        </sec>
        <sec id="sec-1-1-2">
          <title>Retrieve (Evidence Retrieval)</title>
        </sec>
        <sec id="sec-1-1-3">
          <title>Comprehend (Question Answering)</title>
          <p>found in the details of the methodology, experimental 2. Dataset
setup, and results sections. Furthermore, questions may
require synthesis of document passages to produce an The Question Answering on Scientific Research Papers
abstractive answer rather than simply extracting a con- (QASPER) dataset consists of 1,585 NLP papers sourced
tiguous span. Reading and manually cross-referencing from arXiv, and is accompanied by 5,049 questions from
the results of several papers is a labor-intensive approach NLP readers and corresponding answers from NLP
practito glean specific knowledge from scientific documents. tioners. Papers inQASPER are cited by their arXiv DOIs,
Therefore, efective tools to help automate knowledge which we used to fetch the original PDF documents as
discovery are sorely needed. input to our system, as our work is focused on knowledge</p>
          <p>A promising approach to extracting knowledge from extraction at the PDF level.
scientific publications is document-level question answer- QASPER contains 7,993 answers categorized by
aning (QA): using an open set of questions to comprehend swer type: Extractive (4142), Abstractive (1931), Yes/No
ifgure captions, tables, and accompanying text [ 6]. Tradi- (1110), and Unanswerable (810). Using only theExtractive,
tionally, the NLP community has focused on usingclean Abstractive and Yes/No answers, we match our model
texts as context to their QA systems. However, this is prediction to the most similar answer when a question
not representative of the vast majority of scholarly infor-has more than one answer, and report our performance
mation found in structured documents. As QA garners accordingly.
interest from the computer vision community, DocVQA QASPER is ideal for evaluating our proposed
frame[7] and VisualMRC [8] have extended document QA to work because it provides: (1) paragraph text and table
extracting evidence from single images, paving the way information to evaluate our layout-analysis model (in
to extend contexts from text to visual sources. its ability to cleanly extract document regions); (2)
ev</p>
          <p>A foundational challenge in building robust document idence paragraphs to validate, and optionally finetune,
QA systems is ensuring well-formed contexts, which en- our evidence retrieval model (in its ability to retrieve
tails accurate text extraction and requires adaptation to good context paragraphs); and (3) ground-truth answers
new document layouts. Nonetheless, even when text can to assess the accuracy of our QA model (in its ability to
be cleanly extracted, there still remains the crucial task answer the question given the context).
of identifying question-relevant paragraphs for answer
prediction. 3. Methodology</p>
          <p>Our contribution is ageneral-purpose system for
QA on full documents in their original PDF form, Document QA on raw PDFs is necessary towards
authat addresses the key challenges of scientific document tomating knowledge extraction from scientific corpora
QA: (1) accurate text extraction from unseen layouts, (2) and has remained an unaddressed problem. To address
evidence retrieval (i.e., context selection), and (3) robust this, we propose a flexible information extraction tool
QA. to alleviate laboriously searching for answers grounded
in evidence. Our system combines: (1) a robust text
detext extraction (§ 3.2), evidence retrieval (§ 3.3), and QA weights:
tion. We decompose this problem into three subtasks: tween  and passage ∈ 
Our work addresses evidence retrieval at the PDF level. that BM25 is a strong baseline across many datasets</p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>Thus, our document QA task is defined as: given a ques</title>
        <p>[14, 15]. Given a question containing tokens 1, … ,  
tion and a PDF document, predict the answer to the ques- and a set of passages , the BM25 retrieval score 
beis defined using TF-IDF token
(§ 3.4).
images, has its semantic regions identified and their
cor</p>
        <p>First, the PDF document, represented as a series of 
25
,
responding text content extracted as passages. Second,

=1
= ∑ log(</p>
        <p>| |
 (  ,  )
)</p>
        <p>(  , )( 1 + 1)
 1(1 −  + 
||
) + (  , )
correspond to the respective tasks ofDetect, Retrieve and
all architecture is shown in Figure2. These components  are constants.</p>
      </sec>
      <sec id="sec-1-3">
        <title>Comprehend, or DRC, which is also the name of our pro</title>
        <p>Irrelevant passages are filtered out so only the most
relethe passages are ranked by their relevance to the question. where | | is the number of passages in the corpus;||
a context and question, the answer is predicted. The
overvant passages are used as contexts for QA. Finally, given passages with token  ; (  , ) is the term frequency of
  in passage  ;</p>
        <p>is the average passage length. 1 and
is the length of the passage; (
 ,  )
is the number of
posed system.
3.2. Detect</p>
      </sec>
      <sec id="sec-1-4">
        <title>The first step of our pipeline is to extract text from</title>
      </sec>
      <sec id="sec-1-5">
        <title>PDF documents. Libraries such aspdfminer.six [9] and</title>
        <p>TesseractOCR [10] extract text from documents indiscrim- the passage  ∈ 
inately, including unwanted page numbers and footnotes,retrieval score  is defined as the dot product of the two
tector for visually rich documents, (2) explicit passage 3.3. Retrieve
retrieval for evidence selection, and (3) multi-format
answer prediction. We used pretrained open-source ma- Evidence retrieval identifies relevant passages by ranking
chine learning models that are efective in a zero-shot
setting. We also finetuned these models to improve our
system’s end-to-end performance.
3.1. Problem Description
them according to their similarity to the question. We
considered several architectures.</p>
        <p>Lexical Retriever</p>
      </sec>
      <sec id="sec-1-6">
        <title>BM25 [13] ranks questions and passages based on token-matching between sparse representations of the question and passage. Prior work has shown</title>
        <p>Document layout analysis models are trained to seg- which has been trained on additional data, as Karpukhin
ment a document into its constituent components (e.g., et al. [16] has shown that the additional data improves
ing boxes and segmentation masks. Predicted regions are a retrieval score  where  (, ; Φ)
labels. It supports prediction of semantic region bound- and passage  separately, cross-encoders [19] compute
[CLS] encodes both
which would need to be filtered out before the extracted
text can be used as context paragraphs. Thus, prior to
text extraction, document layout analysis should be
performed to detect targeted regions (from which text is to
be extracted).
paragraphs, figures, and tables). The Document Image</p>
      </sec>
      <sec id="sec-1-7">
        <title>Transformer (DiT) [11] is designed for layout analysis</title>
        <p>and text detection.DiT uses a masked image modeling
objective to pretrain a Vision Transformer1[2] without
then passed to OCR tools for text extraction.</p>
      </sec>
      <sec id="sec-1-8">
        <title>In the Detect stage ofDRC, the pdf2image library first</title>
        <p>converts each page of the document to images. For each
image, DiT detects the bounding boxes for paragraphs.
The text within each bounding box is then extracted using where 
Dual Encoder</p>
      </sec>
      <sec id="sec-1-9">
        <title>The Dense Passage Retriever D(PR)</title>
        <p>[16] learns via a contrastive training objective with
inbatch negatives and hard negatives chosen byBM25.</p>
        <p>For question  and a set of passages , DPR measures
question-passage similarity with a dual-encoder
architecture, where   encodes the question and   encodes</p>
        <p>to the same latent space [17, 18]. The
resulting embeddings:


,;Φ</p>
        <p>=   (; Φ  )  (; Φ  )
Φ = [Φ , Φ ] denotes the retriever question and
passage encoder parameters. We used the DPRmulti variant,
are candidates for question evidence.
pdfminer.six. The extracted texts form passages which that classifies whether  is relevant to  . Many
crossretrieval generalizability.</p>
        <p>Cross-Encoder Instead of embedding the questio n
question and passage using the CLS token representation
of their concatenation:


,;Φ</p>
        <p>= softmax ( (, ; Φ) [CLS] +  )
and  are the weight and bias in the final layer
encoders have since been proposed and a comparative
analysis was performed 2[0], where the ELECTRA-base
[21] cross-encoder (ELECTRA_CE) was declared as the</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Experimental Setup</title>
      <p>best cross-encoder due to its stability and efectiveness
across datasets. Thus, the ELECTRA_CE model (trained
on MS MARCO [22]) was selected as our starting cross- 4.1. QASPER Baselines
encoder.</p>
      <p>Once the passages have been ranked, the top - most
relevant passages are used as contexts in the QA stage.</p>
      <sec id="sec-2-1">
        <title>Following Dasigi et al.[24], in our QASPER experiments,</title>
        <p>we use the Longformer-Encoder-Decoder L(ED) [25]
as the baseline model for evidence retrieval and QA.</p>
        <p>This model uses a modification of self-attention from
3.4. Comprehend the Transformer architecture 2[6] to encode longer
seThe final stage of DRC is comprehending a document’s quences more eficiently. To jointly answer questions
contents viamulti-format question answering. These for- and decide whether a context is relevant in providing
mats correspond to answer types, which can be extrac- answer evidence,LED optimizes a multi-task objective.
tive, abstractive, or Boolean. (An extractive answer is a In addition to answer generationL, ED adds a
classificaspan of text taken verbatim from an evidence passage.tion head (termedevidence scafold ) that operates over
An abstractive answer is a generated span not quoted each paragraph to predict binary labels (evidence or
nonverbatim from the evidence. A Boolean answer is a binary evidence). Since we discarded unanswerable questions
prediction: yes or no.) from QASPER, we retrain LED on the remaining
ques</p>
        <p>For comprehension, we useUnifiedQA [23], a genera- tions and evaluate with and without evidence
scafoldtive question-answering model that has been pretrained ing. The retrainedLED serves as a fairer competitor to
on 20 datasets and can predict all answer formats with UnifiedQA, which was not pretrained on unanswerable
a single architecture. The answer type returned byUni- questions.
fiedQA depends on the way the question is phrased.</p>
        <p>For each of the  relevant passages (from theRetrieve 4.2. Text Extraction with Layout Analysis
stage), we pair the passage with the question as input to
UnifiedQA to predict an answer. At the end, we hav e
answers – one for each of th e passages.</p>
      </sec>
      <sec id="sec-2-2">
        <title>We use DiT with pdfminer.six for selective text extraction.</title>
        <p>First, a pretrainedDiT model predicts the bounding boxes
of paragraphs on each page. Then,pdfminer.six extracts
text within the bounding boxes. We denote this
twostep procedure asDiT+pdfminer.six, and compare against
TesseractOCR, which takes an image as input and returns
the text found within the image, as well aspdfminer.six’s
high-level extractor (pdfminer.six*), which takes a PDF as
input and exploits PDF metadata to extract texts within sured by the  1 score between the predicted outputs and
4.3. Retriever-QA Implementation Details
For BM25, we create an inverted index onQASPER
validation and test sets using Pyserini2[7] with default
parameters ( 1=0.9,  =0.4). For DPR and ELECTRA_CE, we
the target labeled inQASPER. Adopting the same
notawe also evaluate its precision and recall.
tion as Dasigi et al.[24], we name the  1 scores for our
evidence retrieval and QA as Evidenc e-1 and Answer - 1,
respectively. For text extraction, in addition to it s1 score,</p>
      </sec>
      <sec id="sec-2-3">
        <title>Since each question inQASPER is labeled with its answer(s) and accompanying evidence, it is possible to eval</title>
        <p>4.4. Evaluation Metrics
We evaluate DRC’s text extraction, retrieval, and QA
stages separately. For each stage, performance is
meastart with pretrained models from Hugging Face, then uate both our QA and evidence retrieval stages using this
ifnetune them per hyperparameters shown in Table 1. In
ifnetuning DPR and ELECTRA_CE, we sample batches
single dataset. For QA, Answer- 1 is calculated between
the tokens in the predicted answer and the tokens in the
containing a 1:4 ratio of positive to negative evidence target answer. For our evidence retrieval stage, which
passages.</p>
      </sec>
      <sec id="sec-2-4">
        <title>For UnifiedQA, we use the unifiedqa-v2-t5-large</title>
        <p>1363200 model from Hugging Face. We finetune it in
ranks passages by their relevance to a given question,
Evidence - 1 is calculated between a fixed percentage of
the top ranking passages and the set of passages labeled
a weakly supervised manner using evidence passages as evidence inQASPER.
ranked by ELECTRA_CE but with the original questions
and answers fromQASPER. The choice to use retrieved</p>
        <p>QASPER also contains the plain text for each of its
documents, organized so that text from paragraphs and
passages (instead of the human-labeled evidence passages tables are separated. We use this plain text to evaluate
from QASPER) should make our system more robust to
the eficacy of our text extraction to extract only the
noisy context paragraphs. We show that a pretrained text primary content of PDF document. The precision, recall,
extractor and evidence retriever can adapUtnifiedQA to
the domain ofQASPER papers without labeled evidence. the tokens in a document’s extracted text and its tokens in
and  1 score for text extraction are calculated between</p>
      </sec>
      <sec id="sec-2-5">
        <title>QASPER’s plain text version. For all of our experiments, tokenization is performed at the word level, using our pretrainedUnifiedQA model’s tokenizer.</title>
        <p>Table 6 ELECTRA_CE and UnifiedQA models are not
fineComparison between ELECTRA cross-encoders against LED tuned. DRC achieves an overall +2.31 improvement
baselines in terms of Evidence- 1. in Answer - 1 over LED-base without scafolding on
Model VEavli.denceT-es1t aQpApSrPoEacRh’s, wteestasppplilty. Twoeiamkpsruopveeruvpisoinonthteo ffinuelltyunzeeroU-snhio-t
fiedQA: we sample extracted passages according to their
ELECTRA_CE 31.75 36.37 retrieval scores from the pretrainedELECTRA_CE model,
ELECTRA_CE-ft 31.58 36.12 assuming that higher ranked passages are correlated with
LLEEDD--bbaassee-InfoNCE 2234..9940 2390..8650 selection probability for answer prediction. Thus, we are
LED-large 31.25 39.37 able to finetune UnifiedQA without access to
humanlabeled contexts, since labeled question-answer pairs are
generally unavailable for large technical corpora. This
5. Results ifnetuning approach yields a +6.25 improvement to
LEDbase in overall Answer - 1 on the test dataset.</p>
        <p>We demonstrateDRC’s efectiveness on document QA by To analyze Answer - 1 performance when
groundmeasuring its end-to-end performance. We also evaluate truth question-passage pairs are available, we consider an
its constituent components on text detection, evidence re-ELECTRA_CE retriever finetuned on QASPER’s training
trieval, and QA tasks against existingQASPER baselines. set. We then finetune UnifiedQA through weak
superFor evidence retrieval, we study the benefits of having a vision using the now improved retrieverD.RC with a
separate retrieval process, in contrast to the evidence se- finetuned ELECTRA_CE shows modest gains over the
lection scafold for LED. Furthermore, we explore DRC’s zero-shot system but still lesser performance compared to
performance in both zero-shot and finetuned settings, to the pretrainedELECTRA_CE with a weakly-supervised
assess its performance under varying degrees of access UnifiedQA. This suggests that downstream QA
perforto labeled data. mance is better improved by adapting to the target
domain QASPER documents, than by receiving more
relevant passages.
5.1. End-to-End QA System Combining a finetuned ELECTRA_CE retriever with a
Table 2 shows DRC’s performance in terms of Answer-1. weakly supervisedUnifiedQA model shows the greatest
In these experiments, we extract text from documents improvements overLED-base without scafolding, +7.19
using eitherTesseractOCR or DiT and rank passages us- in Answer - 1 on the and test dataset for all answer types.
ing ELECTRA_CE. We then pass highly-ranked passages We observe that using finetuned ELECTRA_CE for weak
to UnifiedQA for answer prediction. First, we study the supervision shows worse performance on Boolean
quesinfluence of the text detection model on Answer- 1 per- tions than using the pretrainedELECTRA_CE to weakly
formance by comparingTesseractOCR to DiT. While we supervise UnifiedQA. This discrepancy is likely due to
observe that using DiT reports higher Answer - 1 than the small proportion of Boolean samples in the validation
TesseractOCR across all answer types, the diference is and test datasets compared to other formats, 13% and 15%
negligible. respectively.</p>
        <p>Next, we examineDRC in the zero-shot setting where Across all experiments, DRC demonstrates superior
performance toLED while solving a more dificult task: DPR’s inner product between question and passage or
DRC starts from PDFs whileLED starts from clean texts. BM25’s weighted term matching.</p>
        <p>DRC bridges an essential gap in real-world applications To analyze how retrievers perform with ground-truth
for scientific knowledge extraction because PDFs are question-passage pairs, we also evaluate passage retrieval
directly processed as input. In the following discussion,with DPR and ELECTRA cross-encoders finetuned on
we validate our individual system components. QASPER. Here, ELECTRA_CE outperformsDPR on the
test data for = 5%, 10% and 20% by an average recall
5.2. Text Detection of +3.95, +6.58, and +6.69, respectively. Notably,DPR has
higher recall for =1%. We conjecture that this may be
We evaluate three diferent methods for text extraction due to DPR’s contrastive objective utilizing hard
negafrom PDF files: (1) DiT+pdfminer.six, (2) pdfminer.six*, tive sampling, but further analysis on the relationship
and (3) TesseractOCR. Table 3 reports the average pre- between training objective and ranking is needed.
cision, recall, and 1 between the extracted tokens and
those in the ground-truth text. 5.3.1. Comparison to Evidence Selection Scafold</p>
        <p>For paragraph extraction, DiT+pdfminer.six has
better precision than pdfminer.six* (+19.33) and Tesserac- To compare against LED’s evidence scafold, we now
tOCR (+19.01). We attribute this improvement to ex- treat ELECTRA_CE as a binary classifier. Akin to LED’s
tracting fewer unwanted artifacts (e.g., page numbers, evidence scafold, we use the [CLS] representation of the
headers, footers, and footnotes). For text within ta- question-passage pair as input to a single layer neural
bles, only DiT+pdfminer.six is efective of-the-shelf. network to estimate the probability that the passage is
relpdfminer.six* and TesseractOCR do not disambiguate be- evant as evidence to the question and use a classification
tween text in and outside of tablesp.dfminer.six* and threshold of 0.5 [19]. Table 6 illustrates the evidence
clasTesseractOCR would sufice if text is contained only in sification performance of zero-shot and finetuned
ELECtables, or only in paragraphs, but not a mixture of the TRA_CE models against LED variations. Evidence- 1
two because the text from tables and paragraphs will be scores are computed using the extracted passages
classiinterspersed. ifed as evidence with respect to the ground-truth set
labeled inQASPER. We observe that the diference between
5.3. Evidence Passage Retrieval the zero-shot and finetuned ELECTRA_CE models is
negligible. On the test split, zero-shotELECTRA_CE shows
We compare DPR, ELECTRA_CE, and BM25 by their a notable Evidence- 1 improvement overLED-base
augability to rank passages by relevance to questions. Table4 mented with InfoNCE loss 2[9], but is outperformed by
shows the recall of evidence passages within various LED-large. This agrees with findings from Dasigi et al.
percentages of the top ranked passages, averaged over [24] that LED-large generally outperformsLED-base for
all questions inQASPER, for the retrievers in both zero- retrieval but not QA. Thus, we consider onlyLED-base
shot and finetuned settings. As questions are posed for in our subsequent experiments with downstream QA.
a specific document, our retrievers consider a variable
number of passages per question because documents 5.4. Question Answering
vary in length. Since top- penalizes longer documents
when  is small, we measure recall using top- % for Here, we focus on the efect of using retrieval
mecha ∈ {1, 5, 10, 20} . nisms on QASPER’s plain text passages for answer
pre</p>
        <p>In the zero-shot setting,BM25 outperformsDPR by an diction. We report Answer- 1 scores for extractive,
abaverage of +15.58 gain in recall on the test data. These stractive, and Boolean answer types. We also finetune
results support findings from Sciavolino et al. [15], who DPR and ELECTRA_CE models on QASPER’s train split
reported that DPR trained on Natural Questions 2[8] and compare againstLED variations.
underperformedBM25 when faced with the new ques- Table 5 shows UnifiedQA’s Answer- 1 using the
tion patterns and entities found in their EntityQuestionshighest-ranked passage from each retriever. We observe
dataset. Thus, DPR requires finetuning and is less gener- that UnifiedQA (first 5 rows) generally yields higher
alizable than BM25, which has no trainable parameters. Answer- 1 scores across answer types, datasets, and
re</p>
        <p>ELECTRA_CE, with average recall gains of +7.2 and trievers thanLED baselines (last 2 rows). The exception is
+13.78 over BM25 and DPR, respectively, is the clear win- extractive answers from the test set whereBM25 reports
ner. We hypothesize thatELECTRA_CE’s success is due a lower Answer- 1 than LED-base. Similarly, a zero-shot
to the explicit interaction between every token of the DPR retriever performs worse than bothLED models.
question and passage through its cross-attention mecha- Among zero-shot retrievers, ELECTRA_CE yields the
nism, ofering a more expressive similarity function than best performance on the overall test set with a +7.92 and
+6.54 Answer- 1 increase over LED with and without
60
1
rF55
e
w
sn50
A
45
40</p>
        <p>Retriever
ELECTRA_CE-FT
DPR-FT
ELECTRA_CE
BM25
DPR
our experiments include:
1. DiT demonstrates superior text extraction
perfor</p>
        <p>mance to pdfminer.six and TesseractOCR.
2. Zero-shot ELECTRA_CE ofers the best
retrieval performance for all top %- (where  ∈
{1, 5, 10, 20}).
3. DRC adapts to new domains through
weaklysupervised training on evidence passages leading
to substantially improved answer prediction over
LED baselines.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>6. Conclusion</title>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <sec id="sec-4-1">
        <title>We introducedDRC, an end-to-end QA system for au</title>
        <p>
          tomating manual knowledge extraction from scientific This work was performed under the auspices of the U.S.
PDF documents. We showed thatDRC greatly improves Department of Energy by Lawrence Livermore National
over existing baselines, which act on clean texts and Laboratory under Contract DE-AC52-07NA27344.
sidestep the challenge of PDF-to-text extraction. Through
extensive experiments, we evaluate our pipeline
components in both zero-shot and finetuned settings. In practice, References
datasets as comprehensive asQASPER are few and may [1] M. Fadaee, O. Gureenkova, F. Rejon Barrera,
not be feasible for niche domains. In such cases, a fully C. Schnober, W. Weerkamp, J. Zavrel, A new
neuzero-shot pipeline is mandatory for document QA, and ral search and insights platform for navigating and
DRC can be weakly supervised to adapt to specific do- organizing AI research, in: Proceedings of the First
mains. Our DRC sets a new benchmark forQASPER and Workshop on Scholarly Document Processing,
Asserves as a proof of concept for an end-to-end document sociation for Computational Linguistics, Online,
QA system, from PDF to answer. Key takeaways from
tion Processing Systems, volume 34, Curran puting Machinery, New York, NY, USA, 2021, p.
Associates, Inc., 2021, pp. 25968–25981. URL: 2356–2362. URL: https://doi.org/10.1145/3404835.
https://
          <xref ref-type="bibr" rid="ref13">proceedings.neurips.cc/paper/2021</xref>
          /file/ 3463238. doi:10.1145/3404835.3463238.
da3fde159d754a2555eaa198d2d105b2-Paper.pd.f [28] T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins,
[19] R. Nogueira, K. Cho, Passage re-ranking with bert, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin,
        </p>
        <p>
          ArXiv abs/1901.04085 (2019). M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova,
[20] X. Zhang, A. Yates, J. Lin, Comparing score ag- L. Jones, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le,
gregation approaches for document retrieval with S. Petrov, Natural questions: a benchmark for
quespretrained transformers, in: Advances in Informa- tion answering research, Transactions of the
Assotion Retrieval: 43rd European Conference on IR Re- ciation of Computational Linguistics (2019).
search, ECIR 2021, Virtual Event, March 28 – April [29] A. Caciularu, I. Dagan, J. Goldberger, A. Cohan,
1, 2021, Proceedings, Part II, Springer-Verlag, Berlin, Utilizing evidence spans via sequence-level
conHeidelberg, 2021, p. 150–163. URL: https://doi. trastive learning for long-context question
answerorg/10.1007/978-3-030-72240-1_11. doi:10.1007/ ing, arXiv preprint arXiv:2112.08777 (2021).
978-3-030-72240-1_11. [30] J. Johnson, M. Douze, H. Jégou, Billion-scale
simi[21] K. Clark, M. Luong, Q. V. Le, C. D. Manning, ELEC- larity search with gpus, IEEE Transactions on Big
TRA: pre-training text encoders as discriminators Data 7 (2019) 535–547.
rather than generators, in: 8th International
Conference on Learning Representations, ICLR 2020,
Addis Ababa, Ethiopia, April 26-30,
          <xref ref-type="bibr" rid="ref1 ref15 ref2">2020,
OpenReview.net, 2020</xref>
          . URL:https://openreview.net/forum?
id=r1xMH1BtvB.
[22] P. Bajaj, D. Campos, N. Craswell, L. Deng, J. Gao,
        </p>
        <p>X. Liu, R. Majumder, A. McNamara, B. Mitra,
T. Nguyen, Ms marco: A human generated machine
reading comprehension dataset, arXiv preprint
arXiv:1611.09268 (2016).
[23] D. Khashabi, Y. Kordi, H. Hajishirzi,
Unifiedqav2: Stronger generalization via broader
cross-format training, arXiv preprint
arXiv:2202.12359 (2022). https://huggingface.</p>
        <p>co/allenai/unifiedqa-v2-t5-large-136320.0
[24] P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith,</p>
        <p>M. Gardner, A dataset of information-seeking
questions and answers anchored in research papers, in:</p>
        <p>NAACL, 2021.
[25] I. Beltagy, M. E. Peters, A. Cohan, Longformer:</p>
        <p>The long-document transformer, arXiv preprint
arXiv:2004.05150 (2020).
[26] A. Vaswani, N. Shazeer, N. Parmar, J.
Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I.
Polosukhin, Attention is all you need, in: I. Guyon,
U. V. Luxburg, S. Bengio, H. Wallach, R.
Fergus, S. Vishwanathan, R. Garnett (Eds.),
Advances in Neural Information Processing
Systems, volume 30, Curran Associates, Inc., 2017.</p>
        <p>URL: https://proceedings.neurips.cc/paper/2017/
file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pd.f
[27] J. Lin, X. Ma, S.-C. Lin, J.-H. Yang, R. Pradeep,</p>
        <p>R. Nogueira, Pyserini: A python toolkit for
reproducible information retrieval research with
sparse and dense representations, in:
Proceedings of the 44th International ACM SIGIR
Conference on Research and Development in
Information Retrieval, SIGIR ’21, Association for
Com</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <year>2020</year>
          , pp.
          <fpage>207</fpage>
          -
          <lpage>213</lpage>
          . URL: https://aclanthology.org
          <article-title>/ supervised pre-training for document image trans-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2020.sdp-
          <volume>1</volume>
          .23. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .sdp-
          <volume>1</volume>
          .23. former,
          <source>arXiv preprint arXiv:2203.02378</source>
          (
          <year>2022</year>
          ). [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , Talk to papers: Bringing neu- [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          , D. Weis-
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>Proceedings of the 58th Annual Meeting of the M. Minderer</article-title>
          , G. Heigold,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          , J. Uszkoreit,
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Association for Computational Linguistics: Sys- N. Houlsby</surname>
          </string-name>
          ,
          <article-title>An image is worth 16x16 words: Trans-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>tional Linguistics</surname>
          </string-name>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>30</fpage>
          -
          <lpage>36</lpage>
          . URL: ternational Conference on Learning Representa-
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          https://aclanthology.org/
          <year>2020</year>
          .acl-demos.5. doi:10. tions,
          <year>2021</year>
          . URL: https://openreview.net/forum?id=
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <volume>18653</volume>
          /v1/
          <year>2020</year>
          .
          <article-title>acl-demos.5</article-title>
          . YicbFdNTTy. [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Parisot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zavrel</surname>
          </string-name>
          , Multi-objective representation [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Robertson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zaragoza</surname>
          </string-name>
          ,
          <article-title>The probabilistic rele-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>learning for scientific document retrieval, in: Pro- vance framework: Bm25 and beyond</article-title>
          , Found. Trends
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>ceedings of the Third Workshop on Scholarly Doc- Inf. Retr</source>
          .
          <volume>3</volume>
          (
          <year>2009</year>
          )
          <fpage>333</fpage>
          -
          <lpage>389</lpage>
          . URL:https://doi.org/10.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>ument Processing</surname>
          </string-name>
          ,
          <source>Association for Computational</source>
          <volume>1561</volume>
          /1500000019. doi:
          <volume>10</volume>
          .1561/1500000019.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Linguistics</surname>
          </string-name>
          , Gyeongju, Republic of Korea,
          <year>2022</year>
          . URL: [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Thakur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rücklé</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Srivastava,
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          https://aclanthology.org/
          <year>2022</year>
          .sdp-
          <volume>1</volume>
          .9.
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          , Beir: A heterogeneous benchmark [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kardas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Czapla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Stenetorp</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Ruder, for zero-shot evaluation of information retrieval</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>Proceedings of the 2020 Conference on Empirical Track on Datasets and Benchmarks</source>
          , volume
          <volume>1</volume>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>Association for Computational Linguistics, Online, neurips</article-title>
          .cc/paper/2021/file/
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <year>2020</year>
          , pp.
          <fpage>8580</fpage>
          -
          <lpage>8594</lpage>
          . URL: https://aclanthology. 65b9eea6e1cc6bb9f0cd2a47751a186f-Paper-round2.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          org/
          <year>2020</year>
          .emnlp-main.
          <volume>692</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          . pdf.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          emnlp-main.
          <volume>692</volume>
          . [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sciavolino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Sim[5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sotudeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goharian</surname>
          </string-name>
          ,
          <article-title>On generat- ple entity-centric questions challenge dense re-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>ing extended summaries of long documents, arXiv trievers</article-title>
          ,
          <source>in: Proceedings of the 2021 Conference</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          preprint arXiv:
          <year>2012</year>
          .
          <volume>14136</volume>
          (
          <year>2020</year>
          ).
          <source>on Empirical Methods in Natural Language Pro</source>
          [6]
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>u</year>
          . Borchmann,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pietruszka</surname>
          </string-name>
          , T. Stanislawek, cessing, Association for Computational Linguis-
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <article-title>Due: End-to-end document understanding bench- lic</article-title>
          ,
          <year>2021</year>
          , pp.
          <fpage>6138</fpage>
          -
          <lpage>6148</lpage>
          . URL: https://aclanthology.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          mark, in: J.
          <string-name>
            <surname>Vanschoren</surname>
          </string-name>
          , S. Yeung (Eds.), Proceed- org/
          <year>2021</year>
          .emnlp-main.
          <volume>496</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>ings of the Neural Information Processing Systems emnlp-main.496.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>Track on Datasets and Benchmarks</source>
          , volume
          <volume>1</volume>
          ,
          <year>2021</year>
          . [16]
          <string-name>
            <given-names>V.</given-names>
            <surname>Karpukhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Oguz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . Wu,
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          neurips.cc/paper/2021/file/ trieval for open
          <article-title>-domain question answering</article-title>
          , in:
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <source>069059b7ef840f0c74a814ec9237b6ec-Paper-round2. Proceedings of the 2020 Conference on Empirical</source>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          pdf.
          <source>Methods in Natural Language Processing (EMNLP)</source>
          , [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mathew</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Karatzas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jawahar</surname>
          </string-name>
          ,
          <article-title>Docvqa: A Association for Computational Linguistics</article-title>
          , Online,
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <article-title>dataset for vqa on document images</article-title>
          ,
          <source>in: Proceed- 2020</source>
          , pp.
          <fpage>6769</fpage>
          -
          <lpage>6781</lpage>
          . URL: https://aclanthology.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>ings of the IEEE/CVF Winter Conference on Appli- org/2020.emnlp-main.550. doi:10</source>
          .18653/v1/
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>cations of Computer Vision</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>2200</fpage>
          -
          <lpage>2209</lpage>
          . emnlp-main.
          <volume>550</volume>
          . [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tanaka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Nishida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yoshida</surname>
          </string-name>
          , Visualmrc: Ma- [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bromley</surname>
          </string-name>
          , I. Guyon,
          <string-name>
            <given-names>Y.</given-names>
            <surname>LeCun</surname>
          </string-name>
          , E. Säckinger, R. Shah,
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>cial Intelligence</source>
          <volume>35</volume>
          (
          <year>2021</year>
          )
          <fpage>13878</fpage>
          -
          <lpage>13888</lpage>
          . URL:https: tor (Eds.),
          <source>Advances in Neural Information Process-</source>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          //ojs.aaai.org/index.php/AAAI/article/view/1763 5. ing Systems, volume
          <volume>6</volume>
          , Morgan-Kaufmann,
          <year>1993</year>
          . [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shinyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Guglielmetti</surname>
          </string-name>
          , Pdfminer.sixh,ttps: URL: https://proceedings.neurips.cc/paper/1993/
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          //github.com/pdfminer/pdfminer.si,x2020. file/288cc0ff022877bd3df94bc9360b9c5d-Paper.p.
          <fpage>df</fpage>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <article-title>An overview of the tesseract ocr engine</article-title>
          , [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Reddy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          , C. Dyer,
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <source>ume 02</source>
          , ICDAR '07, IEEE Computer Society, USA, domain question answering, in: M. Ranzato,
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <year>2007</year>
          , p.
          <fpage>629</fpage>
          -
          <lpage>633</lpage>
          . A.
          <string-name>
            <surname>Beygelzimer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Dauphin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          , Dit: Self- Vaughan (Eds.),
          <source>Advances in Neural Informa-</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>