<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>J. Barrow, R. Patel, M. Kharkovski, B. Davies, R. Schmitt, Safepassage: High-fidelity information extraction with
black box llms, CoRR abs/</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1101/2025.09.11.675507</article-id>
      <title-group>
        <article-title>A Collection of Systematic Reviews in Computer Science</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pierre Achkar</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tim Gollub</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Potthast</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bauhaus-Universität Weimar</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Kassel University</institution>
          ,
          <addr-line>hessian.AI, ScaDS.AI</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Leipzig University</institution>
          ,
          <addr-line>Fraunhofer ISI Leipzig</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2510</year>
      </pub-date>
      <volume>00276</volume>
      <fpage>71</fpage>
      <lpage>77</lpage>
      <abstract>
        <p>Systematic reviews are the standard method for synthesizing scientific evidence, but their creation requires substantial manual efort, particularly during retrieval and screening. While recent work has explored automating these steps, evaluation resources remain largely confined to the biomedical domain, limiting reproducible experimentation in other domains. This paper introduces SR4CS, a large-scale collection of systematic reviews in computer science, designed to support reproducible research on Boolean query generation, retrieval, and screening. The corpus comprises 1,212 systematic reviews with their original expert-designed Boolean search queries, 104,316 resolved references, and structured methodological metadata. For controlled evaluation, the original Boolean queries are additionally provided in a normalized, approximated form operating over titles and abstracts. To illustrate the intended use of the collection, baseline experiments compare the approximated expert Boolean queries with zero-shot LLM-generated Boolean queries, BM25, and dense retrieval under a unified evaluation setting. The results highlight systematic diferences in precision, recall, and ranking behavior across retrieval paradigms and expose limitations of naive zero-shot Boolean generation. SR4CS is released under an open license on Zenodo, 1 together with documentation and code,2 to enable reproducible evaluation and future research on scaling systematic review automation.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Systematic reviews</kwd>
        <kwd>Boolean queries</kwd>
        <kwd>Information retrieval</kwd>
        <kwd>Computer science</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Systematic reviews are the standard method for synthesizing evidence on focused research questions.
They are valued for their completeness, transparency, and reproducibility, and their results often
serve as a guide for future research, policy, and practice [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Conducting a systematic review is
timeconsuming and involves steps such as defining research questions, retrieving candidate papers, screening
for relevance, and synthesizing findings [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Retrieving candidate papers is a vital step, typically
accomplished by translating the review topic into complex Boolean queries to support interpretability
and reproducibility. These queries define the pool of eligible candidate studies and directly influence
the outcome of the review because any loss in recall during retrieval can hardly be recovered [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        In recent years, increasing eforts have been made to automate systematic reviews to reduce the
cost of their creation, both end-to-end and in individual phases such as query formulation, screening,
and synthesis. Work on automatic Boolean query generation ranges from computational adaptations
of expert-designed strategies [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to recent investigations using large language models (LLMs), which
show potential but also exhibit substantial trade-ofs in precision and recall [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. LLMs have also been
explored for screening, where recall-oriented calibration can yield promising zero-shot performance [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
and agent-based systems are emerging to orchestrate multiple stages of the systematic review
worklfow [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Despite this progress, the reliability and reproducibility of automated methods remain open
research questions that require controlled evaluation.
      </p>
      <p>
        However, evaluation resources for systematic review automation remain heavily concentrated in
the biomedical domain. Datasets such as the SIGIR 2017 SysRev Query Collection [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the CLEF
TAR 2019 corpus [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and the Seed Studies collection [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] provide valuable test collections for medical
1https://doi.org/10.5281/zenodo.17163932
2https://github.com/webis-de/scolia26-sr4cs
SCOLIA 2026, the Second International Workshop on Scholarly Information Access, ECIR 2026, 2nd April 2026, Delft, Netherland
0009-0007-0791-9078 (P. Achkar); 0000-0003-1737-6517 (T. Gollub); 0000-0003-2451-0665 (M. Potthast)
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
94 26M MEDLINE records Retrieval; screening
123 PubMed baseline (2019) Retrieval; screening
40 PubMed retrieved sets + Retrieval (seed-based);
      </p>
      <p>snowballing screening; citation chasing</p>
      <sec id="sec-1-1">
        <title>1,212 104k references (89k with abstracts)</title>
      </sec>
      <sec id="sec-1-2">
        <title>Retrieval; screening</title>
        <p>systematic reviews, but there is no comparable large-scale resource for other domains. In the absence of
domain-specific benchmarks, evaluating retrieval and screening methods often requires costly expert
involvement, constraining scale and reproducibility.</p>
        <p>To address this gap, we introduce SR4CS, the first large-scale test collection of systematic reviews
in computer science, paired with the original Boolean search strategies reported by the authors and
curated reference pools. Each review includes the original Boolean query, alongside an approximation
normalized to a unified title-and-abstract-only retrieval setting for controlled evaluation. In addition,
each review is associated with structured methodological metadata, including research objectives,
inclusion and exclusion criteria, databases searched, and temporal constraints. In total, SR4CS contains
1,212 systematic reviews with 104,316 resolved references, of which 89,447 include abstracts. The
dataset is publicly available together with documentation and code, enabling reproducible evaluation of
Boolean query translation, retrieval efectiveness, and screening methods in computer science.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Existing evaluation resources for retrieval and screening in systematic reviews have largely focused
on the biomedical domain. While these resources vary in scale, design, and supported tasks, they all
provide reusable test collections for evaluating retrieval efectiveness, screening, and related automation
methods. Table 1 summarizes the most relevant datasets and contrasts them with SR4CS.</p>
      <p>
        The SIGIR 2017 SysRev Query Collection [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is based on ∼ 26 million MEDLINE records derived from
94 Cochrane reviews published between 2014 and 2016. The reported search strategies were converted
into executable queries, and included and excluded references were mapped to MEDLINE identifiers,
enabling controlled experiments on retrieval and screening prioritization in the biomedical domain.
      </p>
      <p>
        The Seed Studies Collection [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] was developed to support research on seed-driven search strategies.
It comprises 40 medical systematic review topics curated by information specialists and includes the
original PubMed Boolean queries, seed studies, retrieved records, and final included references. By
explicitly distinguishing genuine seed studies from pseudo-seeds, the collection enables more realistic
evaluation of automatic query formulation, screening prioritization, and citation chasing.
      </p>
      <p>
        The CLEF TAR 2019 corpus [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], developed as part of the CLEF lab on technology-assisted reviews,
is also grounded in Cochrane reviews and PubMed. It includes expert queries, retrieved records, and
relevance judgments for 123 topics across several medical review types, and supports evaluation of
protocol querying and screening prioritization.
      </p>
      <p>SR4CS extends these eforts beyond biomedicine by providing the first large-scale test collection of
systematic reviews in computer science. It preserves the original Boolean search strategies as reported
in the reviews, while also providing normalized, approximated variants for controlled evaluation. With
1,212 systematic reviews and resolved reference pools, SR4CS supports reproducible evaluation of</p>
      <sec id="sec-2-1">
        <title>3https://github.com/ielab/SIGIR2017-SysRev-Collection 4https://github.com/ielab/sysrev-seed-collection 5https://github.com/CLEF-TAR/tar</title>
        <p>Boolean query translation, retrieval efectiveness across paradigms, and screening methods in a domain
where comparable evaluation resources have previously been lacking.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The SR4CS Corpus</title>
      <p>The corpus construction comprises three main stages: data collection and filtering, parsing and
information extraction, and resolving references. Figure 1 provides an overview of the pipeline.
"Systematic Review"
11,317
results</p>
      <p>Filtering</p>
      <p>Filtering
104,316
3.1. Data Collection
A set of candidate systematic reviews was retrieved from DBLP by searching titles for “systematic
review”, yielding 11,317 results. We applied filter rules to retain only genuine systematic reviews
that required clear wording (e.g., “systematic review of/on”), peer review procedures, open access,
and valid DOIs. Outliers due to extreme page counts were removed. After filtering, 1,339 candidate
reviews remained. In the following steps, only reviews that explicitly stated their Boolean queries were
considered for inclusion in the final dataset.
3.2. Parsing and Extraction
The PDF files of the systematic reviews were downloaded and converted into Markdown format for
further processing. To achieve the conversion, Nanonets-OCR-s,6 a modern OCR model that produces
semantically enriched Markdown text rather than plain text, was used. Unlike traditional OCR tools,
it can recognize LaTeX equations, describe figures and tables, and preserve document structure. This
functionality was particularly important because many reviews contain Boolean queries and other
important information in figures or tables.</p>
      <p>
        The documents were then processed in a zero-shot settings using GPT-4.1 Mini to extract a series
of structured fields: (i) databases used, (ii) Boolean queries reported, (iii) year range, (iv) language
restrictions, (v) inclusion and exclusion criteria, (vi) main topic, (vii) objective, (viii) research questions,
and (ix) whether snowballing or citation chasing was performed. This approach follows recent work
demonstrating the feasibility of LLM-based zero-shot data extraction in scientific documents [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ].
      </p>
      <p>
        To assess extraction quality, the extracted fields were manually inspected for errors and
inconsistencies. Additionally, we randomly selected 10% of the reviews and manually extracted the same
information independently. In 85% of the sample cases, the fields extracted by the LLM matched the
manual extraction at the field level, and the errors found were minimal, such as minor discrepancies
in the wording of the inclusion/exclusion criteria or the specification of certain field values (e.g.,
language restrictions) based on indirect clues rather than explicit statements. These findings indicate that
zero-shot LLM-based extraction provides a suficiently reliable basis for large-scale corpus construction,
while still leaving room for more robust, feedback-driven extraction methods in future extensions.
Finally, we excluded reviews that did not explicitly state their Boolean queries, resulting in a final
dataset of 1,212 systematic reviews.
3.3. Reference Resolution
References were extracted from each systematic review via a hybrid pipeline combining AnyStyle7 and
GROBID,8 a configuration proven to be highly efective for structured citation analysis [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. First,
nonscientific items such as reports, interviews, and patents were removed, excluding 10,101 entries. The
remaining references were enriched with abstracts from public bibliographic APIs, including Crossref,
OpenAlex, Semantic Scholar, PubMed, Europe PMC, and the arXiv API. During resolution, DOI matches
were prioritized, while references without a DOI were handled through a high-threshold title-author
match and consistency checks for the publication year to minimize false positives. Overall, this process
produced a consolidated reference pool on which the SR4CS corpus is based.
3.4. Corpus Statistics
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments and Intended Use</title>
      <p>The SR4CS corpus is released to support reproducible research on retrieval and screening for systematic
reviews in computer science. The dataset is designed for three main tasks: (i) evaluating automatic
Boolean query generation against approximations of expert-designed Boolean queries, (ii) comparing
retrieval paradigms under realistic systematic review topics, and (iii) investigating screening behavior
using curated reference pools. To illustrate its use, we present baseline experiments for Tasks (i) and (ii),
evaluating Boolean, probabilistic (BM25), and dense retrieval methods using titles and abstracts only.</p>
      <sec id="sec-4-1">
        <title>7https://github.com/inukshuk/anystyle 8https://github.com/kermitt2/grobid</title>
        <p>0.238
0.083
0.262
0.373
4.1. Approximating Expert-Designed Boolean Queries
A Boolean retrieval baseline is constructed by approximating the expert-designed queries reported in the
original reviews. Since these queries rely on database-specific syntax and metadata fields, normalization
is required to obtain a unified Boolean formulation operating over document titles and abstracts.</p>
        <p>Query rewriting. Reported Boolean queries are converted into valid SQLite FTS5 MATCH syntax
using GPT-4.1-mini in a zero-shot setting. The query structure is preserved, restricted to the title
and abstract fields, and stripped of unsupported metadata filters. A random 25% sample is manually
verified, with only seven queries requiring minor corrections, primarily for verbatim phrases.</p>
        <p>Index and execution. A SQLite FTS5 index is created using the unicode61 tokenizer with diacritics
folded and prefix indexing (2–10) to account for common spelling and formatting variations. All refined
queries are executed against this index; when multiple Boolean queries are associated with a review,
their result sets are merged via union.
4.2. Alternative Retrieval Baselines
Using the Boolean approximation as a reference point, we evaluate three additional baselines operating
on the same corpus and fields: (i) zero-shot Boolean queries generated from the review title and objective
using GPT-4.1-mini, (ii) a keyword-based BM25 baseline, and (iii) a dense semantic retriever based
on the all-MiniLM-L6-v2 Sentence-BERT model.</p>
        <p>For both BM25 and dense retrieval, queries are constructed from the review title, the objective, and
their concatenation. Results are reported for the best-performing variant (title + objective). Retrieval
is capped at 1,000 documents per query to reflect a fixed screening budget and ensure comparability
across retrieval methods.
4.3. Evaluation and Results
All retrieval outputs are compared against the curated reference pools to compute precision, recall, F1,
and F3, macro-averaged across reviews. For ranked retrieval methods, we additionally report MAP,
P@10, and R@100.</p>
        <p>Although Boolean retrieval is inherently unranked, Boolean result sets are deterministically ranked
using SQLite’s internal BM25 scoring function for the purpose of ranking-based evaluation.</p>
        <p>Table 4 reports set-based screening efectiveness, while Table 5 reports ranking efectiveness across
all evaluated methods. Together, these results provide a consistent and reproducible reference point for
evaluating Boolean query generation and alternative retrieval paradigms within the SR4CS corpus.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>
        The development of SR4CS builds on recent advances in document processing and natural language
technologies for large-scale extraction from scientific publications. OCR parsers based on vision–language
models improve handling of complex PDF layouts [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], LLMs support the extraction of methodological
information from scientific texts [
        <xref ref-type="bibr" rid="ref13 ref16">13, 16</xref>
        ], and GROBID continues to perform strongly in reference
extraction [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. By integrating these components into a unified pipeline, SR4CS demonstrates a
reproducible approach to constructing systematic review corpora and provides a transferable methodology
for building evaluation resources across scientific domains.
      </p>
      <p>The retrieval results clarify how diferent paradigms behave when evaluated under the same
titleand-abstract-only constraints imposed during corpus construction. Approximated versions of
expertdesigned Boolean queries achieve the highest precision (0.352), confirming that Boolean logic remains
the most efective means of controlling screening efort; the moderate recall (0.342) primarily reflects
the approximation process, which strips database-specific operators, controlled vocabularies, and
citation-based expansion techniques rather than shortcomings of the original strategies. Dense
retrieval, in contrast, delivers substantially higher recall (0.702) and the strongest ranking efectiveness
(MAP 0.281, P@10 0.545), demonstrating robustness to vocabulary mismatch and terminological
variation in metadata-constrained settings. The zero-shot GPT-4.1- Mini Boolean baseline performs poorly
across all metrics (recall 0.099), providing a clear negative result: naive zero-shot Boolean generation
with an of-the-shelf model fails to recover either the coverage or selectivity of expert-designed queries,
indicating the need for more structured and adaptive query construction approaches. Overall, these
results establish SR4CS as a targeted diagnostic benchmark for exposing precision–recall trade-ofs and
failure modes in retrieval and automated query generation.</p>
      <p>
        Several directions for future extensions of SR4CS are possible. First, deeper integration with open
scholarly infrastructures such as OpenAlex would support reference coverage, metadata enrichment,
and query execution within an open and reproducible ecosystem, enabling end-to-end studies that rely
on open identifiers and transparent metadata. Second, extending the corpus-construction workflow
to additional domains beyond computer science (e.g., education and social science) would broaden
the benchmark’s scope and help assess how retrieval and query-generation methods transfer across
diferences in terminology and reporting conventions. Third, future work could explore agentic
LLMbased information extraction frameworks that use iterative feedback and explicit grounding in source
text to improve extraction fidelity and consistency [
        <xref ref-type="bibr" rid="ref17">17, 18</xref>
        ]; such approaches may also make it easier to
attach verifiable evidence to extracted fields and to diagnose error modes in downstream evaluation.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>We introduced SR4CS, a large-scale test collection of systematic reviews in computer science, designed
to support reproducible research on Boolean query generation, retrieval, and screening. By providing
approximated versions of expert-designed Boolean queries, curated reference pools, and a controlled
evaluation setting over titles and abstracts, SR4CS enables systematic comparisons of retrieval paradigms
and query formulation strategies beyond the biomedical domain. The dataset is released with
documentation and code to facilitate reuse, benchmarking, and future work on systematic review automation,
retrieval evaluation, and query generation in computer science and related fields.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used DeepL Write for sentence polishing and grammar
correction.
2510.00276. arXiv:2510.00276.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Liberati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Altman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tetzlaf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mulrow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gøtzsche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ioannidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Devereaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kleijnen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moher</surname>
          </string-name>
          ,
          <article-title>The prisma statement for reporting systematic reviews and meta-analyses of studies that evaluate health care interventions: Explanation and elaboration</article-title>
          ,
          <source>PLoS Med</source>
          . (
          <year>2009</year>
          ). doi:doi:10.1371/journal.pmed.
          <volume>100010</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Lamé</surname>
          </string-name>
          ,
          <article-title>Systematic literature reviews: An introduction</article-title>
          ,
          <source>Proc. of Design Soc.: Int. Conf. on Engineering Design</source>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .1017/dsi.
          <year>2019</year>
          .
          <volume>169</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lefebvre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Glanville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Briscoe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Littlewood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Marshall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.-I.</given-names>
            <surname>Metzendorf</surname>
          </string-name>
          , A. NoelStorr, T. Rader,
          <string-name>
            <given-names>F.</given-names>
            <surname>Shokraneh</surname>
          </string-name>
          , J. Thomas,
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Wieland</surname>
          </string-name>
          ,
          <article-title>Searching for and selecting studies</article-title>
          , John Wiley &amp; Sons, Ltd,
          <year>2019</year>
          . doi:https://doi.org/10.1002/9781119536604.ch4. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781119536604.ch4.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>MacFarlane</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
            Russell-Rose,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Shokraneh</surname>
          </string-name>
          ,
          <article-title>Search strategy formulation for systematic reviews: Issues, challenges and opportunities</article-title>
          , Intel. Sys. with
          <string-name>
            <surname>Applications</surname>
          </string-name>
          (
          <year>2022</year>
          ). doi:https://doi.org/10.1016/j.iswa.
          <year>2022</year>
          .
          <volume>200091</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Scells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          ,
          <article-title>A comparison of automatic boolean query formulation for systematic reviews</article-title>
          , Inf. Retr. J. (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1007/S10791-020-09381-1.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Scells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          , G. Zuccon,
          <article-title>Can chatgpt write a good boolean query for systematic review literature search?</article-title>
          ,
          <source>in: Proc. of SIGIR</source>
          <year>2023</year>
          , ACM,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .1145/3539618.3591703.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Scells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          , G. Zuccon,
          <article-title>Zero-shot generative large language models for systematic review screening automation</article-title>
          ,
          <source>in: Proc. of ECIR</source>
          <year>2024</year>
          , LNCS, Springer,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>031</fpage>
          -56027-9\_
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Sami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Rasheed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kemell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Waseem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kilamo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nguyen-Duc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Systä</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Abrahamsson</surname>
          </string-name>
          ,
          <article-title>System for systematic literature review using multiple AI agents: Concept and an empirical evaluation</article-title>
          ,
          <source>CoRR</source>
          (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .48550/ARXIV.2403.08399. arXiv:
          <volume>2403</volume>
          .
          <fpage>08399</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Scells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Deacon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Geva</surname>
          </string-name>
          ,
          <article-title>A test collection for evaluating retrieval of studies for inclusion in systematic reviews</article-title>
          ,
          <source>in: Proc. of SIGIR</source>
          <year>2017</year>
          , ACM,
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .1145/3077136.3080707.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kanoulas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Spijker</surname>
          </string-name>
          ,
          <article-title>CLEF 2019 technology assisted reviews in empirical medicine overview</article-title>
          , in: W.N.
          <source>of CLEF</source>
          <year>2019</year>
          ,
          <article-title>CEUR-WS</article-title>
          .org,
          <year>2019</year>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2380</volume>
          /paper_250.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Scells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Zuccon, From little things big things grow: A collection with seed studies for medical systematic review literature search</article-title>
          ,
          <source>in: Proc. of SIGIR</source>
          <year>2022</year>
          , ACM,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .1145/3477495. 3531748.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Polak</surname>
          </string-name>
          , D. Morgan,
          <article-title>Extracting accurate materials data from research papers with conversational language models and prompt engineering - example of chatgpt</article-title>
          ,
          <source>CoRR</source>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .48550/ARXIV.2303.05352. arXiv:
          <volume>2303</volume>
          .
          <fpage>05352</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Gartlehner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kahwati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hilscher</surname>
          </string-name>
          , I. Thomas,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kugley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Crotty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viswanathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nussbaumer-Streit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Booth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Erskine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Konet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Chew</surname>
          </string-name>
          ,
          <article-title>Data extraction for evidence synthesis using a large language model: A proof-of-concept study</article-title>
          , Research Synthesis Methods (
          <year>2024</year>
          ). doi:https://doi.org/10.1002/jrsm.1710. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/jrsm.1710.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Backes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Iurshina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Shahid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mayr</surname>
          </string-name>
          ,
          <article-title>Comparing free reference extraction pipelines</article-title>
          ,
          <source>Int. J. Digit. Libr</source>
          . (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .1007/S00799-024-00404-6.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Poznanski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rangapur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Borchardt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dunkelberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Huf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rangapur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wilhelm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          , L. Soldaini,
          <article-title>olmocr: Unlocking trillions of tokens in pdfs with vision language models</article-title>
          ,
          <year>2025</year>
          . doi:
          <volume>10</volume>
          .48550/ARXIV.2502. 18443. arXiv:
          <volume>2502</volume>
          .
          <fpage>18443</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Lai</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Estill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <article-title>Language models for data extraction and risk of bias assessment in complementary medicine</article-title>
          ,
          <source>npj Digit</source>
          .
          <source>Medicine</source>
          (
          <year>2025</year>
          ). doi:
          <volume>10</volume>
          .1038/S41746-025-01457-W.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Dual-llm adversarial framework for information extraction from research literature</article-title>
          , bioRxiv (
          <year>2025</year>
          ). URL: https: //www.biorxiv.org/content/early/2025/09/16/
          <year>2025</year>
          .09.11.675507.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>