<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>From Philosophy to NLU: Evolving Definitions of Research Hypotheses</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>JianWu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SarahRajtmaje</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Old Dominion University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The Pennsylvania State University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the theyrpmothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe diferences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as we move toward a machine-interpretable scholarly record.</p>
      </abstract>
      <kwd-group>
        <kwd>natural language processing</kwd>
        <kwd>natural language understanding</kwd>
        <kwd>natural language inference</kwd>
        <kwd>hypothesis extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The word “hypothesis” has been used variably with diferent meanings, over decades and centuries,
across the social, natural, and formal sciences—from its conceptual roots in ancient Greek philosophy
[
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ] to the development of hypothesis testing as statistical me4t,h5o]dasn[d subsequent evolution
of the term in diferent fields reflecting their unique questions and approac6h]e.sO[f course, ambiguity
around language and variable use of terminology is pervasive within and outside of science. Language
has always been an impoverished tool for representation and expression of comple7x, i8d,e9a]s. [
In many cases though, this is not a problem. Members of a particular community develop shared
understanding of the meaning of a given term in context, and this allows them to communicate
efectively toward collective goals.
      </p>
      <p>We argue that ambiguity and variability around the definition of hypotheses, which was once
acceptable—even perhaps productive—is now a critical concern in light of natural language processing
(NLP) and natural language understanding (NLU) tasks requiring quantitative operationalization of
hypotheses, in particular, hypothesis extraction (detection/identification), verification, and generation.
Emerging technical work in these fields often do not include explicit definitions of hypotheses, claims,
or evidence, instead relyindegfacto on benchmark datasets to provide implicit definitions.</p>
      <p>Following, we survey the definitions and operationalizations of hypotheses, focusing on research
hypotheses engaged in the hypothesis mining literature. Research hypotheses are hypotheses designed
for systematic investigation within a research framework. In this paper, we do not distinguish between
research hypothesis and scientific hypothesis. In principle, these two terms have diferent scopes, but in
practice, they are often used interchangeably in modern hypothesis mining papers. The related tasks are
particularly in the areas of natural language inference (NLI), hypothesis extraction, scientific hypothesis
evidencing and scientific claim verification, and scientific hypothesis generation. We highlight important
Japan
∗The two authors made equal contributions to this paper.
https://www.cs.odu.edu/~jwu(/J. Wu); https://www.rajtmajerlab.n(eSt./Rajtmajer)</p>
      <p>CEUR</p>
      <p>ceur-ws.org
diferences and discuss the challenges these diferences impose on knowledge assembly and aggregation.
Our work is motivated by the vision of a computable scholarly record—a verifiable and extensible
knowledge base synthesizing computationally and data-enabled disco1v0e]r. ieTsh[is vision, we
suggest, will be enabled by machine-readable hypotheses well-structured in predictable formats.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <sec id="sec-2-1">
        <title>2.1. Conceptual origins</title>
        <p>The wordhypothesis derives from Greek and means, literaallpyu,tting under or supposition. Ancient
Greek philosophers used the term to describe a foundational assumption upon which to build out
further reasoning. Plato uses the term in several of his dialogues with this intention–namely, as a claim
accepted temporarily in order to explore its implications. Of particular interest to Plato was whether a
hypothesis could support consistent and coherent conclu1s1io,1n2s][. Aristotle also engaged with the
term. Aristotle viewed hypotheses more cautiously, being skeptical of relying on hypotheses without
empirical verification. He delineated tentative assumptions from axioms or first principles, insisting
that scientific knowledge must be based on demonstrated causes, not just assumed pre13m]i.ses [</p>
        <p>During the scientific revolution of the 16th and 17th centuries, the concept of hypothesis evolved
significantly. Galileo and Newton began using hypotheses as formal components of the scientific method,
emphasizing the importance of testing through observation and exper1i4m, e1n5]t. [René Descartes
also contributed to this shift, promoting skepticism and the formulation of testable prop1o6,si1t7i]o. ns [</p>
        <p>
          By the 19th and 20th centuries, the hypothesis had become a central pillar in science. Prominent
philosophers of science Karl Popper and Thomas Kuhn centered hypotheses within the scientific process,
both agreeing that hypotheses must engage with empirical data in some way, i.e., should be testable,
observable1[
          <xref ref-type="bibr" rid="ref19 ref8">8, 19</xref>
          ]. However, Kuhn and Popper’s views on scientific advancement difered in important
ways and their respective views on the role of hypotheses reflected these diferences. Kuhn, a historian,
viewed periods of science through the lenspaorfadigms and hypotheses as statements that operates
within a paradigm (vs. free-floating assumptions to be directly tested in isol2a0t]i.oPno)p[per, on
the other hand, highlighted the asymmetry between verificationfalasnificdation —hypotheses cannot
be proven true, only proven false. Central to Popper’s thinking is his assertion that confirmations
for a theory are easy to find if we look for them. Confirming evidence should only count when it is
the result of a genuine test of the theory (i.e., we conclude that the theory withstood an attempt to
disconfirm it) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. In essence, Popper argued that Kuhn’s hypotheses risked self-fulfillment. Popperian
theories underlie the current open science movement triggered by the replication crisis, including eforts
promoting development of strontegst,able hypotheses, and preregistration to delineate exploratory vs.
confirmatory findings [
          <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Modern forms</title>
        <p>In the last decades, there has been a growing interest in hypothesis mining in scientific literature, mostly
in the fields of NLP and NLU, but also in interdisciplinary fields between social science and AI. The
exact definitions of hypotheses involved are not always provided in the context of research problems,
and the specific forms and expressions vary across papers. Here, we categorize modern hypotheses
into several types, which may deviate from traditional definitions2,3e]..g., [</p>
        <p>
          Ideas as hypotheses. As defined in Kuhn &amp; Hawkins [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], an idea is a realization or hypothesis
that can challenge and shift paradigms within a scientific community. Several recent papers about
hypothesis generation adopted this conceptualization and treated ideas as hyp2o4t,h2e5s]e.sI n[
Kumar et al.2[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], the authors build a dataset contaiFnuitnugre Research Ideas (FRIs) and then consider
the generated FRIs to be hypotheses. The structure of these ideas includes: premises; a traditional
research hypothesis; and its context. In Wang et25a]l,.t[he authors do not discern the term ideas from
hypotheses and use them interchangeably in certain contexts, but the ground truth data shows that the
an astronaut is a kind of human
        </p>
        <p>a human is a kind of animal
Question: Why do astronauts need oxygen in the backpacks of their spacesuits?
an animal requires oxygen to breathe
an astronaut is a kind of animal
a vacuum does not contain oxygen
space is a vacuum
an astronaut requires oxygen</p>
        <p>to breath
spacesuit backpacks contain</p>
        <p>oxygen
there is no oxygen in space
Answer: to help astronauts breathe in outer space</p>
        <p>H: an astronaut
requires the
oxygen in a
spacesuit
backpack to
breath
generated content contains preliminary and broad notations intended to inspire further investigation,
which is aligned with the concept of ideas. An example is shown in T1.able</p>
        <p>
          Claims as hypotheses. The classical definition ocflaim is the conclusion or assertion that you want
your audience to accep2t6[]. Adopting this definition, in scientific literature, a claim can be defined as a
specific assertion reported as a finding of the paper. A paper can make more than one claim, and a claim
may contain one or multiple sentences. One definition of hypothesis is a claim that has not been tested
[27]. In Alipourfard et al2.8[], authors labelcalaim trace for each paper in their corpora, and each
trace contains four claims. Hypotheses and evidence are treated as two types of claims. In the recent
SciHyp dataset2[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] developed for hypothesis detection and classification in Computer Science papers,
many hypotheses in the ground truth are claims manually extracted throughout the full text. This
ambiguity also occurs in scientific hypothesis evidencing (SH3E0;])[and scientific claim verification
(SCV; [31]) tasks. Both tasks aim to discern the relationship (or stance) between a hypothesis (in SHE,
or a claim in SCV) and a candidate piece of evidence.
        </p>
        <p>Hypothesis-proposals. In recent works about hypothesis generation, models are built to
generate not only a hypothesis but also a series of related sections such as its background, justification,
and test procedures, resulting ipnroaposal-style document32[]. The hypothesis-proposal increases
the transparency of hypothesis generation and provides a guide for testing. However, the specific
format/sections of the proposal difer by model. An example3i2n] i[s shown in Table1.</p>
        <p>Formal expressions. In early work, a research hypothesis is broken down into three dimensions,
namely contexts, variables, and relations1h6i]p.sE[ach hypothesis is associated with a target variable
and a set of independent variables, and relationships refer to the interactions between a given set of
variables under a given context that produces the hypothesis. A hypothesis is then naturally expressed
with asemantic tree in which the nodes represent variables and the edges represent relationships.</p>
        <p>
          In more recent works in NLU, papers have expressed research hypotheses in various ways, depending
on the focal tasks. For example, in hypothesis generation tasks, the generated hypothesis may be
composed of multiple declarative statements, in which one serves as the main hypothesis and the others
provide additional context or details3(3s]eaen[d [25] in Table1). In the hypothesis evidencing task,
hypotheses can be written as questio3n0s],[which can be converted into hypotheses in declarative
form. A research hypothesis can also be decomposed as a question and an answer, e.g., the SciTail
dataset3[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Their entailment relation can be further explained useinntagilamnent tree showing how
the hypothesis follows from the text cor3p5u]s([Figure1).
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Scientific hypothesis-related tasks and datasets</title>
      <sec id="sec-3-1">
        <title>3.1. Natural language inference</title>
        <p>
          Natural language inference (NLI, i.e., recognizing textual entailment (RTE)) involves assessing whether
a given textual premise entails or implies a given hypoth36e,si3s7[]. Most NLI datasets, such as SNLI
[38] and RTE-6 [39], are in open domains3[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. SciTail is one of the few datasets built for scientific NLI
[34]. Hypotheses are expressed in single declarative sentences (see 1T)a.ble
        </p>
        <p>Another scientific NLI datasetEisntailmentBank, built for multi-step scientific inference (Fig1u)r.e
The task is to generate an entailment tree given a hypothesis. The tree shows a hierarchical supportive
Question: Is there an association
between social media use and bad
mental health outcomes?</p>
        <p>studies indicating an association
Study1 Study2 Study3
studies indicating little or no association
Study4 Study5</p>
        <p>studies showing mixed evidence
Study6 Study7 Study8</p>
        <p>H1: There is an association between social media use and
bad mental health out- comes.</p>
        <p>H2: There is little or no association between social media
use and bad mental health out- comes.</p>
        <p>H3: There is little or no association between social media
use and bad mental health out- comes.
structure of claims toward a hypothesis. Other scientific NLI datasets include M4e0d]iaNnLdI [BioNLI
[41] in the medical and biomedical domains, and e-SNLI-3V8E] [for visual scientific inference.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Hypothesis and claim extraction</title>
        <p>
          Here, the goal is to automatically identify hypotheses from a scientific document. Because hypotheses
can be viewed as claims prior to testing, the input, output, and methods for extracting hypotheses and
claims are similar. The input document can be an abstract42,,e.4g3.,] [or a full paper2[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. In White et
al. [44], authors propose and apply a schema for annotating sentences in full text of scientific articles
into 9 types: hypothesis; goal; motivation; background; method; experiment; result; observation; and
conclusion4[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. In the dataset used for training DeepCause, a model for hypothesis extraction, selected
claims identified from the full text are labeled as hypoth2e3s]e. s [
        </p>
        <p>Claim extraction may benefit from structured abstracts which contain, e.g., a dFeidnidciantgsed
section (see Lancet45[]). Yet still, the specific sections difer across journals. For example4,2i]n, [
ifndings orproposed items are labeled as claims and one abstract may contain multiple claim1)s. (Table</p>
        <p>It is worth noting that papers in computer science and several other domains often claim
findings or contributions without explicitly stating hypothese2s5,]e..Wg.,e[consider these findings or
contributionpsseudo-hypotheses.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Scientific hypothesis evidencing and scientific claim verification</title>
        <p>Scientific hypothesis evidencing (SHE) 3[0] is the task of automatically identifying evidence from
scientific publications in support of or refutation of a given hypothesis. This task is similar to another
task called Scientific claim verification (SCV), both reflecting a model’s reasoning capability. The main
diference is that in SHE, the hypotheses are usually high-level research questions (2F)i.gIunrSeCV
datasets, the hypotheses are usually lower-level claims in specific contexts. However, certain cases
are in between. Both tasks can be divided into two subtasks: identifying evidence candidates from a
corpus of documents; and discerning the relationship between the hypothesis (claim) and an evidence
candidate. Most research focuses on the second subtask, in which the relationships are classified into
exclusive categories, nameSlUyPPORT, REFUTE, andNEI (not enough information) or their variants.</p>
        <p>
          In Koneru et al.30[] authors build a dataset for the task of SHE using community-driven annotations
of studies in social sciences. The input is a hypothesis and an abstract (i.e., candidate evidence), and the
output is a label indicating whether the abstract entails, contradicts, or is inconclusive to the hypothesis.
In Wadden et al.3[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], authors build an SCV dataset, the input of which is a (claim, abstract) pair (see an
example in Table1). In the Covid-Fact datase4t6][, claims are obtained by filtering titles of social media
posts. While, in the HealthVer dataset, claims are manually extracted from questions and snippets
returned by search engine4s7[].
        </p>
        <p>
          DiscoveryBench 4[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is a benchmark designed for a task caldlaetda-driven discovery. Similar to SCV,
the goal is to verify a hypothesis, originally expressed in the form of a research question. Nevertheless,
instead of using abstracts as evidence, the verification is grounded in data. In an example task, a model is
given a dataset in the form of a spreadsheet and a research question. The model is expected to generate
a stepwise workflow that tries to answer the question using the given data. The final output, treated
as a hypothesis, is decomposed into a semantic tree containing (context, variable, relationship) and
Input keywords: heat transfer performance, soft lithography, etc.
        </p>
        <p>Output hypothesis: We hypothesize that integrating biomimetic materials with
microfluidic chips will significantly enhance their heat transfer performance
and biocompatibility, making them ideal for advanced biomedical applications.</p>
        <p>Specifically, we propose that the lamellar structure of biomaterials, inspired by
keratin scales, can be engineered into microfluidic chips using soft lithography
techniques to improve their mechanical behavior and heat transfer eficiency
under cyclic loading conditions.</p>
        <p>Output other sections: outcome, mechanisms, design principles, unexpected
properties, comparison, and novelty, and their expanded versions.</p>
        <p>Input: Tweet pairs in the Tweet Popularity dataset [51].</p>
        <p>Output: Tweets with named entities like people, places, or organizations tend
to get more retweets by being more specific.</p>
        <p>Input seed term: diverse relational edge embedding
Input background: the task of converting a natural lanuage question into
an executable sql query , known as text - to -sql, is an important branch of
semantic parsing . the state - of - the -art graph -based encoder has been
successfully used in this task but does not model the question syntax well.</p>
        <p>Output: We propose a novel graph-based encoder that uses a diverse relational
edge embeddings to model the question syntax.</p>
        <p>Gen</p>
        <p>Keywords
(SciAgents [33])</p>
        <p>Proposal
Gen
compared against the gold standard hypothesis (see an example in1T).aSbclieClaimHunt is a dataset
recently built for SCV, in which a small amount of claims are manually extracted from the discussion
and conclusion sections of research papers in computer science. Most claims are generated by LLMs. Of
note, a fraction of claims are not self-contained and require reference to the context of the source paper.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Scientific hypothesis generation</title>
        <p>The goal of scientific hypothesis generation is to automatically create new, testable scientific hypotheses
or research ideas that identify novel relationships, phenomena, or gaps in existing kno2w4]l.edge [
Recent advancements in LLMs, e.g., Llam5a2][and GPT [53], ofer promise. Here too, existing literature
is inconsistent with respect to task formulatioinnp,ui.te.a, ndoutput (see examples in Table1).</p>
        <p>
          We identify four types of input for the hypothesis generation tasks:
(1) keywords, concise descriptions of the topics or central concepts of the model, such as in 2S5c]iMon [
and SciAgents3[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ];
(2) goals, brief discourse outlining research goals. In AI co-scien32t]i,stgo[als can be a request to
propose a novel hypothesis, suggest special requirements, or ask a question. In Si54e]t, aLLl.M[s are
provided “topics”, such as “novel prompting methods that can better quantify uncertainty or calibrate
the confidence of large language models”, which serve as goals. In Pu e5t5a],la.n[ input is described
as an “objective”, which is equivalent to a goal;
(3) data, i.e., a dataset, based on which the model is requested to generate a hypothes5is0,]e;.g., [
(4) background, context, rationale, or theoretical foundation of a hypothesis. For example, in the SciMon
framework 2[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], the input includes seed terms, including concepts and keywords, and background
context, which contains problems, motivations, or focus points. The Mamba fram56e]wuosreks [the
same ground truth as SciMo2n5][. The MOOSE framework 5[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] uses background and inspiration
derived from the raw web corpus as the input.
        </p>
        <p>Likewise, we identify three types of output of hypothesis generation tasks:
(1) traditional hypothesis, usually expressed as a single or multiple declarative sentence5s0,,e5.g8.],. [
(2) ideas, enriched hypotheses as shown in Secti2o.n2. For example, inSciMon, generated ideas may
contain claims, methods, and objectives extracted from abstracts.
(3) hypothesis proposals, comprehensive structured hypothesis description documents. For example, the
output of SciAgent3s3[] is a document containing hypotheses, outcomes, mechanisms, design principles,
unexpected properties, comparison, and novelty, each having its expanded version. AI co-sc3i2e]ntist [
also outputs a structured document but with diferent sections: introduction, recent findings, related
research, rationale, specificity, experimental design, and validation. Si5e4t]raelq. u[est LLM agents
to generate an “idea”, containing several components (e.g., problem, existing methods, motivation,
proposed method, experiment plan), similar to a research proposal. Whereas, the Piflow fram5e5w] ork [
requires LLM agents to generate a “hypothesis structure” consisting of rationale, hypothesis, reiterate,
and an experimental candidate.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion and Conclusions</title>
      <p>Our work here intends to be more descriptive than prescriptive. We outline the various definitions
and instances of hypotheses in existing scientific literature (and beyond). In particular, we focus on
definitions of hypotheses and related concepts in recent work in NLU.</p>
      <p>We hope that this work may raise awareness within the hypothesis mining community about
standardization of corpus-level tools, e.g., knowledge graphs representing connections amongst interdisciplinary
hypotheses, or hypothesis generation models across multiple domains. In lieu of standardization, the
inclusion of explicit, clear definitions of hypotheses (formal, where possible) could improve alignment
and assembly.</p>
      <p>For more than two decades, many in the research community have advocated for open data, open
materials, preregistration, and other best practices as central to thesevairsciohanbolef aand interpretable
scholarly record. With recent technological advances, this vision is on the horizon. The ultimate
goals of an interpretable scholarly record are: robust and eficient scientific progress; thoughtful
allocation of community resources toward important open problems; and honest dialogue with public
and policymakers. How–precisely–a queryable scholarly corpus comes together is an open question.
Here, we suggest that dissemination of clear, consistent, well-specified machine-readable hypotheses,
claims, and evidence are critical to this mission.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This project is partially supported by Open Philanthropy.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>In accordance with the principles of responsible AI use, we disclose that generative AI was used solely
for language editing and did not contribute to the scientific content or analysis presented in this work.
research ideas?, CoRR abs/2409.06185 (2024). URLh:ttps://doi.org/10.48550/arXiv.2409.061.85
doi:10.48550/ARXIV.2409.06185. arXiv:2409.06185.
[25] Q. Wang, D. Downey, H. Ji, T. Hope, SciMON: Scientific Inspiration Machines Optimized for
Novelty, in: L. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of
the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok,
Thailand, August 11-16, 2024, Association for Computational Linguistics, 2024, pp. 279–299. URL:
https://doi.org/10.18653/v1/2024.acl-long..1d8oi:10.18653/V1/2024.ACL-LONG.18.
[26] S. E. Toulmin, The uses of argument, Cambridge university press, 2003.
[27] T. Heger, A. Algergawy, M. Brinner, J. M. Jeschke, B. König-Ries, D. Mietchen, S. Zarrieß,
Natural language hypotheses in scientific papers and how to tame them: Suggested steps for
formalizing complex scientific claims, in: Robust Argumentation Machines: First
International Conference, RATIO 2024, Bielefeld, Germany, June 5–7, 2024, Proceedings,
SpringerVerlag, Berlin, Heidelberg, 2024, p. 3–19. URhLt:tps://doi.org/10.1007/978-3-031-63536-6_.1
doi:10.1007/978-3-031-63536-6_1.
[28] N. Alipourfard, B. Arendt, D. M. Benjamin, N. Benkler, M. Bishop, M. Burstein, M. Bush, J. Caverlee,
Y. Chen, C. Clark, et al., Systematizing confidence in open research and evidence (score), SocArXiv
(2021).
[29] R. Vasu, C. Sarasua, A. Bernstein, Scihyp: A fine-grained dataset describing hypotheses and their
components from scientific articles, in: International Semantic Web Conference, Springer, 2024,
pp. 134–152.
[30] S. D. Koneru, J. Wu, S. Rajtmajer, Can large language models discern evidence for scientific
hypotheses? case studies in the social sciences, in: N. Calzolari, M. Kan, V. Hoste, A. Lenci,
S. Sakti, N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational
Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20-25 May, 2024, Torino,
Italy, ELRA and ICCL, 2024, pp. 2787–2797. URLh:ttps://aclanthology.org/2024.lrec-main..248
[31] D. Wadden, S. Lin, K. Lo, L. L. Wang, M. van Zuylen, A. Cohan, H. Hajishirzi, Fact or fiction:
Verifying scientific claims, in: B. Webber, T. Cohn, Y. He, Y. Liu (Eds.), Proceedings of the
2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online,
November 16-20, 2020, Association for Computational Linguistics, 2020, pp. 7534–7550. URL:
https://doi.org/10.18653/v1/2020.emnlp-main.6.0d9oi:10.18653/V1/2020.EMNLP-MAIN.609.
[32] J. Gottweis, W. Weng, A. N. Daryin, T. Tu, A. Palepu, P. Sirkovic, A. Myaskovsky, F. Weissenberger,
K. Rong, R. Tanno, K. Saab, D. Popovici, J. Blum, F. Zhang, K. Chou, A. Hassidim, B. Gokturk,
A. Vahdat, P. Kohli, Y. Matias, A. Carroll, K. Kulkarni, N. Tomasev, Y. Guan, V. Dhillon, E. D.
Vaishnav, B. Lee, T. R. D. Costa, J. R. Penadés, G. Peltz, Y. Xu, A. Pawlosky, A. Karthikesalingam,
V. Natarajan, Towards an AI co-scientist, CoRR abs/2502.18864 (2025). UhtRtLp:s://doi.org/10.
48550/arXiv.2502.18864. doi:10.48550/ARXIV.2502.18864. arXiv:2502.18864.
[33] A. Ghafarollahi, M. J. Buehler, Sciagents: Automating scientific discovery
through bioinspired multi-agent intelligent graph reasoning, Advanced
Materials 37 (2025) 2413523. URL: https://advanced.onlinelibrary.wiley.com/doi/
abs/10.1002/adma.202413523. doi:https://doi.org/10.1002/adma.202413523.
arXiv:https://advanced.onlinelibrary.wiley.com/doi/pdf/10.1002/adma.202413523.
[34] T. Khot, A. Sabharwal, P. Clark, SciTaiL: A Textual Entailment Dataset from Science Question
Answering, Proceedings of the AAAI Conference on Artificial Intelligence 32 (2018)h.UttRpLs::
//ojs.aaai.org/index.php/AAAI/article/view/12.0d2o2i:10.1609/aaai.v32i1.12022.
[35] B. Dalvi, P. Jansen, O. Tafjord, Z. Xie, H. Smith, L. Pipatanangkura, P. Clark, Explaining answers
with entailment trees, in: M. Moens, X. Huang, L. Specia, S. W. Yih (Eds.), Proceedings of the
2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual
Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Association for Computational
Linguistics, 2021, pp. 7358–7370. URL:https://doi.org/10.18653/v1/2021.emnlp-main.5.8d5oi:10.
18653/V1/2021.EMNLP-MAIN.585.
[36] I. Dagan, D. Roth, F. Zanzotto, M. Sammons, Recognizing textual entailment: Models and
applications, Springer Nature, 2022.
[37] S. Storks, Q. Gao, J. Y. Chai, Recent advances in natural language inference: A survey of benchmarks,
resources, and approaches, arXiv preprint arXiv:1904.01172 (2019).
[38] V. Do, O.-M. Camburu, Z. Akata, T. Lukasiewicz, e-snli-ve: Corrected visual-textual entailment
with natural language explanations, arXiv preprint arXiv:2004.03744 (2020).
[39] L. Bentivogli, P. Clark, I. Dagan, D. Giampiccolo, The fith pascal recognizing textual entailment
challenge., TAC 7 (2009) 1.
[40] C. Shivade, Mednli—a natural language inference dataset for the clinical domain, (No Title) (2017).
[41] M. Bastan, M. Surdeanu, N. Balasubramanian, BioNLI: Generating a biomedical NLI dataset using
lexico-semantic constraints for adversarial examples, in: Y. Goldberg, Z. Kozareva, Y. Zhang
(Eds.), Findings of the Association for Computational Linguistics: EMNLP 2022, Association for
Computational Linguistics, Abu Dhabi, United Arab Emirates, 2022, pp. 5093–5104.hUttRpLs::
//aclanthology.org/2022.findings-emnlp.3.7d4o/i:10.18653/v1/2022.findings-emnlp.374.
[42] T. Achakulvisut, C. Bhagavatula, D. Acuna, K. Kording, Claim extraction in biomedical publications
using deep discourse model and transfer learning, arXiv preprint arXiv:1907.00962 (2019).
[43] X. Wei, M. R. U. Hoque, J. Wu, J. Li, Claimdistiller: Scientific claim extraction with supervised
contrastive learning, in: Proceedings of Joint Workshop of the 4th Extraction and Evaluation of
Knowledge Entities from Scientific Documents (EEKE2023) and the 3rd AI + Informetrics (All2023)
co-located with the JCDL 2023, Santa Fe, New Mexico, United States, June 26 – June 30, 2023, 2023,
pp. 3487–3496. URL: https://ceur-ws.org/Vol-3451/paper11.p.df
[44] E. White, K. B. Cohen, L. Hunter, The CISP annotation schema uncovers hypotheses and
explanations in full-text scientific journal articles, in: Proceedings of BioNLP 2011 Workshop, BioNLP ’11,
Association for Computational Linguistics, USA, 2011, p. 134–135.
[45] Lancet, Information for Authorhst,tps://www.thelancet.com/pb-assets/Lancet/authors/
tl-info-for-authors-1740074875577.p,d2f025.
[46] A. Saakyan, T. Chakrabarty, S. Muresan, COVID-fact: Fact extraction and verification of real-world
claims on COVID-19 pandemic, in: C. Zong, F. Xia, W. Li, R. Navigli (Eds.), Proceedings of the 59th
Annual Meeting of the Association for Computational Linguistics and the 11th International Joint
Conference on Natural Language Processing (Volume 1: Long Papers), Association for
Computational Linguistics, Online, 2021, pp. 2116–2129. URhLt:tps://aclanthology.org/2021.acl-long..165/
doi:10.18653/v1/2021.acl-long.165.
[47] M. Sarrouti, A. Ben Abacha, Y. Mrabet, D. Demner-Fushman, Evidence-based fact-checking
of health-related claims, in: M.-F. Moens, X. Huang, L. Specia, S. W.-t. Yih (Eds.), Findings of
the Association for Computational Linguistics: EMNLP 2021, Association for Computational
Linguistics, Punta Cana, Dominican Republic, 2021, pp. 3499–3512. UhRtLt:ps://aclanthology.org/
2021.findings-emnlp.297./doi:10.18653/v1/2021.findings-emnlp.297.
[48] B. P. Majumder, H. Surana, D. Agarwal, B. D. Mishra, A. Meena, A. Prakhar, T. Vora, T. Khot,
A. Sabharwal, P. Clark, Discoverybench: Towards data-driven discovery with large language
models, in: The Thirteenth International Conference on Learning Representations, ICLR 2025,
Singapore, April 24-28, 2025, OpenReview.net, 2025. URhL:ttps://openreview.net/forum?id=
vyflgpwfJW.
[49] K. BahadarKhan, A. A Khaliq, M. Shahid, A morphological hessian based approach for retinal blood
vessels segmentation and denoising using region based otsu thresholding, PLOS ONE 11 (2016)
1–19. URL: https://doi.org/10.1371/journal.pone.01589.9d6oi:10.1371/journal.pone.0158996.
[50] Y. Zhou, H. Liu, T. Srivastava, H. Mei, C. Tan, Hypothesis generation with large language models,
in: L. Peled-Cohen, N. Calderon, S. Lissak, R. Reichart (Eds.), Proceedings of the 1st Workshop
on NLP for Science (NLP4Science), Association for Computational Linguistics, Miami, FL, USA,
2024, pp. 117–139. URL: https://aclanthology.org/2024.nlp4science-.1d.1o0i/:10.18653/v1/2024.
nlp4science-1.10.
[51] C. Tan, L. Lee, B. Pang, The efect of wording on message propagation: Topic- and
authorcontrolled natural experiments on Twitter, in: K. Toutanova, H. Wu (Eds.), Proceedings of
the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long
Papers), Association for Computational Linguistics, Baltimore, Maryland, 2014, pp. 175–185. URL:
https://aclanthology.org/P14-10.1d7o/i:10.3115/v1/P14-1017.
[52] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P.
Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. Canton-Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes,
J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan,
M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M. Lachaux, T. Lavril, J. Lee,
D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton,
J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan,
B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur,
S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, T. Scialom, Llama 2: Open foundation and
finetuned chat models, CoRR abs/2307.09288 (2023). URhLt:tps://doi.org/10.48550/arXiv.2307.092.88
doi:10.48550/ARXIV.2307.09288. arXiv:2307.09288.
[53] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam,
G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh,
D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark,
C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei, Language models are few-shot
learners, in: H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, H. Lin (Eds.), Advances in Neural
Information Processing Systems 33: Annual Conference on Neural Information Processing Systems
2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URhLt:tps://proceedings.neurips.cc/paper/
2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.h t.ml
[54] C. Si, D. Yang, T. Hashimoto, Can LLMs Generate Novel Research Ideas? A Large-Scale Human
Study with 100+ NLP Researchers, in: The Thirteenth International Conference on Learning
Representations, ICLR 2025, Singapore, April 24-28, 2025, OpenReview.net, 2025. UhRtLt:ps:
//openreview.net/forum?id=M23dTGWC Z.y
[55] Y. Pu, T. Lin, H. Chen, PiFlow: Principle-aware Scientific Discovery with Multi-Agent Collaboration,
CoRR abs/2505.15047 (2025). URL:https://doi.org/10.48550/arXiv.2505.150.4d7oi:10.48550/ARXIV.
2505.15047. arXiv:2505.15047.
[56] M. Chai, E. Herron, E. Cervantes, T. Ghosal, Exploring scientific hypothesis generation with
mamba, in: L. Peled-Cohen, N. Calderon, S. Lissak, R. Reichart (Eds.), Proceedings of the 1st
Workshop on NLP for Science (NLP4Science), Association for Computational Linguistics, Miami,
FL, USA, 2024, pp. 197–207. URL: https://aclanthology.org/2024.nlp4science-.1d.1o7i/:10.18653/
v1/2024.nlp4science-1.17.
[57] Z. Yang, X. Du, J. Li, J. Zheng, S. Poria, E. Cambria, Large language models for automated
opendomain scientific hypotheses discovery, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Findings of the
Association for Computational Linguistics: ACL 2024, Association for Computational Linguistics,
Bangkok, Thailand, 2024, pp. 13545–13565. URLh:ttps://aclanthology.org/2024.findings-acl..804/
doi:10.18653/v1/2024.findings-acl.804.
[58] B. Qi, K. Zhang, H. Li, K. Tian, S. Zeng, Z.-R. Chen, B. Zhou, Large language models are zero shot
hypothesis proposers, arXiv preprint arXiv:2311.05965 (2023).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hawkins</surname>
          </string-name>
          ,
          <article-title>The structure of scientific revolutions</article-title>
          ,
          <source>American Journal of Physics</source>
          <volume>31</volume>
          (
          <year>1963</year>
          )
          <fpage>554</fpage>
          -
          <lpage>555</lpage>
          . URL:https://doi.org/10.1119/1.196966.0 doi:10.1119/1.1969660. arXiv:https://pubs.aip.org/aapt/ajp/articlepdf/31/7/554/12111921/554_1_online.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Popper</surname>
          </string-name>
          , Science as falsification,
          <source>Conjectures and refutations 1</source>
          (
          <year>1963</year>
          )
          <fpage>33</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>I. Lakatos</surname>
          </string-name>
          ,
          <article-title>History of science and its rational reconstructions, in: PSA: Proceedings of the biennial meeting of the philosophy of science association</article-title>
          , volume
          <year>1970</year>
          , Cambridge University Press,
          <year>1970</year>
          , pp.
          <fpage>91</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>R. A</surname>
          </string-name>
          . Fisher,
          <article-title>Statistical methods for research workers, 5, Oliver</article-title>
          and Boyd,
          <year>1928</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Fisher</surname>
          </string-name>
          ,
          <article-title>Statistical methods and scientific induction</article-title>
          ,
          <source>Journal of the Royal Statistical Society Series B: Statistical Methodology</source>
          <volume>17</volume>
          (
          <year>1955</year>
          )
          <fpage>69</fpage>
          -
          <lpage>78</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Glass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hall</surname>
          </string-name>
          ,
          <article-title>A brief history of the hypothesis</article-title>
          ,
          <source>Cell</source>
          <volume>134</volume>
          (
          <year>2008</year>
          )
          <fpage>378</fpage>
          -
          <lpage>381</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Fauconnier</surname>
          </string-name>
          ,
          <article-title>Mappings in thought and language</article-title>
          , Cambridge University Press,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Malt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Majid</surname>
          </string-name>
          ,
          <article-title>How thought is mapped into words</article-title>
          ,
          <source>Wiley Interdisciplinary Reviews: Cognitive Science</source>
          <volume>4</volume>
          (
          <year>2013</year>
          )
          <fpage>583</fpage>
          -
          <lpage>597</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Yarkoni</surname>
          </string-name>
          ,
          <article-title>The generalizability crisis</article-title>
          ,
          <source>Behavioral and Brain Sciences</source>
          <volume>45</volume>
          (
          <year>2022</year>
          )
          <article-title>e1</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>V.</given-names>
            <surname>Stodden</surname>
          </string-name>
          ,
          <article-title>On emergent limits to knowledge-or, how to trust the robot researchers: A pocket guide</article-title>
          ,
          <source>Harvard Data Science Review</source>
          <volume>6</volume>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Karasmanēs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Karasmanis</surname>
          </string-name>
          ,
          <article-title>The hypothetical method in Plato's middle dialogues</article-title>
          ,
          <source>Ph.D. thesis</source>
          , University of Oxford,
          <year>1987</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H. W.</given-names>
            <surname>Ausland</surname>
          </string-name>
          ,
          <article-title>Socrates' dialectical use of hypothesis</article-title>
          , in: New Perspectives on Platonic Dialectic, Routledge,
          <year>2022</year>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Barnes</surname>
          </string-name>
          , Posterior analytics (
          <year>1994</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Galilei</surname>
          </string-name>
          ,
          <article-title>Dialogue concerning the two chief world systems</article-title>
          , Berkeley: University of California Press.·[1638] (
          <year>1914</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>I. Newton</surname>
          </string-name>
          ,
          <article-title>Philosophiae naturalis principia mathematica</article-title>
          , volume
          <volume>1</volume>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Brookman</surname>
          </string-name>
          ,
          <year>1833</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Descartes</surname>
          </string-name>
          ,
          <article-title>A discourse on method</article-title>
          ,
          <source>JM Dent &amp; Sons Limited</source>
          ,
          <year>1912</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sakellariadis</surname>
          </string-name>
          ,
          <article-title>Descartes's use of empirical data to test hypotheses</article-title>
          ,
          <source>Isis</source>
          <volume>73</volume>
          (
          <year>1982</year>
          )
          <fpage>68</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Quinn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Dunham</surname>
          </string-name>
          ,
          <article-title>On hypothesis testing in ecology and evolution</article-title>
          ,
          <source>The American Naturalist</source>
          <volume>122</volume>
          (
          <year>1983</year>
          )
          <fpage>602</fpage>
          -
          <lpage>617</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Leplin</surname>
          </string-name>
          ,
          <article-title>A novel defense of scientific realism</article-title>
          , Oxford University Press,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>D.</given-names>
            <surname>Shapere</surname>
          </string-name>
          ,
          <article-title>The structure of scientific revolutions</article-title>
          ,
          <source>The Philosophical Review</source>
          <volume>73</volume>
          (
          <year>1964</year>
          )
          <fpage>383</fpage>
          -
          <lpage>394</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Rajtmajer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Errington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. G.</given-names>
            <surname>Hillary</surname>
          </string-name>
          ,
          <article-title>How failure to falsify in high-volume science contributes to the replication crisis</article-title>
          ,
          <source>Elife</source>
          <volume>11</volume>
          (
          <year>2022</year>
          )
          <article-title>e78830</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Nosek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Ebersole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>DeHaven</surname>
          </string-name>
          , D. T. Mellor,
          <article-title>The preregistration revolution</article-title>
          ,
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>115</volume>
          (
          <year>2018</year>
          )
          <fpage>2600</fpage>
          -
          <lpage>2606</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abdullaev</surname>
          </string-name>
          , Deepcause:
          <article-title>Hypothesis extraction from information systems papers with deep learning for theory ontology learning (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ghosal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekbal</surname>
          </string-name>
          ,
          <article-title>Can large language models unlock novel scientific</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>