<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Do We Need Subject Matter Experts? A Case Study of Measuring Up GPT-4 Against Scholars in Topic Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kabir Manandhar Shrestha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>(Melbourne Data Analytics Platform)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katie Wood</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>(University of Melbourne Archives)</institution>
          ,
          <addr-line>David Goodman</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Melbourne</institution>
          ,
          <addr-line>Parkville VIC 3010, Melbourne</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Assessing the quality of topics extracted from large text datasets presents a significant challenge in the ifeld of computational social science. This research examines the effectiveness of coherence metrics, the GPT-4 model, and evaluations by subject matter experts (SMEs) using a set of speeches by former Australian Prime Minister Malcolm Fraser. Our primary objective was to analyze the evolution of Fraser's rhetoric. By comparing topics identified by coherence metrics and GPT-4 to those deemed meaningful by SMEs, we found that GPT-4 not only performs on par with traditional coherence metrics but also offers a scalable alternative for comprehensive topic evaluations. However, SMEs provide unparalleled depth and contextual understanding, proving indispensable in situations demanding meticulous accuracy. In situations where SMEs aren't available, our approach does show that GPT-4 can be employed for topic evaluation, albeit with some margin of error.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Computational Social Science</kwd>
        <kwd>Topic Modeling</kwd>
        <kwd>Latent Dirichlet Allocation (LDA)</kwd>
        <kwd>Coherence Metrics</kwd>
        <kwd>Large Language Models (LLMs)</kwd>
        <kwd>GPT-4</kwd>
        <kwd>Evaluation Metrics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The digital era, marked by the fusion of vast data sets and advanced computational techniques,
has instigated significant shifts in social and cultural research, ushering in the possibility of
novel approaches to studying human behavior and communication [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. Historically, within
Humanities and Social Sciences, close reading has been highly valued. This method focuses on
small passages, emphasizing lexical choice and structure, as well of course as analyzing entire
texts. However, with the integration of computational methods in Social Sciences, the term
distant reading emerged [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], emphasizing a holistic view of texts or document collections. Topic
modeling, in particular, has gained prominence as a method for this kind of analysis [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. The
adoption of these computational techniques has enabled humanists and social scientists to analyze
data on a much broader scale, offering the exciting possibility of entirely fresh perspectives on
human behavior and communication [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Central to this transformation is computational social science, which harnesses these techniques
to decode intricate patterns and trends previously elusive [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Within this domain, topic modeling
stands out as an adept tool for uncovering latent themes in extensive text corpora, shedding light
on cultural evolution, thematic dynamics over time, and its broader applications in computational
social science [
        <xref ref-type="bibr" rid="ref3 ref8 ref9">3, 8, 9</xref>
        ].
      </p>
      <p>
        Historically, topic evaluations have predominantly relied on coherence metrics, emphasizing
coherence, relevance, and interpretability. However, as the discipline evolves, there’s an increasing
awareness that these metrics may not capture the full depth or importance of certain topics [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and
that we need Subject Matter Experts (SMEs) for precision. While SMEs provide invaluable depth,
their engagement can be challenging due to availability and expense. In light of these challenges,
we explored the potential of GPT-4, a large language model, to emulate the discernment of SMEs
in topic evaluations as well.
      </p>
      <p>Our research is centered on the speeches of former Australian Prime Minister Malcolm Fraser,
with dual objectives: tracing the trajectory of Fraser’s rhetoric and providing a comprehensive
view of his political odyssey1. To this end, we employed topic models and evaluated their quality
through three distinct lenses: intrinsic coherence metrics, insights from subject matter experts
(SMEs), and the computational prowess of GPT-4. While this paper focuses on topic evaluation
and comparing coherence metrics and GPT-4 with SMEs, readers interested in a deeper dive into
our broader research question can refer to Section 8 for our dedicated website and additional
resources.</p>
      <p>Exploring the novel application of GPT-4 in topic evaluation, we examine its capabilities in
assessing topics. We compare GPT-4’s assessments with traditional measures of coherence and
judgments made by SMEs. Although GPT-4 emerges as a scalable alternative that mirrors the
effectiveness of coherence metrics, it occasionally misses subtle details that SMEs recognize.
Conversely, the automated analyses may pick up latent patterns of which the SME’s were unaware.
Based on these findings, we advocate for a collaborative approach to topic evaluation, emphasizing
the synergy of human expertise and automated insights. This approach provides researchers with
a roadmap for incorporating various evaluation techniques tailored to their specific challenges.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>Malcolm Fraser, the former Prime Minister of Australia, served from 1975 to 1983. Although
he was leader of the conservative Liberal Party of Australia, upon leaving office Fraser became
estranged from the party and highly critical of its policies and rhetoric., Fraser became a prominent
commentator on a range of issues, such as human rights, multiculturalism in Australia, Indigenous
affairs, and foreign policy.</p>
      <p>
        To make Fraser’s long and varied political life more accessible and digestible, this study
employed topic modeling to analyze his radio speeches. The radio speeches (given in english)
1https://library.unimelb.edu.au/asc/collections/highlights/collections/malcolmfraser
are a unique resource, in that they span his entire parliamentary career, from 1954 to 1983,
are relatively consistent in format and wide-ranging in subject matter. The team included the
curator of the Malcolm Fraser Collection, and an historian, who assessed the utility of the
induced topics solely based on their expert judgment. Topic modeling, rooted in text mining, has
been instrumental across fields from computational social science to digital humanities, enabling
scholars to identify latent patterns in vast textual data [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Central to the realm of topic modeling is the Latent Dirichlet Allocation (LDA) algorithm [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
LDA is a probabilistic model that operates on the assumption that each document consists of a
combination of different topics, and conversely, each topic is made up of various words. Through
careful analysis of how words are distributed across documents, LDA attempts to uncover the
underlying topics that could have generated these observed documents [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In the digital age,
with the rapid growth of vast document archives, the utility of LDA and other topic models
becomes even more pronounced, offering algorithmic solutions to manage, organize, and annotate
large text collections [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. LDA’s adaptability and robustness have cemented its position as a
cornerstone in topic modeling, proving its mettle across diverse datasets and research contexts
[
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ].
      </p>
      <p>
        Within computational disciplines, topic coherence scores are the accepted measure of topic
quality and semantic interpretability with the most common quantitative intrinsic evaluations
of topics being pointwise mutual information (PMI) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and normalised pointwise mutual
information (NPMI) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. There are many variations to measure topic coherence, for example
Mimno et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] employs log conditional probability (LCP) rather than the accepted (N)PMI,
Röder et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] proposes a measure, CV, that combines an indirect cosine measure with NMPI,
and Aletras and Stevenson [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] experiments with representing the topic words as context vectors
instead of their citation or surface form. Despite the variations, these measures commonly rely
on how closely associated the top n words are with each other according to a general reference
corpus [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In addition, rather than focusing on a specific topic, where a scholar or scholars may
have in-depth insights, like our Malcolm Fraser project, in determining measures of coherence
and topic quality, topics are often judged by crowd sourced participants [
        <xref ref-type="bibr" rid="ref17 ref20 ref21">20, 17, 21</xref>
        ]. However, in
our instance, we have SMEs in this field who would be more equipped in judging topics rather
than crowd-sourced workers. We aim to investigate the impact of incorporating SMEs in the topic
modeling process by comparing the results of their evaluations to those produced by common
quantitative measure.
      </p>
      <p>
        The last decade has witnessed a paradigm shift in the realm of natural language processing
(NLP). This evolution in language understanding can be traced back to the Turing Test’s inception
in the 1950s, with the journey transitioning from statistical models to the modern pre-trained
models that harness the Transformer architecture [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. These models, characterized by their
massive number of parameters and extensive training data, have set new benchmarks across
several NLP tasks [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. As these models scaled, they began to exhibit unique abilities, such as
in-context learning, leading to the term "large language models (LLM)" [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. The research surge
in LLMs has reinvigorated discussions on the potential of artificial general intelligence (AGI)
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], emphasizing the blend of human expertise and computational prowess in language modeling
[
        <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
        ].
      </p>
      <p>
        The GPT (Generative Pre-trained Transformer) series [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], developed by OpenAI, stands as
a testament to the rapid advancements in this domain. GPT-4 [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], the latest in this lineage, is
trained on extensive datasets, enabling it to produce human-like text with remarkable accuracy. Its
transformer-based architecture excels in capturing textual nuances and context. Beyond just text
generation, the potential applications of GPT-4 are vast. From serving as virtual assistants, aiding
in content creation, to more specialized tasks like medical diagnosis assistance and legal document
analysis, GPT-4’s prowess has been demonstrated across domains [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. For our research purpose,
we assess if GPT-4 can act as a subject matter expert. While doing so, we also assess if it fails to
align with Subject Matter Experts’ point of view and if it provides some misleading information
in our case-study as it has in the past [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Topic Analysis Framework</title>
      <sec id="sec-3-1">
        <title>3.1. Data and Objectives</title>
        <p>Our research aimed to capture the evolving political landscape in Australia by analyzing the
topics addressed by Malcolm Fraser over time. All speeches were delivered in English, reflecting
the primary language spoken by the Prime Minister of Australia. We focused on the speeches
from the decade spanning 1960 to 1970 because of the relatively low error rate in transcribed
documents during these years. This is because Subject Matter Experts (SMEs) were involved in
the transcription process.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Preprocessing for Quality Enhancement</title>
        <p>The 270 speeches analyzed comprised a total of 245,701 tokens. Prior to constructing the topic
models, we executed several preprocessing steps to enhance the quality of the results. Stop words
were eliminated to prevent them from influencing the topic formation process. Additionally, we
identified relevant bi-grams and tri-grams to preserve essential word combinations. Moreover,
we pruned words and phrases that appeared in less than 10 documents or were present in over
50% of the documents. These measures collectively aimed to enhance the representativeness and
clarity of the identified topics. After preprocessing, the total vocabulary size was 1425.
wool; industry; board; grower; woolgrower; marketing;
promotion; conference; levy
operation; enemy; viet cong; battalion; regiment; province;
village; task force; troop
school; education; university; technical; student;
assistance; high; science; study; training
concern; affect; problem; important; involve; opportunity;
importance; individual; return
operation; parliament; house; minister; day; party; question;
speaker; business; sit
scheme; argument; price; buy; fact; show; put; fund; market;
plan</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Topic Modelling Technique</title>
        <p>We employed Latent Dirichlet Allocation (LDA) for our topic modeling needs. LDA’s clear
interpretability, scalability, and broad acceptance in academic research made it a fitting choice
for analyzing Malcolm Fraser’s speeches. For our analysis, we utilized the MALLET toolkit2, a
prominent implementation of LDA, to generate three topic models: 30, 60, and 90 topics in a grid
search approach in our 270 speeches. We chose three models to provide a range of granularity in
the topics, allowing for a more comprehensive analysis of the corresponding topics discussed by
Fraser over time.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Model Selection</title>
        <p>Selecting the optimal topic model was a crucial step in our methodology. To ensure an informed
decision, we enlisted the expertise of SMEs to assess the three generated models. The evaluation
process involved analyzing the coherence and diversity of each topic model. Ultimately, based on
their assessments, the SMEs concluded that the model with 60 topics best suited the content of
Fraser’s speeches, striking a balance between granularity and comprehensiveness.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Expert-Based Topic Categorization: Foundation for Comparative</title>
      </sec>
      <sec id="sec-3-6">
        <title>Analysis</title>
        <p>Our journey into comparative analysis was underpinned by a pivotal phase where the expertise of
SMEs came to the forefront. With the aid of the 60-topic model chosen by our SMEs, each topic
was carefully categorized, setting the stage for assessing different evaluation methods. These
SMEs, equipped with a profound understanding of Malcolm Fraser’s political journey, worked in
cohesion to separate meaningful topics from the rest. A meaningful topic was one that provided
practical, valuable, and coherent insights into various aspects of Fraser’s political trajectory.</p>
        <sec id="sec-3-6-1">
          <title>2https://mimno.github.io/Mallet/topics</title>
          <p>20
30
40</p>
          <p>PMI
NPMI
LCP
PMI
NPMI
LCP
PMI
NPMI</p>
          <p>LCP
ORACLE-NPMI</p>
          <p>RANDOM
MAJORITY
On the other hand, other topics encompassed themes that, while coherent, did not seamlessly
harmonize with the overarching narrative of Fraser’s political journey. SMEs further enriched our
analysis by appending subject labels to topics, where feasible. These labels encompassed diverse
subjects, ranging from the Vietnam War and Wool Marketing to Education, reflecting the wide
spectrum of themes in Fraser’s speeches. Examples of both meaningful and other topics can be
seen in Table 1.</p>
          <p>This classification procedure takes center stage in our study, serving as the foundation against
which the performance of both quantitative metrics and GPT-4 is compared. The SMEs’ keen
evaluation of topic significance forms a pivotal reference point, enabling us to gauge the accuracy
and subtleties of alternative approaches. This strategic alignment underscores the SMEs’ role not
merely as annotators but as navigators guiding our exploration of automated evaluation methods.
An overall summary of our framework is seen in Figure 1.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Comparative Analysis of Coherence Metrics, Subject Matter</title>
    </sec>
    <sec id="sec-5">
      <title>Experts, and GPT-4: A Multifaceted Evaluation</title>
      <p>In this section, we delve into an extensive comparative analysis that assesses the efficacy of
coherence metrics, the discernment of subject matter experts (SMEs), and the automated capabilities of
GPT-4. Our investigation aims to determine the extent to which these evaluation methodologies
can effectively identify the significance of topics derived from our Latent Dirichlet Allocation
(LDA) model.</p>
      <sec id="sec-5-1">
        <title>4.1. Coherence Metrics: A Quantitative Lens on Topic Significance</title>
        <p>
          To determine how well the quantitative evaluation of topics measured up against the judgments of
our SMEs, we calculated 3 coherence scores: Pointwise Mutual Information (PMI), Normalized
Pointwise Mutual Information (NPMI), and Local Coherence Probability (LCP) as calculated
in Lau and Baldwin [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. PMI measures the pointwise mutual information between each pair
of words in a topic, while NPMI normalizes this score by dividing it by the logarithm of the
probability of the two words occurring together. LCP, on the other hand, measures the probability
of a given word following another word within a certain distance in the same document. These
coherence scores were used as a predictor of topic quality.
        </p>
        <p>
          For the reference corpus, we created a corpus that consisted of 2,500 Wikipedia articles, 5,000
news corpus articles, and 484 Malcolm Fraser speeches that were not used in the modeling of the
topics nor the decade from which the topics had been induced. The reference corpus serves as a
baseline for evaluating the quality of the topics [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] which allowed us to evaluate the coherence
of the topics generated by the LDA model against the opinions of subject matter experts (SMEs).
The corpus had approximately 99k tokens with 47k, 32k, and 20k being made up of the news
corpus, Wikipedia, and the remaining Malcolm Fraser speeches, respectively. Specifically, we
wanted to ascertain if high coherence scores correlated with meaningful topics, so we ranked
them according to their coherence scores (highest to lowest) and labeled the top n-ranked topics
(n = 20, 30, 40) as meaningful. The reason for the selection of these values could be based on a
trade-off between the number of topics that can be labeled as meaningful and the quality of the
labeled topics. The results of coherence as a predictor of meaningful topics are shown in Table 2.
As an upper bound, we also labeled the top 37 topics of the ranked NPMI, which coincides with
the number of topics the SMEs deemed meaningful, shown as ORACLE-NPMI. This table shows
that the coherence scores for NPMI and PMI consistently performed better than our RANDOM
and MAJORITY class baselines, with NPMI achieving an f1-score of 0.81 for n=40, equivalent
to the f1-score for ORACLE-NPMI. The metric LCP did not always surpass these baselines and
came nowhere near ORACLE-NPMI. Save for LCP, using quantitative topic coherence measures
to ascertain quality topics would have resulted in comparable topics to the ones chosen by SMEs
for n=40.
        </p>
        <p>
          Figure 2 shows the coherence score for the NPMI metric with the blue topics, on the left,
deemed meaningful, and the red topics were those marked as other. As can be seen in this graphic,
the meaningful topics consistently have higher coherence scores, which coincides with them being
good predictors of meaningful topics as deemed by SMEs. We performed the Kruskal-Wallis
test [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] to determine whether there is an inherent difference in the meaningful and other topics
seen in Figure 2. We calculated H = 17.93 and p-value = 2.29− 5 using the scipy Python library.
Therefore, given the small p-value, we accept the null hypothesis that the distribution of the two
groups is not the same.
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. GPT-4 as a Surrogate SME: Automating Topic Classification</title>
        <p>The evaluation of topic models often hinges on the coherence and relevance of the topics generated.
Traditional metrics like NPMI and PMI offer quantitative measures, but as our findings suggest,
they occasionally fall short in capturing the nuanced understanding that Subject Matter Experts
(SMEs) bring to the table. For instance, topics such as ‘USA’ and ‘Foreign Policy’, despite their
apparent significance, might be overlooked based on these metrics alone, as seen in Figure 2.</p>
        <p>The real-world application of topic modeling, especially in specialized domains, often requires
the insights of SMEs. Their deep domain knowledge can reveal subtleties and intricacies that
might elude quantitative metrics. However, the engagement of SMEs is not without its challenges.
Their expertise, though invaluable, might not always be readily available. Moreover, the evaluation
process can be labor-intensive and time-consuming. A pivotal challenge is determining the optimal
value for n, the number of top-ranked topics labeled as meaningful. This choice is inherently
subjective and can significantly influence the outcome of the evaluation, potentially introducing
variability and inconsistencies in the insights derived.</p>
        <p>To circumvent these challenges and introduce a scalable, automated approach to topic
classification, we turned to the capabilities of large language models, specifically GPT-4. The underlying
hypothesis was simple yet profound: Could a model like GPT-4, with its vast training data and
sophisticated architecture, emulate the discernment typically associated with SMEs?</p>
        <p>Our methodology involves tasking GPT-4 with the classification of the 60 identified topics. As
prompts, we provided the model with a detailed context of our collaboration, explaining the aim,
the modeling process, and the SMEs’ evaluation criteria. The model was then presented with
the top 10 words for each of the 60 topics just as SMEs’ were and was tasked with classifying
the topics as meaningful or other. To ensure the robustness of our approach and to account
Temperature
0
0.25
for potential model variability, each topic undergoes classification vfie times. This iterative
methodology not only provides multiple perspectives on each topic but also allows us to assess
the internal consistency of GPT-4’s classifications. We chose the maximum result from vfie runs.</p>
        <p>Recognizing the potential variability in GPT-4’s outputs, we explored different temperature
settings. The temperature parameter in GPT-4 influences the randomness of the model’s outputs. A
lower temperature makes the model’s predictions more deterministic, while a higher temperature
introduces more variability. By experimenting with a range of temperature values, from 0 to 1, we
aim to identify the optimal setting that maximizes the model’s accuracy while preserving its ability
to discern meaningful topics. Notably, while GPT-4 offers two primary parameters to influence
its output randomness - temperature and top_p - we chose to adjust only the temperature setting.
This decision was based on the GPT-4 API documentation from OpenAI, which recommends
adjusting either the temperature or top_p, but not both3.</p>
        <p>Table 3 offers a comparative analysis, stating the performance metrics of GPT-4 against
ORACLE-NPMI for the classification of meaningful and other topics. This comparison
underscores the potential of GPT-4 as a surrogate SME. The results indicate that GPT-4, at a
temperature of 0.5, can match the F Score of the Oracle NPMI with a score of 0.81, and achieves
an F score of 0.8 at all other temperatures. Importantly, while overall accuracy is crucial, our
primary focus was on correctly identifying as many meaningful topics as possible, given their
significance in describing Fraser’s narrative. Notably, at all temperatures, GPT-4 surpassed the
ORACLE-NPMI’s recall. The ability to achieve this performance without prior knowledge of the
optimal value for n underscores the promise of this approach.</p>
        <sec id="sec-5-2-1">
          <title>3https://platform.openai.com/docs/api-reference/chat/create</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion</title>
      <p>Only SMEs</p>
      <p>SMEs and GPT-4
SMEs and ORACLE-NPMI</p>
      <p>Canberra Development</p>
      <p>Politics
Parliament</p>
      <p>Wannon</p>
      <p>USA
Foreign Policy</p>
      <p>New Guinea
National Service</p>
      <p>ALP
canberra; capital; decision; move;
development; building; years ago; grow;
hear; home
programme; point; view; remark;
speak; line; attack; question; full;
press
opposition; debate; majority; attack;
week; thing; question; show; move;
important
fraser; company; operate; malcolm
fraser; today; worth; wannon; order;
support; press
united states; american; president;
interest; world; thing; washington;
meet
prime minister; leader; world;
international; difference; president; united
nation; whitlam; asian; put
new guinea; agreement; full; council;
position; territory; white; bring; long;
argument
call; service; period; provide;
introduce; full; condition; general;
provision; additional
cost; speech; calwell; general; thing;
kind; promise; hard; governor; pay</p>
      <p>The evaluation of topic models, especially in the context of historical speeches like those
of Malcolm Fraser, presents a unique set of challenges and considerations. Our comparative
analysis between NPMI, GPT-4, and SMEs has shed light on the intricacies of topic classification,
revealing both the strengths and limitations of each approach. Our findings as highlighted in
Table 4 provide a comprehensive understanding of the differences between the approaches. In
the table, we specifically showcase the results from GPT-4 at a temperature value of 0.5, which
achieved the highest F1-score of 0.81 while also suprassing the recall of ORACLE-NPMI at 0.86.
Let’s delve deeper:
• Performance Metrics and Their Implications:
1. NPMI’s Robustness: NPMI’s strong performance in terms of precision, recall, and
F1-score reinforces its reputation as a dependable quantitative coherence measure. Its
alignment with SMEs on topics like "National Service" and "ALP" underscores its
value in topic modeling tasks and its ability to capture significant political and policy
narratives. A key factor contributing to NPMI’s success in identifying these topics is
the reference corpus used. A well-curated reference corpus enhances the accuracy
and relevance of coherence-based topic modeling by capturing context.
2. GPT-4’s Competence: The results showing GPT-4’s performance on par with NPMI
is both remarkable and promising. Its alignment with SMEs on topics such as
"Wannon" and "Foreign Policy" indicates the vast potential of large language models
in domain-specific tasks. However, while GPT-4 can emulate the capabilities of
specialized metrics, it may not possess the same depth of understanding as SMEs, as
evident in its unique topic identifications.
3. Patterns in GPT-4’s Misclassifications:
a) Context-Specific Topics: Topics such as "Canberra Development," "National
Service," and "ALP" as seen in Table 4 are deeply rooted in the Australian
context. GPT-4 might have struggled with these due to their specificity and the
potential lack of emphasis on the Australian political and developmental context
in its training data.
b) Broad and Abstract Topics: "Politics" and "Parliament" are more abstract
and encompassing. While SMEs can easily discern the nuances of such topics,
GPT-4 might have found it challenging to pinpoint their significance amidst other
potential topics.
• Scalability, Consistency, and Real-world Applicability: In practical scenarios, handling
large volumes of data demands scalable solutions. GPT-4’s capacity to efficiently process
extensive datasets without sacrificing consistency makes it an invaluable tool for
largescale text analysis tasks. This scalability becomes even more essential when obtaining
SME insights is challenging due to various constraints. GPT-4’s consistent outputs are
commendable. However, the occasional variations we observed in topic labeling highlight
the importance of iterative methodologies and cross-validation in topic classification
endeavors.
• The Indispensable Role of SMEs: The unique topics identified by SMEs, such as
"Canberra Development" and "Politics," emphasize their unparalleled depth of understanding.
Beyond mere numbers, SMEs contribute qualitative insights, adding layers of interpretation
and context that enrich the overall analysis. Their expertise becomes indispensable when
dissecting content like Fraser’s speeches, where historical, political, and cultural nuances
significantly influence the overarching narrative.
• Future Directions and Considerations: One of the most promising avenues for future
research lies in the synergy between SMEs and models like GPT-4. SMEs bring
unparalleled depth and understanding, especially when determining the granularity of topics
needed to address broader research questions, such as discerning shifts in Fraser’s rhetoric
over time. Once this granularity is established, leveraging GPT-4 can be invaluable. The
model’s scalability and consistency can be employed to assess the significance of these
topics, streamlining the process and reducing the manual effort required by SMEs albeit
with some error.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Limitations</title>
      <p>
        A notable limitation of our study is its reliance on Subject Matter Experts (SMEs) possessing
profound knowledge of the dataset. Individuals lacking this specialized expertise might encounter
challenges in selecting the optimal number of topics or in evaluating the clarity and significance
of each topic. The Latent Dirichlet Allocation (LDA) method we employed conceptualizes
each document as a blend of topics, presuming each word originates from one of these topics.
However, this assumption might not hold true universally across diverse document types or topics,
potentially influencing the quality of the topics identified. Exploring alternative methodologies,
such as BERTopic [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] and Topic2Vec [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ], could provide valuable comparative insights. While
we did not venture into these methods within the scope of this study, they present promising
avenues for future research.
      </p>
      <p>Additionally, our analysis is temporally constrained, focusing on the period from 1960 to 1970
due to the limited availability of Fraser’s speeches. Extending the analysis to encompass speeches
from varied time frames could validate the generalizability of our findings.</p>
      <p>Another dimension worth highlighting is our exclusive utilization of the GPT-4 Large Language
Model (LLM). We did experiment with other models like Llama 2 chat and PALM 2, but
encountered challenges. Specifically, these models struggled to classify topics using the prompts
that were effective for GPT-4. The prompt length was also a limiting factor, which was not an
issue with GPT-4. Crafting uniform prompts compatible with several Large Language Models
would have necessitated significant time and effort. It’s essential to clarify that the primary
objective of this paper was not a comparative analysis of multiple LLMs, but rather an exploration
of the feasibility of using LLMs as potential replacements for traditional coherence metrics. The
results derived from GPT-4 affirmatively indicate this possibility.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Conclusion</title>
      <p>In this study, we embarked on a journey to understand the comparative efficacy of subject matter
experts (SMEs), coherence metrics, and the capabilities of GPT-4 in evaluating topic models.
While our initial expectation leaned towards SMEs being the gold standard for topic evaluation in
a specific domain, our findings revealed that common measures like coherence metrics performed
commendably, often aligning closely with expert judgments. Nevertheless, relying solely on
these metrics could have led to overlooking some pivotal topics. The introduction of GPT-4 into
the evaluation mix not only showcased its potential as a scalable and consistent evaluator but
also highlighted its ability to match, and in some instances, rival traditional coherence metrics.
Yet, even with GPT-4’s impressive performance, the study underscores the irreplaceable value
of SMEs’ nuanced understanding. In essence, while tools like coherence metrics and GPT-4
offer promising avenues for topic evaluation, they should complement, not replace, the insights
of SMEs. Our study advocates for a balanced amalgamation of automated methods and human
expertise for comprehensive topic model evaluation, emphasizing that GPT-4 can serve as a
standalone evaluator in the absence of SMEs.</p>
    </sec>
    <sec id="sec-9">
      <title>8. Online Resources</title>
      <p>• Website of our collaboration - Political Rhetoric Project
• Malcolm Fraser’s Speeches</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bassignana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramponi</surname>
          </string-name>
          , Preface to the
          <source>Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI)</source>
          ,
          <source>in: Proceedings of the Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI</source>
          <year>2023</year>
          )
          <article-title>co-located with 22th International Conference of the Italian Association for Artificial Intelligence (AI* IA</article-title>
          <year>2023</year>
          ),
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lazer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pentland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Adamic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Aral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-L.</given-names>
            <surname>Barabasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brewer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Christakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Contractor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fowler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gutmann</surname>
          </string-name>
          ,
          <article-title>Life in the network: The coming age of computational social science 323 (</article-title>
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Paths study on knowledge convergence and development in computational social science: Data metric analysis based on web of science</article-title>
          ,
          <source>Complexity</source>
          <year>2022</year>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          . doi:
          <volume>10</volume>
          .1155/
          <year>2022</year>
          /3200371.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Moretti</surname>
          </string-name>
          ,
          <source>Conjectures on world literature, New Left Review</source>
          <volume>1</volume>
          (
          <year>2000</year>
          )
          <fpage>54</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , B. Adams,
          <article-title>The wild process: constructing multi-scalar environmental narratives</article-title>
          , Ubiquity Press Ltd, United Kingdom,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , M. Mistica, G. Fairclough, T. Baldwin, Inferring Value:
          <article-title>A Multiscalar Analysis of Landscape Character Assessments</article-title>
          , Ubiquity Press Ltd, United Kingdom,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lazer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pentland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Adamic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Aral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-L.</given-names>
            <surname>Barabasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brewer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Christakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Contractor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fowler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gutmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jebara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Macy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Van Alstyne</surname>
          </string-name>
          ,
          <article-title>Computational social science</article-title>
          ,
          <source>Science</source>
          <volume>323</volume>
          (
          <year>2009</year>
          )
          <fpage>721</fpage>
          -
          <lpage>723</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Peponakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kapidakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Doerr</surname>
          </string-name>
          , E. Tountasaki,
          <article-title>From calculations to reasoning: History, trends and the potential of computational ethnography and computational social anthropology</article-title>
          , Social Science Computer Review (
          <year>2023</year>
          )
          <article-title>089443932311676</article-title>
          . doi:
          <volume>10</volume>
          .1177/ 08944393231167692.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Boyd-Graber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mimno</surname>
          </string-name>
          , Applications of topic models,
          <source>Foundations and Trends® in Information Retrieval</source>
          <volume>11</volume>
          (
          <year>2017</year>
          )
          <fpage>143</fpage>
          -
          <lpage>296</lpage>
          . doi:
          <volume>10</volume>
          .1561/1500000030.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Boyd-Graber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gerrish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Blei</surname>
          </string-name>
          , Reading tea leaves:
          <article-title>How humans interpret topic models</article-title>
          ,
          <source>Neural Information Processing Systems</source>
          <volume>32</volume>
          (
          <year>2009</year>
          )
          <fpage>288</fpage>
          -
          <lpage>296</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Jelodar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <article-title>Latent dirichlet allocation (lda) and topic modeling: models, applications, a survey (</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>D. M. Blei</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>M. I. Jordan</given-names>
          </string-name>
          , Latent dirichlet allocation,
          <source>J. Mach. Learn. Res</source>
          .
          <volume>3</volume>
          (
          <year>2001</year>
          )
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>D. M. Blei</surname>
          </string-name>
          ,
          <article-title>Probabilistic topic models</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>55</volume>
          (
          <year>2012</year>
          )
          <fpage>77</fpage>
          -
          <lpage>84</lpage>
          . URL: https://api.semanticscholar.org/CorpusID:753304.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Alghamdi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Alfalqi</surname>
          </string-name>
          ,
          <article-title>A survey of topic modeling in text mining</article-title>
          ,
          <source>International Journal of Advanced Computer Science and Applications</source>
          <volume>6</volume>
          (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .14569/IJACSA.
          <year>2015</year>
          .
          <volume>060121</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Silwattananusarn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kulkanjanapiban</surname>
          </string-name>
          ,
          <article-title>A text mining and topic modeling based bibliometric exploration of information science research</article-title>
          ,
          <source>IAES International Journal of Artificial Intelligence (IJ-AI</source>
          )
          <volume>11</volume>
          (
          <year>2022</year>
          )
          <fpage>1057</fpage>
          -
          <lpage>1065</lpage>
          . doi:
          <volume>10</volume>
          .11591/ijai.v11.
          <year>i3</year>
          .
          <fpage>pp1057</fpage>
          -
          <lpage>1065</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Lau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Grieser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          ,
          <article-title>Automatic evaluation of topic coherence, in: Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Los Angeles, California,
          <year>2010</year>
          , pp.
          <fpage>100</fpage>
          -
          <lpage>108</lpage>
          . URL: https://aclanthology.org/ N10-1012.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Lau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          , Machine reading tea leaves:
          <article-title>Automatically evaluating topic coherence and topic model quality, in: Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Gothenburg, Sweden,
          <year>2014</year>
          , pp.
          <fpage>530</fpage>
          -
          <lpage>539</lpage>
          . URL: https://aclanthology.org/E14-1056. doi:
          <volume>10</volume>
          .3115/v1/
          <fpage>E14</fpage>
          -1056.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Mimno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          , E. Talley,
          <string-name>
            <given-names>M.</given-names>
            <surname>Leenders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <article-title>Optimizing semantic coherence in topic models</article-title>
          ,
          <source>in: Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Edinburgh, Scotland, UK.,
          <year>2011</year>
          , pp.
          <fpage>262</fpage>
          -
          <lpage>272</lpage>
          . URL: https://aclanthology.org/D11-1024.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Röder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Both</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hinneburg</surname>
          </string-name>
          ,
          <article-title>Exploring the space of topic coherence measures</article-title>
          ,
          <source>Proceedings of the Eighth ACM International Conference on Web Search and Data Mining</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>N.</given-names>
            <surname>Aletras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stevenson</surname>
          </string-name>
          ,
          <article-title>Evaluating topic coherence using distributional semantics</article-title>
          ,
          <source>in: Proceedings of the 10th International Conference on Computational Semantics (IWCS</source>
          <year>2013</year>
          )
          <article-title>- Long Papers, Association for Computational Linguistics</article-title>
          , Potsdam, Germany,
          <year>2013</year>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>22</lpage>
          . URL: https://aclanthology.org/W13-0102.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Lau</surname>
          </string-name>
          , T. Baldwin,
          <article-title>The sensitivity of topic coherence evaluation to topic cardinality, in: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</article-title>
          , San Diego, California,
          <year>2016</year>
          , pp.
          <fpage>483</fpage>
          -
          <lpage>487</lpage>
          . URL: https://aclanthology.org/ N16-1057. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N16</fpage>
          -1057.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: NIPS</source>
          ,
          <year>2017</year>
          . URL: https://api.semanticscholar. org/CorpusID:13756489.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Thirunavukarasu</surname>
          </string-name>
          ,
          <article-title>Large language models will not replace healthcare professionals: curbing popular fears and hype</article-title>
          ,
          <source>Journal of the Royal Society of Medicine</source>
          <volume>116</volume>
          (
          <year>2023</year>
          )
          <article-title>1410768231173123</article-title>
          . doi:
          <volume>10</volume>
          .1177/01410768231173123.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          , P. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nie</surname>
          </string-name>
          , J. rong
          <string-name>
            <surname>Wen</surname>
          </string-name>
          ,
          <article-title>A survey of large language models</article-title>
          ,
          <source>ArXiv abs/2303</source>
          .18223 (
          <year>2023</year>
          ). URL: https://api.semanticscholar.org/CorpusID:257900969.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>L.</given-names>
            <surname>Floridi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chiriatti</surname>
          </string-name>
          , Gpt-3
          <article-title>: Its nature, scope, limits, and consequences</article-title>
          ,
          <source>Minds and Machines</source>
          <volume>30</volume>
          (
          <year>2020</year>
          )
          <fpage>681</fpage>
          -
          <lpage>694</lpage>
          . URL: https://api.semanticscholar.org/CorpusID:228954221.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ethayarajh</surname>
          </string-name>
          ,
          <article-title>How contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings</article-title>
          ,
          <source>in: Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2019</year>
          . URL: https://api.semanticscholar.org/CorpusID: 202120592.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Narasimhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salimans</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Improving language understanding by generative pre-training (</article-title>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>OpenAI</surname>
          </string-name>
          , Gpt-4
          <source>technical report, ArXiv abs/2303</source>
          .08774 (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>N.</given-names>
            <surname>Jethani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Genes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Major</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Jaffe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cardillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Heilenbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bonanni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Clayburn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Khera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sadler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Prasad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schlacter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montgomery</surname>
          </string-name>
          , E. Kim,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Razavian</surname>
          </string-name>
          ,
          <article-title>Evaluating chatgpt in information extraction: A case study of extracting cognitive exam dates and scores (</article-title>
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .1101/
          <year>2023</year>
          . 07.10.23292373.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Kruskal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. A.</given-names>
            <surname>Wallis</surname>
          </string-name>
          ,
          <article-title>Use of ranks in one-criterion variance analysis</article-title>
          ,
          <source>Journal of the American Statistical Association</source>
          <volume>47</volume>
          (
          <year>1952</year>
          )
          <fpage>583</fpage>
          -
          <lpage>621</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>M.</given-names>
            <surname>Grootendorst</surname>
          </string-name>
          , Bertopic:
          <article-title>Neural topic modeling with a class-based tf-idf procedure</article-title>
          ,
          <year>2022</year>
          . arXiv:
          <volume>2203</volume>
          .
          <fpage>05794</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>L.-Q.</given-names>
            <surname>Niu</surname>
          </string-name>
          , X.-Y. Dai,
          <article-title>Topic2vec: Learning distributed representations of topics</article-title>
          ,
          <year>2015</year>
          . arXiv:
          <volume>1506</volume>
          .
          <fpage>08422</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>