<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Models to Fill Relevance Judgment Holes?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zahra Abbasiantaeb</string-name>
          <email>z.abbasiantaeb@uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chuan Meng</string-name>
          <email>c.meng@uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leif Azzopardi</string-name>
          <email>leif.azzopardi@strath.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohammad Aliannejadi</string-name>
          <email>m.aliannejadi@uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>EMTCIR '24: The First Workshop on Evaluation Methodologies, Testbeds and Community for Information Access Research</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Amsterdam</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Strathclyde, Glasgow</institution>
          ,
          <addr-line>Scotland</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Incomplete relevance judgments limit the reusability of test collections. When new systems are compared against previous systems used to build the pool of judged documents, they often do so at a disadvantage due to the “ holes” in test collection (i.e., pockets of un-assessed documents returned by new systems). In this paper, we aim to extend test collections by employing Large Language Models (LLM) to fill these holes by leveraging and grounding existing judgments. We explore this problem in the context of Conversational Search (CS) using TExt Retrieval Conference (TREC) Interactive Knowledge Assistance Track (iKAT) collection, where information needs are highly dynamic and the responses (and, the results retrieved) are much more varied (leaving bigger holes). While previous work has shown that automatic judgments from LLMs result in highly correlated rankings, we find it substantially lower correlates when human plus automatic judgments are used (regardless of LLM, one/two/few shot, or fine-tuned). We further find that, depending on the LLM employed, new runs will be highly favored (or penalized), and this efect is magnified proportionally to the size of the holes. Instead, one should generate the LLM annotations on the whole document pool to achieve more consistent rankings with human-generated labels. Further work is needed to prompt engineer and fine-tune LLMs to align with human judgment, thereby, improving the methods' accuracy.</p>
      </abstract>
      <kwd-group>
        <kwd>Conversational search</kwd>
        <kwd>Large language models</kwd>
        <kwd>Judgments</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Building reusable test collections in a cost-eficient manner to evaluate current and future systems has
been a long-standing challenge in the field of information retrieval ( IR) [1]. The predominant strategy
for creating such collections has been through the use of pooling [2, 3] – where a subset of documents,
taken from various systems, is assessed for relevance. This is a compromise away from the “ideal test
collection” with complete relevance assessments which is infeasible and impractical. While the pooling
strategy is fairly robust [4, 5, 6, 7], it leads to various evaluation biases (e.g., [8, 5]) where systems that
did not contribute to the pool, can be significantly disadvantaged. This is because documents that have
not been judged are considered irrelevant. Therefore, the fewer judged/assessed documents returned in
a ranking, the lower the retrieval performance ceiling [9]. However, the fewer the judgments required
to compare systems the cheaper the test collection.</p>
      <p>While researchers have tried to address these trade-ofs in various ways, either by proposing new
metrics and methodologies for compensating for the un-assessed documents (e.g., [5, 10, 9]) and/or
developing new pooling and judgment strategies (e.g., [11, 12, 13]), “holes” in the pools still remain [14,
CEUR</p>
      <p>ceur-ws.org</p>
      <p>
        However, with the advances in the development of powerful Large Language Models (LLMs) and other
neural-based models, new opportunities arise for building scalable, robust, and reusable test collections
at a lower cost. LLMs ofers the possibility to: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) assess large volumes of documents reasonably cheaply,
especially compared to human judgments, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) do so in a consistent and independent but potentially
biased manner, (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) if the LLM and prompt are fixed and shared, then judgments can be collected at
diferent times under the same conditions, and, (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) typically at a higher quality than “typical” crowd
workers.
      </p>
      <p>Indeed, recent studies [16, 17, 18, 19] have shown the efectiveness of using LLMs to automatically
generate relevance judgments in the scenario of ad-hoc search. These works demonstrate that
LLMbased judgments exhibit a high correlation with human judgments. Thomas et al. [18], Faggioli et al.
[19] prompted commercial LLMs (e.g., GPT-3.5, GPT-3.5/4) to generate relevance judgments. However,
commercial LLMs come with limitations like non-reproducibility, non-deterministic outputs, and
potential data leakage between pre-training and evaluation data, impeding their utility in scientific
research [20]. MacAvaney and Soldaini [15] and Khramtsova et al. [17] prompted an open-source LLMs,
Flan-T5 [21], for generating relevance judgments. While open-source LLMs are less efective, they do
ofer the potential for the development of reproducible and reusable test collections at scale. This led to
eforts by Meng et al. [16], who fine-tuned an open-source LLM, Llama [22] using parameter-eficient
ifne-tuning ( PEFT) [23] to better condition the LLM for performing the task of assigning relevance
judgments. They found that more complete test collections could be produced, with high quality, at a
lower cost.</p>
      <p>While this prospect is very appealing, it is fraught with new, unexplored challenges. Of interest, in
this work is the notion of grounding. Training systems on the judgments of LLMs, and then evaluating
those systems on subsequent test collections, based on judgments from LLMs, creates a potentially
dangerous cycle that may amplify and re-enforce existing biases inherent in LLMs. Grounding the
LLMs based judgment given human judgments provides a mechanism to condition the LLMs to be
more aligned with human annotators, reducing the risk of AI falling into a negative feedback loop [24].
To this end, MacAvaney and Soldaini [15] focus on a setting wherein the LLM is given one relevant
example to help ground the subsequent judgments. In this paper, we draw up this direction in the
context of Conversational Search (CS) and aim to build/augment test collections with grounded LLM
based judgments.</p>
      <p>Conversational search (CS) is defined as responding to the user’s information needs in the context of
the conversation [25, 26, 27, 28]. In CS the user’s information need depends on the query, the context of
the conversation, and the user’s personal preferences [29]. This results in highly dynamic, non-linear
conversational trajectories – where a user’s information need could be answered quite diferently
depending on the system’s interpretation because the needs are evolving and change in response to the
information presented. To meet these changing information needs, systems are likely to pull in a wider
range of documents. This could result in many more unassessed documents, creating bigger gaps when
evaluating new systems, which would significantly reduce the reusability of test collections [ 30]. So, in
this work, we explore whether LLMs can be used to augment and extend CS test collections which are
grounded by human annotations, in order to evaluate new, future systems.</p>
      <p>
        In this paper, we leverage both commercial and open-source LLMs in zero- and few-shot, as well as
ifne-tuning manners, to automatically generate relevance judgments in the CS scenario. Specifically,
for commercial LLMs, we use the GPT-3.5 model with diferent prompts. We try one-shot, two-shot,
and zero-shot prompts. For open-source LLMs, we consider three setups: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) we directly use the Llama
checkpoint released by Meng et al. [16], which has undergone fine-tuning based on human-labeled
relevance judgments on MS MARCO, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) we directly prompt Llama-3 [31] in a one-shot way, and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) we
ifrst use partial human-labeled relevance judgments in a CS dataset to fine-tune Llama-3 and then test
it using the rest of the relevance judgments in the dataset.
      </p>
      <p>We use the TREC iKAT 2023 [29] dataset which is a personalized CS benchmark. In our experiments,
we re-create the relevance judgments of the TREC iKAT 2023 benchmark using various prompts and
techniques. We compare the generated judgments with the oficial TREC iKAT 2023 relevance labels in
terms of various metrics. In particular, we are interested in answering the following research questions:
To answer the RQs, we conduct a set of experiments where for RQ1, we create a training and test set
of relevance labels and compare Llama-1, Llama-3, and GPT-3.5 in zero-shot, few-shot, and fine-tuning
settings. For RQ2, we use GPT-3.5 to generate relevance labels on the oficial TREC iKAT 2023 pool
and use it to rank the oficial TREC runs. To answer RQ3, we conduct multiple experiments where at
each experiment, we remove the relevance labels of each run from the pool, mimicking the case where
that model is not included in the original pool. We then generate the relevance labels using GPT-3.5
and use those labels to assess the new model.</p>
      <p>Our results show that ranking of the retrieval models in CS datasets using human- and LLM-generated
annotations are highly correlated, although they have a low agreement in binary- and graded-level, in
line with the findings of Faggioli et al. [19] on ad-hoc search. In addition, the correlation of diferent
IR metrics converges as we add more retrieval systems to the comparison pool. We show that by
ifne-tuning the open-source LLMs we can achieve a higher agreement between LLM-generated and
human judgments. However, higher agreement in terms of binary and graded judgments does not
necessarily result in a higher correlation in the ranking of the runs. Moreover, we show that in the case
of adding a new retrieval model, filling the holes with a zero-shot Llama model results in less significant
shifts in the ranking of the corresponding retrieval model, compared to using one-shot GPT-3.5, perhaps
due to higher agreement of the ratings, and because the one-shot GPT-3.5 model is biased to predict
higher relevance scores.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <sec id="sec-2-1">
        <title>2.1. The choice of LLMs</title>
        <p>
          As a commercial closed-source LLM, we consider the GPT-3.5 (gpt-3.5-turbo-0125) model with values
of 0 and 1 for the temperature ( ) and top_p=1. The value of 0 for temperature means that the model
outputs the tokens with the highest probability and has no randomness. For open-source LLMs, we
consider the following three setups: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) we directly use the Llama-1 (7B) checkpoint released by Meng
et al. [16],1 which has undergone fine-tuning using the human-labeled relevance judgments from the
development set of MS MARCO [32]; (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) we directly prompt Llama-3 (8B) [31] in a one-shot way;
specifically, besides the original version of Llama-3 (8B), we also consider its instruction-tuned version
(Llama-3-inst); (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) we first use partial human-labeled relevance judgments in a CS dataset to fine-tune
the two variants Llama-3, and then test them using the rest of the relevance judgments in the dataset.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Dataset</title>
        <p>We use the TREC iKAT 2023 [29] benchmark in our experiments. In this benchmark, the relevance
of each query–document pair is assessed and represented with a score in the range of 0–4. Using
1https://github.com/ChuanMeng/QPP-GenRE
the GPT-3.5 model, we judge the relevance of all query–document pairs from the TREC iKAT 2023
collection. For fine-tuning the Llama-3 model [ 31], we divide the TREC iKAT 2023 benchmark into
train, test, and validation sets. First, we randomly remove 20680 irrelevant documents (  &lt; 2 ) from
the existing pool to ensure that the number of relevant (  &gt;= 2 ) and irrelevant documents for each
query are the same. Second, we randomly split the documents for each query between train, test, and
validation sets. We keep the portion of train, test, and validation sets as 70%, 15%, and 15%. The train,
test, and validation sets include 3852, 917, and 709 query-passage pairs, respectively. The distribution
of the labels for the train, test, and validation set are shown in Table 1. We ensure that all user queries
appear in the training set and have at least one positive and one negative document.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Retrieval models</title>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Metrics</title>
        <p>To assess the correlation of the generated pools, we use the baselines and runs submitted to TREC iKAT
2023. There are in total 28 baselines and runs submitted to the TREC iKAT 2023. We use the output of
retrieval for these runs and baselines which are released by the organizers.</p>
        <p>
          To assess the ranking-level performance of our proposed models for relevance judgment, we rank the
retrieval models two times, once based on their performance using the main pool and second based on
the generated pool. We compute and report Kendall’s Tau ( ) metrics to assess the correlation between
the two rankings. To assess the agreement between proposed models for relevance judgment and
humans, we report Cohen’s Kappa agreement at binary and graded levels. We convert the predicted
graded scores (
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">0-4</xref>
          ) to a binary label by considering scores 2-4 as relevant and scores 0-1 as irrelevant.
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. GPT-3.5 prompt design</title>
        <p>We designed three diferent prompts, inspired by relevant work. We use the resolved utterances provided
by humans as a context-independent query in our prompts.
• We design a zero-shot prompt inspired by the prompt used in Thomas et al. [18] which is shown in</p>
        <p>Table 2.
• Our next prompt is a one-shot prompt which includes a relevant document with a relevance score
of 4. This prompt is inspired by the prompt used by MacAvaney and Soldaini [15]. This prompt is
shown in Table 2. We use the canonical response for the corresponding user utterance from the TREC
iKAT 2023 collection as the relevant (perfect) example with a relevance score of 4. The canonical
responses are provided by the organizers and are supposed to be the best possible answer that can be
given to the utterance at every point in the conversation [29].
• The third prompt includes a relevant document and an irrelevant document. We randomly sample
these documents from the TREC iKAT 2023 oficial pool. A document with a relevance score higher
than or equal to 2 is selected as relevant and a document with a relevance score lower than 2 is
randomly selected as irrelevant. This prompt is shown in Table 2.</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.6. Llama fine-tuning and prompt design</title>
        <p>For open-source LLMs, we consider the following setups:
• We use the one-shot prompt shown in Table 2 to prompt Llama-3.
• For fine-tuning Llama-3, we follow Meng et al. [16] to fine-tune Llama-3 using a novel PEFT method,
4-bit QLoRA [23]; the train, test, and validation data used for fine-tuning and inference over the
model is explained above.
• Because the Llama model released by Meng et al. [16] is only trained to generate binary relevance
judgments given a query and a passage on MS MARCO, we leave out the user personal knowledge
when we use the Llama model released by Meng et al. [16].
Document : { document }
Score:
Please only generate an int score between 0 to 4 to say to
what extent the document is relevant to the user question.</p>
        <p>Score lower than 2 means the document is irrelevant.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments &amp; Results</title>
      <p>In this section, we describe in detail the experiment we designed to answer each research question,
followed by the results we obtained by doing the experiments.</p>
      <sec id="sec-3-1">
        <title>3.1. LLM relevance label comparison (RQ1)</title>
        <p>3.1.1. Experimental design
In this section, we aim to answer our first research question, RQ1: How do diferent LLMs compare in
predicting relevance judgments in conversational search? To do so, as described in Section 2, we randomly
sample the human-generated labels into the train, validation, and test sets and use the training data to
ifne-tune Llama-based models. We then compare the performance of the Llama-based models with the
diferent GPT-3.5-based models.
3.1.2. Results
In Table 3, we report the agreement of our proposed models on the test set. The experiments reveal that
we can improve the agreement by fine-tuning the Llama-3-inst model. As can be seen, the fine-tuned
Llama-3-inst achieves the agreement of 0.729 on the binary level.</p>
        <p>We use fine-tuned and zero-shot Llama to predict the test set. We create a small pool based on</p>
        <p>NDCG
0.788
0.852</p>
        <p>
          MRR
0.815
0.810
0.794
0.847
0.873
0.804
0.921
0.735
0.847
0.614
the model’s predictions on the test set. We sorted the TREC iKAT 2023 runs based on their retrieval
performance two times (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) using the LLM-generated assessments and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) using the human-generated
pool. We compare the ranking of runs by computing the correlation between them. Table 4 reports the
result of the relative ranking performance of diferent LLMs, compared to human ranking. Surprisingly,
the Llama-3 model is not performing better than the GPT-3.5 model in this scenario while it has a
higher agreement with human labels on the same test set. This could be due to the diferent labeling
biases that the models have where GPT-3.5 labels could be more diferent from human labels in terms
of absolute numbers, but when we compare diferent documents they are more similar relatively.
        </p>
        <p>We report the binary- and graded-level confusion matrices for prediction of best Llama- and
GPT3.5-based models over the test set in Tables 5 and 6, respectively. Additionally, we report the binary
confusion matrix of the Llama-1 which is fine-tuned on the MS MARCO dataset. According to Table 6,
the fine-tuned Llama-3-inst has a very lower tendency to assign scores 1 and 4 compared to scores 0
and 2. The behavior of the Llama is natural as the train data has less number of 1 and 4 labels compared
to other labels according to Table 1. We do not observe such bias in the one-shot GPT-3.5 as this model
is not fine-tuned on the train data. However, the one-shot GPT-3.5 has predicted a large number of
4 labels compared to the Llama. This bias could be the result of putting the canonical answer in the
prompt as a positive example with a score of 4.</p>
        <p>As can be seen in Table 5, the one-shot GPT-3.5 has more false positives compared to the fine-tuned
Llama over the binary-level labels. Giving one positive example in the prompt might cause this bias.
The distribution of the false positive and false negative are approximately equal for the Llama. This
might be because the number of relevant and irrelevant passages in the training set of LLaMA is equal.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. LLM vs. human labels (RQ2)</title>
        <p>3.2.1. Experimental design
Here, we aim to answer our second research question, RQ2: How do LLM-generated assessments compare
to human-generated assessments in both absolute label prediction and relative ranking of retrieval models
in CS datasets? To do so, we regenerate all the relevance labels of the oficial TREC iKAT 2023 pool
using the three prompts described in Section 2. Inspired by Faggioli et al. [19] and MacAvaney and
Soldaini [15] we aim to test the hypothetical case of having zero or one assessed passage for each query
and rely on LLMs to assess the pool. In this experiment, we evaluate the models based on both the</p>
        <p>Temperature</p>
        <p>Binary</p>
        <p>Graded</p>
        <p>
          NDCG
quality of individual predicted labels and the relative ranking of the models assessed with each of the
generated labels, compared to human labels.
3.2.2. Results
We report the agreement of proposed models with human labels from the complete pool of the TREC
iKAT 2023 benchmark in Table 7. As can be seen, the one-shot prompting of the GPT-3.5 has the highest
agreement with human labels among zero-shot and two-shot prompting. Additionally, setting the
temperature to 0 increases the agreement. A lower value for the temperature parameter means that the
model has less randomness in generating the output while higher values mean that the model is more
creative and has randomness in generation. Our one-shot prompt has the highest agreement in terms
of both binary and graded labels. The better performance of one-shot prompting compared to two-shot
prompts indicates that (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) using the canonical response as a positive example is more useful and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
using two positive and negative examples confuses the GPT-3.5.
        </p>
        <p>We use the pools generated by GPT-3.5 in diferent settings and the oficial pool assessed by humans to
assess the runs and rank them. More correlation between the rankings obtained by the LLM-generated
pool and the human-assessed pool indicates that using LLM-generated assessments is as efective as
using human-generated assessments. Table 8 shows the correlation between the relative ranking of runs
using diferent LLM-generated pools with the human-assessed pool. As can be seen, one-shot prompting
the ChatGPT model significantly outperforms the other settings over Kendall’s Tau correlation metric.
Interestingly, we observe that the temperature of 0 is not always better than the temperature of 1 in
terms of all retrieval metrics.</p>
        <p>In Figure 1, we show the correlation between the relative ranking of LLM-generated and
humangenerated pools using the  best-performing runs. The best-performing model is selected according
to the ranking based on using a human-generated pool. The LLM-generated pool is generated using
the best pool generation model from Table 8, i.e., one-shot labeling ChatGPT using temperature of 0.
Considering the 4 best-performing runs, the relative ranking using the LLM-generated assessments is
the same as using the human-generated assessments over all ranking metrics. Using the LLM-generated
assessments, the relative ranking of the 10 best-performing runs based on NDCG@5 is the same as the
relative ranking of these runs based on human-generated assessments (Kendall’s Tau = 1). According
to Figure 1, as the value of  increases (more runs are included in the comparison), the value of the
correlation converges. This finding represents the reliability of LLM-generated assessments in terms of
relative ranking.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Filling judgment holes (RQ3)</title>
        <p>
          3.3.1. Experimental design
To answer our RQ3: How are new models with diferent levels of holes ranked using LLM-generated
assessments? Can we rely on LLM-generated labels to compare a new model with existing models?, we
simulate the case where a new model is being tested using TREC iKAT 2023 runs. To do so, we do
multiple experiments where in each one we take out all the judgments of one run while keeping
those judgments that are in common with other existing runs. This leads to diferent levels of holes
per run, depending on their similarity to other existing models. We then assess the relevance of the
unjudged passages using GPT-3.5 and use those labels to compute the performance of the model. To
assess performance, we compare the ranking of the model using the original human assessments vs.
GPT-3.5-generated assessments and report the absolute diference in the model’s ranking in the two
cases. This indicates, how reliable LLM-generated labels are in filling the holes for new models. After
removing a run and generating labels for it using GPT-3.5, we do the ranking based on the (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) new pool
(human pool filled by LLM) and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) human pool which includes the human judgments for the holes of
the current run.
3.3.2. Results
The value of the absolute distance of the run in the two rankings based on the portion of the
Unjudged@10 passages for that specific run is shown in Figure 2. We use the one-shot GPT-3.5 model
with a temperature of 0 for hole filling. As can be seen, as the value of Unjudged@10 increases, the
absolute distance increases which means the missing run also increases. This makes sense because we
know that GPT-3.5 is biased to rate the passages with higher scores compared to humans. As a result,
we can conclude that given a new ranking model with a lot of missing judgments (a larger value for
Unjudged@10), it is advisable to recreate the whole pool using the GPT-3.5, rather than augmenting
the existing human-created pool by filling the holes using GPT-3.5.
        </p>
        <p>Interestingly, we see that the results of Llama exhibit a completely diferent trend where the number
of holes does not seem to matter. We see in the plot that Llama can consistently rank the missing run
close to its original ranking and even achieves perfect ranking at some points. This is in line with our
observation in Table 3, where we observed a higher agreement of Llama-generated labels with human
labels, leading to a lower disparity in terms of the absolute value of the labels, which then makes the
augmented labels more reliable.</p>
        <p>2
4
6
8 10 12 14 16 18 20 22 24 26 28</p>
        <p>K
0.2 0.4 0.6 0.8 1.0</p>
        <p>Unjudged@10
Figure 2: Absolute distance between the location of a new run before and after filling holes using GPT-3.5 and
Llama. The X-axis shows the average of unjudged documents among the top 10 documents returned by a new
run.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>In this work, we conducted extensive experiments to study the efect of LLM-generated relevance
judgments on incomplete relevance judgments (aka. “holes”) of the TREC iKAT 2023 collection. We
studied the efectiveness of diferent open-source and closed-source LLMs on generating relevance
assessments on the same set, where we observed that labels by fine-tuned Llama align better with human
labels compared to the labels obtained by few-shot prompting the GPT-3.5 model. In line with previous
work, we observed that automatic judgments from LLMs result in highly correlated model rankings;
however, we found that it substantially correlates lower when human plus automatic judgments were
used when a new model was being assessed on the pool. We further found that, depending on the LLM
employed, new runs will be highly favored (or penalized), and this efect is magnified proportional
to the size of the holes. We conclude that generating automatic labels on the whole pool is more
efective, rather than just the missing holes, as it leads to higher correlation and ensures that the same
labeling biases are applied to all the models. Further work is needed to refine prompt engineering
and fine-tuning of LLMs so they better match and reflect human annotations. This will help align
the models more closely with their intended purpose. Moreover, we plan to simulate various labeling
strategies to study the efectiveness of fine-tuning in more practical scenarios.
ity, in: Proceedings of the 28th Annual International ACM SIGIR Conference on Research and
Development in Information Retrieval, SIGIR ’05, 2005, p. 162–169.
[7] X. Lu, A. Mofat, J. S. Culpepper, The efect of pooling and evaluation depth on ir metrics,</p>
      <p>Information Retrieval Journal 19 (2016) 416–445.
[8] M. Baillie, L. Azzopardi, I. Ruthven, A retrieval evaluation methodology for incomplete relevance
assessments, in: Advances in Information Retrieval, 29th European Conference on IR Research,
ECIR 2007, volume 4425 of Lecture Notes in Computer Science, Springer, 2007, pp. 271–282.
[9] M. Baillie, L. Azzopardi, I. Ruthven, Evaluating epistemic uncertainty under incomplete
assessments, Information processing &amp; management 44 (2008) 811–837.
[10] J. A. Aslam, E. Yilmaz, Inferring document relevance from incomplete information, in: Proceedings
of the Sixteenth ACM Conference on Conference on Information and Knowledge Management,
CIKM ’07, Association for Computing Machinery, 2007, p. 633–642.
[11] B. Carterette, J. Allan, R. Sitaraman, Minimal test collections for retrieval evaluation, in:
Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in
Information Retrieval, SIGIR ’06, Association for Computing Machinery, New York, NY, USA, 2006,
p. 268–275.
[12] A. Lipani, J. Palotti, M. Lupu, F. Piroi, G. Zuccon, A. Hanbury, Fixed-cost pooling strategies based
on ir evaluation measures, in: Advances in Information Retrieval, 2017, pp. 357–368.
[13] A. Mofat, W. Webber, J. Zobel, Strategic system comparisons via targeted relevance judgments,
in: Proceedings of the 30th Annual International ACM SIGIR Conference on Research and
Development in Information Retrieval, SIGIR ’07, 2007, p. 375–382.
[14] E. M. Voorhees, N. Craswell, J. Lin, Too many relevants: Whither cranfield test collections?, in:
Proceedings of the 45th International ACM SIGIR Conference on Research and Development in
Information Retrieval, SIGIR ’22, Association for Computing Machinery, 2022, p. 2970–2980.
[15] S. MacAvaney, L. Soldaini, One-shot labeling for automatic relevance estimation, in: Proceedings
of the 46th International ACM SIGIR Conference on Research and Development in Information
Retrieval, 2023, pp. 2230–2235.
[16] C. Meng, N. Arabzadeh, A. Askari, M. Aliannejadi, M. de Rijke, Query performance prediction
using relevance judgments generated by large language models, arXiv preprint arXiv:2404.01012
(2024).
[17] E. Khramtsova, S. Zhuang, M. Baktashmotlagh, G. Zuccon, Leveraging llms for unsupervised dense
retriever ranking, in: Proceedings of the 47th International ACM SIGIR Conference on Research
and Development in Information Retrieval, SIGIR ’24, Association for Computing Machinery, 2024,
p. 1307–1317.
[18] P. Thomas, S. Spielman, N. Craswell, B. Mitra, Large language models can accurately predict
searcher preferences, in: Proceedings of the 47th International ACM SIGIR Conference on Research
and Development in Information Retrieval, 2024, pp. 1930–1940.
[19] G. Faggioli, L. Dietz, C. L. Clarke, G. Demartini, M. Hagen, C. Hauf, N. Kando, E. Kanoulas,
M. Potthast, B. Stein, et al., Perspectives on large language models for relevance judgment, in:
ICTIR, 2023, pp. 39–50.
[20] R. Pradeep, S. Sharifymoghaddam, J. Lin, Rankzephyr: Efective and robust zero-shot listwise
reranking is a breeze!, arXiv preprint arXiv:2312.02724 (2023).
[21] H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma,
et al., Scaling instruction-finetuned language models, Journal of Machine Learning Research 25
(2024) 1–53.
[22] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal,
E. Hambro, F. Azhar, et al., Llama: Open and eficient foundation language models, arXiv preprint
arXiv:2302.13971 (2023).
[23] T. Dettmers, A. Pagnoni, A. Holtzman, L. Zettlemoyer, Qlora: eficient finetuning of quantized
llms, in: Proceedings of the 37th International Conference on Neural Information Processing
Systems, NIPS ’23, Curran Associates Inc., 2024.
[24] A. J. Peterson, Ai and the problem of knowledge collapse, 2024. arXiv:2404.03502.
[25] F. Radlinski, N. Craswell, A theoretical framework for conversational search, in: Proceedings
of the 2017 Conference on Conference Human Information Interaction and Retrieval, CHIIR ’17,
Association for Computing Machinery, 2017, p. 117–126.
[26] L. Azzopardi, M. Dubiel, M. Halvey, J. Dalton, A conceptual framework for conversational search
and recommendation: Conceptualizing agent-human interactions during the conversational search
process, in: Proceedings of the CAIR’18: Second International Workshop on Conversational
Approaches to Information Retrieval at SIGIR 2018, 2018.
[27] C. Meng, N. Arabzadeh, M. Aliannejadi, M. de Rijke, Query performance prediction: From ad-hoc
to conversational search, in: SIGIR, 2023, p. 2583–2593.
[28] C. Meng, M. Aliannejadi, M. de Rijke, System initiative prediction for multi-turn conversational
information seeking, in: CIKM, 2023, pp. 1807–1817.
[29] M. Aliannejadi, Z. Abbasiantaeb, S. Chatterjee, J. Dalton, L. Azzopardi, Trec ikat 2023: A test
collection for evaluating conversational and interactive knowledge assistants, in: Proceedings
of the 47th International ACM SIGIR Conference on Research and Development in Information
Retrieval, SIGIR ’24, Association for Computing Machinery, 2024, p. 819–829.
[30] Z. Abbasiantaeb, M. Aliannejadi, Generate then retrieve: Conversational response retrieval using
llms as answer and query generators, CoRR abs/2403.19302 (2024). arXiv:2403.19302.
[31] AI@Meta, Llama 3 model card (2024). URL: https://github.com/meta-llama/llama3/blob/main/</p>
      <p>MODEL_CARD.md.
[32] P. Bajaj, D. Campos, N. Craswell, L. Deng, X. L. Jianfeng Gao, R. Majumder, A. McNamara, B. Mitra,
T. Nguyen, M. Rosenberg, X. Song, A. Stoica, S. Tiwary, T. Wang, Ms marco: A human generated
machine reading comprehension dataset, in: NIPS, 2016.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C. W.</given-names>
            <surname>Cleverdon</surname>
          </string-name>
          ,
          <article-title>The cranfield tests on index language devices</article-title>
          ,
          <source>Aslib Proceedings 19</source>
          (
          <year>1967</year>
          )
          <fpage>173</fpage>
          -
          <lpage>194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sparck-Jones</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. J. Van Rijsbergen</surname>
          </string-name>
          ,
          <article-title>Report on the Need for and Provision of an 'Ideal' Information Retrieval Test Collection</article-title>
          ,
          <source>Technical Report British Library Research and Development Report No. 5266</source>
          , Computer Laboratory, University of Cambridge,
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. K.</given-names>
            <surname>Harman</surname>
          </string-name>
          ,
          <article-title>TREC: Experiment and Evaluation in Information Retrieval (Digital Libraries</article-title>
          and Electronic Publishing), The MIT Press,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Cormack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Palmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L. A.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <article-title>Eficient construction of large test collections</article-title>
          ,
          <source>in: Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR '98,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery,
          <year>1998</year>
          , p.
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Buckley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          ,
          <article-title>Retrieval evaluation with incomplete information</article-title>
          ,
          <source>in: Proceedings of the 27th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '04</source>
          ,
          <year>2004</year>
          , p.
          <fpage>25</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zobel</surname>
          </string-name>
          ,
          <article-title>Information retrieval system evaluation: efort, sensitivity</article-title>
          , and reliabil-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>