<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>User Simulation⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xiao Fu</string-name>
          <email>xiao.fu.20@ucl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aldo Lipani</string-name>
          <email>aldo.lipani@ucl.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noriko Kando</string-name>
          <email>Noriko.Kando@nii.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Conversation Information Retrieval, Evaluation, User Simulation, Large Language Models</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of Informatics (NII)</institution>
          ,
          <addr-line>Tokyo 101-8430</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University College London (UCL)</institution>
          ,
          <addr-line>Gower Street, London, WC1E 6BT</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recent advancements in the field of Conversational Information Retrieval (CIR) have increased the demand for more sophisticated modelling and evaluation approaches. This paper introduces a novel framework for user simulation in CIR, aimed at enhancing the modeling and evaluation of user interactions. Additionally, this study explores the potential integration of large language models (LLMs) within this domain. Furthermore, the paper anticipates future developments in CIR, particularly in the context of the widespread use of LLMs. The study emphasizes the necessity for robust evaluation paradigms that go beyond traditional methods to efectively measure the success of CIR systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Recent advancements in Machine Learning (ML), Natural Language Processing (NLP), and the
proliferation of smart devices have significantly enhanced conversational AI. This progress has led to a variety
of commercial conversational services that enable natural spoken interactions, thereby increasing the
demand for more human-centric approaches in information retrieval (IR) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The objective of
Conversational Information Retrieval (CIR) is to facilitate information seeking through multi-turn natural
language dialogues between users and systems, a longstanding yet challenging area of research [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
CIR. Traditional IR systems typically focus solely on the user’s current query, treating each query as an
independent event. These systems do not account for the influence of previous queries on the current
search.
      </p>
      <p>CEUR</p>
      <p>ceur-ws.org</p>
      <p>
        In contrast, CIR systems emphasize the importance of context in the retrieval process. Previous
queries and responses influence the current response. The second part of Figure 1 shows a conversation
between a user and a CIR system from the FaithDial dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], where the interaction is more natural,
and queries form sequences akin to conversations. For instance, the word they in the second turn of
the conversation refers to the shops mentioned in the system’s previous response.
      </p>
      <p>This shift in communication style introduces several diferences between traditional IR and CIR. The
conversation-like query format in CIR allows for richer, more interactive responses. Firstly, CIR systems
can provide more detailed responses compared to traditional IR, which typically returns documents
directly. CIR systems can refine information needs and the search space through user or system
revealment. Secondly, users have more strategic options in their interactions. If unsatisfied with a
response, users can ask follow-up questions to refine their search or request additional information.
This necessitates the introduction of ad-hoc search methods (such as session-based and task-based
search) in CIR, enabling the system to refine the search space based on the conversation, querying
related sections instead of the entire database. Thirdly, context is crucial in CIR for understanding user
queries, unlike traditional IR where queries are treated independently. For example, the meaning of
pronouns can change depending on their order in a conversation.</p>
      <p>
        Currently, CIR systems are predominantly used for simple tasks, as they are not suficiently
effective for complex and exploratory information-seeking conversations. Nonetheless, advancements
in key components, particularly in ML, are driving a trend within the IR community towards more
conversational methodologies [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>
        Despite the progress, several critical questions remain unresolved, presenting challenges to the
development of CIR systems. Zamani et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] identifies four key directions in the CIR field with
potential for significant advancements:
1. Modelling and Producing Conversational Interactions
2. Result Presentation
3. Exploring Under-Explored Conversational Tasks
4. Measuring Interaction Success and Evaluation
      </p>
      <p>This paper primarily focuses on the first and last directions, which are both challenging and
interconnected. The first direction involves modelling and producing conversational interactions, addressing
the uncertainty of information needs between multiple agents through a mixed-initiative approach.
Additionally, understanding long-term conversational interactions and addressing associated privacy
and transparency concerns are critical topics in this direction.</p>
      <p>The final direction pertains to the measurement and evaluation of CIR. Both academia and industry
face limitations due to the absence of a robust definition of success in this field. As CIR continues to
evolve, there is an urgent need for an evaluation paradigm that transcends the traditional Cranfield
Paradigm, especially in frontier tasks such as personalized evaluation and transparency.</p>
      <p>These two directions are highly interconnected. The definition of success relies on the proper
modelling of conversational interactions, while precise measures support the modelling process. This
paper introduces a new potential contribution to this field: introducing user simulation to CIR.</p>
      <p>In this paper, we present a framework for the automated evaluation of CIR systems via user
simulation. This paper also includes the recent advancements within this framework, while exploring the
encountered opportunities and challenges. Section 3 details the two-stage user simulation prototype,
which amalgamates psychological principles and ML to enhance explainability, and utilizes Large
Language Models (LLMs) to sustain high performance. Section 4 focuses on the application of this
prototype in assessing CIR systems. The methodology proposed aims to connect user simulation with
well-established CIR conversation modelling approaches, such as efort and cost, and seeks to align the
simulations closely with real user interactions through indirect assessment. Section 5 discusses the
challenges and future perspectives in this field.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <sec id="sec-2-1">
        <title>2.1. Evaluation</title>
        <p>
          Conversational search, a well-established field, continues to be a popular research topic due to its
relevance for modern devices with small or no screens [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>
          Despite significant progress, the evaluation of CIR remains relatively underdeveloped [
          <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
          ]. While
CIR extends functionalities from traditional IR systems [9], many studies still rely on conventional
metrics such as Mean Average Precision (MAP), Normalized Discounted Cumulative Gain (nDCG),
and Mean Reciprocal Rank (MRR) [
          <xref ref-type="bibr" rid="ref7">7, 10, 11</xref>
          ]. At the heart of these metrics is the concept of relevance,
where the documents retrieved are fitting to the topics (keywords) sought by the user [ 12]. Additionally,
metrics from other domains like ROUGE and BLEU are also utilized [
          <xref ref-type="bibr" rid="ref6">6, 13, 14</xref>
          ]. However, recent research
indicates that without real user interaction, these metrics may not accurately reflect user satisfaction
[15, 16].
        </p>
        <p>
          User satisfaction, a highly abstract and subjective measure, pertains to the overall experience and
interaction with a search system [
          <xref ref-type="bibr" rid="ref8">10, 8</xref>
          ]. It is defined as the fulfillment users achieve in pursuing their
goals [17]. Extensive research has been conducted to understand this measure [10, 15, 18, 19, 20, 21, 22].
Studies like Yilmaz et al. [18] ofer various metrics to reflect user satisfaction, considering factors such as
efort. Some studies further break down user satisfaction into query-level satisfactions [ 19, 20], although
others, such as Järvelin et al. [21], argue that summing query-level satisfaction misses the contextual
journey where user queries are interconnected. Unlike traditional search systems, conversational search
systems allow users to ask follow-up questions to refine their answers [ 22].
        </p>
        <p>
          Despite widespread adoption, measuring user satisfaction remains an open question. Many studies
rely on real-time user participation to gather feedback, providing fresh and realistic insights but
requiring significant resources and participant incentives [ 10, 15]. Alternatively, satisfaction prediction
proxies using deep learning models ofer a computational approach, overcoming temporal and spatial
constraints but demanding high-quality computational resources and datasets [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          Emerging trends in user simulation ofer promising solutions to these challenges [
          <xref ref-type="bibr" rid="ref2 ref5">5, 2</xref>
          ], though prior
studies are not yet comprehensive. Gao et al. [23] noted that earlier simulators relied on randomly
generated scenarios to approximate users’ states of mind. While such models prove beneficial in specific
domains like recommender systems [24], issues of explainability and scrutability remain unresolved.
Our proposed framework addresses these issues by incorporating a user-profile-based personalized
simulation approach, facilitating easier alignment with actual user behaviours and providing a scrutable
means to control the simulation by modifying textual user profiles.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. User Simulation</title>
        <p>Azzopardi et al. [25] define simulation as the imitation of the operation of real-world phenomena.
Simulations enable detailed experimental design and control tailored to specific research questions.
These high-level controls allow for experiments with user simulations to be conducted with several
advantages [25].</p>
        <p>Firstly, what-if experiments can be performed by setting up diferent scenarios [26, 27]. Secondly,
user simulators ensure the repeatability of experimental results. Additionally, user simulations can
achieve these benefits at a low cost [ 28].</p>
        <p>In the IR community, user simulation methods are primarily divided into cognitive and statistical
approaches [29]. Cognitive approaches were among the first used in this field. Belkin [30] described
users, information resources, and IR models, characterizing users by their objectives, problems, and
knowledge. Subsequent studies expanded on this foundation [31, 32, 33].</p>
        <p>In contrast, statistical approaches focus on analyzing user behaviours and satisfaction [29, 34, 35, 36,
37]. These approaches underpin early user simulators based on statistical models [38, 39]. Although
these simulators heavily relied on corpora, they faced limitations such as the diversity of user intent
[29].</p>
        <p>Agenda-based user simulations are popular due to their realistic responses and straightforward
dialogue strategies [40, 41, 42]. The latest trend involves employing deep learning models, including
adversarial generative approaches [43], reinforcement learning [44], and inverse reinforcement learning
to abstract knowledge from data [45].</p>
        <p>
          Evaluating CIR with user simulations is becoming a key trend, enabling eficient and cost-efective
evaluation at various levels of CIR [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Our proposed prototype integrates benefits from the
aforementioned approaches. Initially, it simulates user behaviour guided by statistical signals derived from the
context, enhancing explainability. Subsequently, the prototype employs deep learning models, utilizing
textual user profiles to produce realistic and diverse responses within controlled parameters, thereby
enabling further exploration of the target system. The details of the user simulation prototype are in
the next section.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Modelling and Simulating Users</title>
      <p>In this section, we introduce an two-stage prototype for constructing a robust user simulation for CIR.</p>
      <p>The main structure of the prototype is divided into two parts to control simulated users: the Action
Predictor and the Response Generator, as illustrated in Fig 2.</p>
      <sec id="sec-3-1">
        <title>3.1. Modeling Actions</title>
        <p>The Action Predictor controls the simulated users’ actions during the current conversation. This concept
originates from the psychology community’s notion of priming, which refers to the unconscious
influence of past experiences on current performance or behaviour [50, 51]. Many studies have established
models to explain this mechanism [52, 50]. For instance, Tulving et al. [53] conducted an experiment
where participants viewed a list of 96 words and later completed graphemic word fragments both one
hour and seven days after studying the list. The results demonstrated a significant influence of the
word list on subsequent tests.</p>
        <p>As shown in Table 1, results from Fu and Lipani [46] demonstrate the potential of predicting users’
next actions. The benefits of using lexical and textual patterns to predict users’ next actions include
low cost and ease of interpretation, which are valuable for evaluation.</p>
        <p>In the current prototype, three user actions are modelled based on the dataset available. Stopping
is defined as the action where users opt to end a conversation, typically indicating the conclusion of
the exchange. This action is critical for evaluating efects such as the principle of least efort [</p>
        <p>Sessions of conversation often include</p>
        <p>Following up on queries that build upon previous interactions,
acknowledging missing contexts and references to earlier discussed topics [49]. As noted by Stede and
Schlangen [55], an inquisitive user engaged in an ongoing dialogue may express interest in further
related subjects as a response to the information provided. Switching topics is commonly seen in
information-seeking dialogues, especially when using search systems for data acquisition [56].</p>
        <p>As output from the Action Predictor, a general action as described above will be predicted, with the
details elaborated in the Response Generator.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Personalized Responses</title>
        <p>The Response Generator will generate realistic and diversified responses based on the conversation,
user profiles and the predicted actions from the Action Predictor</p>
        <p>LLMs such as GPTs and Llama, renowned for their sophisticated natural language processing
capabilities, present unique opportunities to enhance Conversational Systems through mechanisms like
pre-training, fine-tuning, and prompting [ 57]. The ability of LLMs to mimic diverse demographic
characteristics ofers a novel approach to simulating user behaviour and preferences [ 58]. High-quality
user simulations, which closely mirror real user behaviour distributions, can significantly advance
CRS development, currently dependent on real data for training, with its inherent constraints and
disadvantages.</p>
        <p>Ramos et al. [59] ofers a valuable method for generating user profiles from the Amazon dataset,
introducing a more compact style of personalization into the user simulator. Table 2 demonstrates the
agreement between simulated and real users’ responses to the same items in the Amazon dataset. Since
the simulated users are based on LLMs and prompts with user profiles, the results suggest the potential
of generating personalized responses based on textual user profiles.
Agreements (Cohen’s   , Randolph’s   and Krippendorf’s  ) between simulated users and actual users in
response.Cohen’s index considers the marginal distribution of categories, Randolph’s assumes a uniform
distribution, and Krippendorf’s ofers a broader approach to assessing agreement. Fair agreements for each metric
(&gt;0.2) are marked as bold. Three simulation methods are evaluated: Real UP, where users are simulated based on
their profiles; Rand UP, involving users simulated with random profiles; and Rand Sc, where scores are randomly
generated based on the dataset’s historical distribution.</p>
        <p>Setting
Real UP
Rand UP
Rand Sc</p>
        <p>In this simulation, responses are tailored based on user profiles, which consist of concise text that
summarizes the attributes of users in a few succinct sentences. This can include motivations for
task-oriented CIR systems. Furthermore, the Response Generator is tasked with handling clarification
questions posed by the CIR system.</p>
        <p>At the conclusion of this phase, the user simulator is equipped to interact with CIR systems. The
subsequent section proposes a linkage as the remaining component of this framework, specifically
addressing the evaluation of the CIR system using this user simulator, given the absence of a direct
indicator from the user simulation on the quality of the target CIR system.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation of CIR</title>
      <p>This section discusses potential methods for applying evaluation tasks based on the user simulation
prototype established in Section 3.</p>
      <p>As introduced in Section 2, the target of evaluation originates from modelling users and conversations.
According to Yilmaz et al. [18], user satisfaction can be reflected by the efort exerted. Efort also forms
the foundation of modelling user actions, particularly stopping behaviours.</p>
      <p>In the IR community, several studies have been conducted to depict stopping behaviours [60, 61, 62, 63].
These studies aim to quantify the feeling of having ”enough.” For instance, users may decide to stop a
conversation when they feel frustrated or satisfied.</p>
      <p>A previous study by Fu and Lipani [46] provided a reliable method for predicting stopping behaviours.
The subsequent step is to explore the relationship between each stopping point and user satisfaction.</p>
      <sec id="sec-4-1">
        <title>4.1. Evaluating User Simulation</title>
        <p>The final part of the evaluation focuses on assessing the user simulation. To efectively perform
evaluation for CIR, the user simulation must exhibit not only a diversity of reasonable responses but
also a high alignment with real users, particularly in reflecting satisfaction or frustration.</p>
        <p>Following the study by Fu et al. [28], where direct and indirect assessments in CIR show substantial
agreement, the alignment between user simulation and real users can be evaluated. Real users will
review the conversations between the user simulation and the system to determine if the simulated
user appears satisfied.</p>
        <p>Figure 3 illustrates the operational flow within the evaluation component of the framework. In this
component, a classifier, integrated with the user simulator, is utilized to predict user satisfaction. This
classifier is aligned with annotations from real users. Feedback from these real annotators is employed
to train both the user simulator and the classifier. The user simulator aims to accurately mimic real
user behaviour through the Actions Predictor and the Response Generator. Similarly, the classifier is
trained to align its judgments with those of real users regarding the same conversation.</p>
        <p>During the development of this framework, numerous emerging trends were observed, particularly
in LLMs. These observations are presented in the following section.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Towards Future</title>
      <p>This represents a significant transformation since 2021, as the term Large Language Models (LLMs)
has gained popularity. This shift has introduced both opportunities and challenges for contemporary
research.
5.1. CIR and RAG
The integration of LLMs is not limited to the IR community; the ML community is also embracing IR
techniques. A prominent development in this space is Retrieval-Augmented Generation (RAG), which
enhances LLMs in domain-specific or knowledge-intensive tasks [ 64].</p>
      <p>RAG involves multiple retrieval processes to enrich the context, going beyond traditional single
retrieval methods. For instance, Self-RAG [65] refines the RAG framework by enabling LLMs to actively
determine the optimal moments and content for retrieval, thus improving the eficiency and relevance
of the sourced information. The key aspect here is not merely multiple retrievals, but the reliance on
the judgment of LLMs, indicating that LLMs can further participate in the processing with minimal
human intervention.</p>
      <p>For the evaluation of RAG, integrating typical LLMs alone may not sufice. A realistic inquiry is how
agents based on LLMs can assess the responses from RAG systems that incorporate retrieved documents.
One feasible approach is using RAG to evaluate itself. Here, the crucial aspects include not only the
quality of the conversation and the retrieval process but also how efectively the documents are presented .</p>
      <p>Moreover, in the specialized domain of CIR, where the systems are relatively light, LLMs can still
serve as experts. The following section provides an example.</p>
      <sec id="sec-5-1">
        <title>5.2. Should We Ask LLMs First?</title>
        <p>The TREC Interactive Knowledge Assistance Track (iKAT) builds upon the foundational work of the
TREC Conversational Assistance Track (CAsT) [66], with a key diference being the addition of personal
context for each user in the dataset. The primary task remains similar to TREC CAsT—retrieving and
ranking documents from the corpus at each turn of the given conversations.</p>
        <p>The best performance in iKAT 2023 introduced a novel approach [67]. In this approach, the LLM
generates an initial answer to the user’s query based on the context of the conversation and the user
profile. This answer is derived through reasoning over the context and the user’s profile, but it is not
grounded in the documents within the collection. Subsequently, the LLM generates a set of five queries
to achieve this answer.</p>
        <p>The controversial aspect of this approach is using the LLM-generated answer as the target without
initial retrieval, followed by employing an IR system to achieve it. This process assumes that LLMs’
answers are suficiently accurate. Alternatively, it suggests that the documents have likely been exposed
to the LLMs, raising concerns of potential data leakage.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.3. How Can We Go Beyond Our Knowledge Borders?</title>
        <p>Training LLMs from scratch is a challenging task for most research groups due to high costs and the
lack of storage and computational resources. The most common practice involves fine-tuning a public
base version of LLMs and accessing them via APIs. Top LLMs are trained on a substantial portion of
internet text documents, making it nearly impossible to prevent data leakage once any public data is
used in a study.</p>
        <p>Furthermore, the widespread use of LLMs will inevitably introduce LLM-generated text back into
the internet, posing a significant challenge that has already raised considerable concerns within the
community.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This paper proposes a framework for evaluating CIR systems automatically using a user simulator
prototype. The framework comprises two main components:
1. A prototype of user simulation that leverages the advancements from both psychology and ML
ifelds to conduct realistic and scrutable simulations targeted at CIR systems.
2. A component that employs sophisticated conversation modelling concepts from the IR community
to provide reasonable feedback aimed at predicting user satisfaction alongside the user simulation
prototype.</p>
      <p>In addition to the framework, this study also presents emerging trends observed with the promising
development of LLMs. The advent of systems such as RAG introduces both opportunities and challenges.
It raises concerns that the current practices within the community utilizing LLMs may lead to increased
data leakages.
[9] A. Anand, L. Cavedon, H. Joho, M. Sanderson, B. Stein, Conversational Search (Dagstuhl Seminar
19461), Dagstuhl Reports 9 (2020). doi:10.4230/DagRep.9.11.34.
[10] J. Kiseleva, K. Williams, A. Hassan Awadallah, A. C. Crook, I. Zitouni, T. Anastasakos, Predicting
user satisfaction with intelligent assistants, in: Proceedings of the 39th International ACM SIGIR
Conference on Research and Development in Information Retrieval, SIGIR ’16, Association for
Computing Machinery, New York, NY, USA, 2016. doi:10.1145/2911451.2911521.
[11] J. Dalton, C. Xiong, J. Callan, Trec cast 2019: The conversational assistance track overview, 2020.
[12] T. Saracevic, Relevance reconsidered, in: Proceedings of the second conference on conceptions of
library and information science (CoLIS 2), 1996, pp. 201–218.
[13] K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, Bleu: A method for automatic evaluation of machine
translation, in: Proceedings of the 40th Annual Meeting on Association for Computational
Linguistics, ACL ’02, Association for Computational Linguistics, USA, 2002. doi:10.3115/1073083.
1073135.
[14] E. Reiter, A structured review of the validity of BLEU, Computational Linguistics 44 (2018).</p>
      <p>doi:10.1162/coli_a_00322.
[15] J. Kiseleva, K. Williams, J. Jiang, A. Hassan Awadallah, A. C. Crook, I. Zitouni, T. Anastasakos,
Understanding user satisfaction with intelligent assistants, in: Proceedings of the 2016 ACM
on Conference on Human Information Interaction and Retrieval, CHIIR ’16, Association for
Computing Machinery, New York, NY, USA, 2016. doi:10.1145/2854946.2854961.
[16] C.-W. Liu, R. Lowe, I. Serban, M. Noseworthy, L. Charlin, J. Pineau, How NOT to evaluate your
dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response
generation, in: Proceedings of the 2016 Conference on Empirical Methods in Natural Language
Processing, Association for Computational Linguistics, Austin, Texas, 2016. doi:10.18653/v1/
D16- 1230.
[17] D. Kelly, Methods for evaluating interactive information retrieval systems with users, Foundations
and Trends® in Information Retrieval 3 (2009). doi:10.1561/1500000012.
[18] E. Yilmaz, M. Verma, N. Craswell, F. Radlinski, P. Bailey, Relevance and efort: An analysis of
document utility, in: Proceedings of the 23rd ACM International Conference on Conference on
Information and Knowledge Management, CIKM ’14, Association for Computing Machinery, New
York, NY, USA, 2014, pp. 91–100. doi:10.1145/2661829.2661953.
[19] J. Kiseleva, E. Crestan, R. Brigo, R. Dittel, Modelling and detecting changes in user satisfaction,
in: Proceedings of the 23rd ACM International Conference on Conference on Information and
Knowledge Management, CIKM ’14, Association for Computing Machinery, New York, NY, USA,
2014, pp. 1449–1458. doi:10.1145/2661829.2661960.
[20] J. Kiseleva, J. Kamps, V. Nikulin, N. Makarov, Behavioral dynamics from the serp’s perspective:
What are failed serps and how to fix them?, in: Proceedings of the 24th ACM International on
Conference on Information and Knowledge Management, CIKM ’15, Association for Computing
Machinery, New York, NY, USA, 2015, pp. 1561–1570. doi:10.1145/2806416.2806483.
[21] K. Järvelin, S. L. Price, L. M. L. Delcambre, M. L. Nielsen, Discounted cumulated gain based
evaluation of multiple-query ir sessions, in: C. Macdonald, I. Ounis, V. Plachouras, I. Ruthven,
R. W. White (Eds.), Advances in Information Retrieval, Springer, Springer Berlin Heidelberg, Berlin,
Heidelberg, 2008, pp. 4–15.
[22] A. Al-Maskari, M. Sanderson, P. Clough, The relationship between ir efectiveness measures
and user satisfaction, in: Proceedings of the 30th Annual International ACM SIGIR Conference
on Research and Development in Information Retrieval, SIGIR ’07, Association for Computing
Machinery, New York, NY, USA, 2007, pp. 773–774. doi:10.1145/1277741.1277902.
[23] J. Gao, M. Galley, L. Li, Neural approaches to conversational ai, in: The 41st international ACM</p>
      <p>SIGIR conference on research &amp; development in information retrieval, 2018, pp. 1371–1374.
[24] S. Zhang, K. Balog, Evaluating conversational recommender systems via user simulation, in:
Proceedings of the 26th acm sigkdd international conference on knowledge discovery &amp; data
mining, 2020, pp. 1512–1520.
[25] L. Azzopardi, K. Järvelin, J. Kamps, M. D. Smucker, Report on the sigir 2010 workshop on the
simulation of interaction, SIGIR Forum 44 (2011) 35–47. URL: https://doi.org/10.1145/1924475.
1924484. doi:10.1145/1924475.1924484.
[26] M. I. Kellner, R. J. Madachy, D. M. Rafo, Software process simulation modeling: why? what?
how?, Journal of Systems and Software 46 (1999) 91–105.
[27] K. Balog, D. Maxwell, P. Thomas, S. Zhang, Report on the 1st simulation for information retrieval
workshop (sim4ir 2021) at sigir 2021, SIGIR Forum 55 (2022). URL: https://doi.org/10.1145/3527546.
3527559. doi:10.1145/3527546.3527559.
[28] X. Fu, E. Yilmaz, A. Lipani, Evaluating the cranfield paradigm for conversational search systems, in:
Proceedings of the 2022 ACM SIGIR International Conference on Theory of Information Retrieval,
2022, pp. 275–280.
[29] P. Erbacher, L. Soulier, L. Denoyer, State of the art of user simulation approaches for conversational
information retrieval, arXiv preprint arXiv:2201.03435 (2022).
[30] N. Belkin, Cognitive models and information transfer, Social Science Information Studies 4 (1984)
111–129. URL: https://www.sciencedirect.com/science/article/pii/014362368490070X. doi:https://
doi.org/10.1016/0143- 6236(84)90070- X, special Issue Seminar on the Psychological Aspects
of Information Searching.
[31] C. C. Kuhlthau, Inside the search process: Information seeking from the user’s perspective, J. Am.</p>
      <p>Soc. Inf. Sci. 42 (1991) 361–371.
[32] P. Ingwersen, K. Järvelin, The Turn: Integration of Information Seeking and Retrieval in Context,
2005. doi:10.1007/1- 4020- 3851- 8.
[33] D. Ellis, A behavioral approach to information retrieval system design, J. Doc. 45 (1989) 171–212.</p>
      <p>URL: https://doi.org/10.1108/eb026843. doi:10.1108/eb026843.
[34] N. Craswell, O. Zoeter, M. Taylor, B. Ramsey, An experimental comparison of click position-bias
models, in: Proceedings of the international conference on Web search and web data mining,
WSDM ’08, ACM, New York, NY, USA, 2008, pp. 87–94. URL: http://doi.acm.org/10.1145/1341531.
1341545. doi:10.1145/1341531.1341545.
[35] O. Chapelle, Y. Zhang, A dynamic bayesian network click model for web search ranking, in: In</p>
      <p>WWW, 2009.
[36] G. E. Dupret, B. Piwowarski, A user browsing model to predict search engine click data from past
observations., in: SIGIR ’08: Proceedings of the 31st annual international ACM SIGIR conference on
Research and development in information retrieval, ACM, New York, NY, USA, 2008, pp. 331–338.</p>
      <p>URL: http://portal.acm.org/citation.cfm?id=1390334.1390392. doi:10.1145/1390334.1390392.
[37] A. Chuklin, P. Serdyukov, M. Rijke, Modeling clicks beyond the first result page, 2013. doi: 10.</p>
      <p>1145/2505515.2507859.
[38] W. Eckert, E. Levin, R. Pieraccini, User modeling for spoken dialogue system evaluation, 1997</p>
      <p>IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings (1997) 80–87.
[39] K. Schefler, S. J. Young, Probabilistic simulation of human-machine dialogues, 2000 IEEE
International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100)
2 (2000) II1217–II1220 vol.2.
[40] K. Komatani, S. Ueno, T. Kawahara, H. Okuno, User modeling in spoken dialogue systems to
generate flexible guidance, User Modeling and User-Adapted Interaction 15 (2005) 169–183.
doi:10.1007/s11257- 004- 5659- 0.
[41] J. Schatzmann, B. Thomson, K. Weilhammer, H. Ye, S. Young, Agenda-based user simulation for
bootstrapping a pomdp dialogue system, 2007, pp. 149–152. doi:10.3115/1614108.1614146.
[42] X. Li, Z. C. Lipton, B. Dhingra, L. Li, J. Gao, Y.-N. Chen, A user simulator for task-completion
dialogues., CoRR abs/1612.05688 (2016). URL: http://dblp.uni-trier.de/db/journals/corr/corr1612.
html#LiLDLGC16.
[43] X. Zhao, L. Xia, Z. Ding, D. Yin, J. Tang, Toward simulating environments in reinforcement
learning based recommendations, ArXiv abs/1906.11462 (2019).
[44] X. Chen, S. Li, H. Li, S. Jiang, Y. Qi, L. Song, Generative adversarial user model for reinforcement
learning based recommendation system, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings
of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine
Learning Research, PMLR, 2019, pp. 1052–1061. URL: https://proceedings.mlr.press/v97/chen19f.
html.
[45] S. Chandramohan, M. Geist, F. Lefèvre, O. Pietquin, User simulation in dialogue systems using
inverse reinforcement learning, in: INTERSPEECH, 2011.
[46] X. Fu, A. Lipani, Priming and actions: An analysis in conversational search systems, SIGIR, 2023.
[47] V. Adlakha, S. Dhuliawala, K. Suleman, H. de Vries, S. Reddy,
TopiOCQA: Open-domain conversational question answering with topic
switching, volume 10, 2022, pp. 468–483. URL: https://doi.org/10.1162/tacl_a_00471.
doi:10.1162/tacl_a_00471.
arXiv:https://direct.mit.edu/tacl/articlepdf/doi/10.1162/tacl_a_00471/2008126/tacl_a_00471.pdf.
[48] R. Anantha, S. Vakulenko, Z. Tu, S. Longpre, S. Pulman, S. Chappidi, Open-domain question
answering goes conversational via question rewriting, Proceedings of the 2021 Conference of the
North American Chapter of the Association for Computational Linguistics: Human Language
Technologies (2021).
[49] C. Qu, L. Yang, C. Chen, M. Qiu, W. B. Croft, M. Iyyer, Open-Retrieval Conversational Question</p>
      <p>Answering, in: SIGIR, 2020.
[50] D. L. Schacter, R. L. Buckner, Priming and the brain, Neuron 20 (1998) 185–195.
[51] J. A. Bargh, T. L. Chartrand, The mind in the middle: A practical guide to priming and automaticity
research. (2014).
[52] C. Janiszewski, R. S. Wyer Jr, Content and process priming: A review, Journal of consumer
psychology 24 (2014) 96–118.
[53] E. Tulving, D. L. Schacter, H. A. Stark, Priming efects in word-fragment completion are independent
of recognition memory., Journal of experimental psychology: learning, memory, and cognition 8
(1982) 336.
[54] G. K. Zipf, Human behavior and the principle of least efort: An introduction to human ecology,</p>
      <p>Ravenio Books, 2016.
[55] M. Stede, D. Schlangen, Information-seeking chat : Dialogue management by topic structure,
2004.
[56] A. Spink, H. Özmutlu, S. Özmutlu, Multitasking information seeking and searching processes,</p>
      <p>JASIST 53 (2002) 639–652. doi:10.1002/asi.10124.
[57] W. Fan, Z. Zhao, J. Li, Y. Liu, X. Mei, Y. Wang, J. Tang, Q. Li, Recommender systems in the era of
large language models (llms), arXiv preprint arXiv:2307.02046 (2023).
[58] G. V. Aher, R. I. Arriaga, A. T. Kalai, Using large language models to simulate multiple humans
and replicate human subject studies, in: International Conference on Machine Learning, PMLR,
2023, pp. 337–371.
[59] J. Ramos, H. A. Rahmani, X. Wang, X. Fu, A. Lipani, Natural language user profiles for transparent
and scrutable recommendations, 2024. arXiv:2402.05810.
[60] D. Maxwell, L. Azzopardi, K. Järvelin, H. Keskustalo, Searching and stopping: An analysis of
stopping rules and strategies, in: Proceedings of the 24th ACM International on Conference
on Information and Knowledge Management, CIKM ’15, Association for Computing Machinery,
New York, NY, USA, 2015, p. 313–322. URL: https://doi.org/10.1145/2806416.2806476. doi:10.1145/
2806416.2806476.
[61] W. S. Cooper, On selecting a measure of retrieval efectiveness part ii. implementation of the
philosophy, Journal of the American Society for information Science 24 (1973) 413–424.
[62] D. H. Kraft, T. Lee, Stopping rules and their efect on expected search length, Information</p>
      <p>Processing &amp; Management 15 (1979) 47–58.
[63] K. R. Nickles, Judgment-based and reasoning-based stopping rules in decision-making under
uncertainty, University of Minnesota, 1995.
[64] Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, H. Wang, Retrieval-augmented
generation for large language models: A survey, arXiv preprint arXiv:2312.10997 (2023).
[65] A. Asai, Z. Wu, Y. Wang, A. Sil, H. Hajishirzi, Self-rag: Learning to retrieve, generate, and critique
through self-reflection, arXiv preprint arXiv:2310.11511 (2023).
[66] M. Aliannejadi, Z. Abbasiantaeb, S. Chatterjee, J. Dalton, L. Azzopardi, Trec ikat 2023: The
interactive knowledge assistance track overview, arXiv preprint arXiv:2401.01330 (2024).
[67] Z. Abbasiantaeb, C. Meng, D. Rau, A. Krasakis, H. A. Rahmani, M. Aliannejadi, Llm-based retrieval
and generation pipelines for trec interactive knowledge assistance track (ikat) 2023, TREC, 2023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bennett</surname>
          </string-name>
          ,
          <article-title>Recent advances in conversational information retrieval</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR '20,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2020</year>
          , p.
          <fpage>2421</fpage>
          -
          <lpage>2424</lpage>
          . URL: https://doi.org/10.1145/3397271.3401418. doi:
          <volume>10</volume>
          .1145/3397271.3401418.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bennett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Craswell</surname>
          </string-name>
          ,
          <article-title>Neural approaches to conversational information retrieval</article-title>
          , volume
          <volume>44</volume>
          ,
          <string-name>
            <surname>Springer</surname>
            <given-names>Nature</given-names>
          </string-name>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Dziri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kamalloo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Milton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Zaiane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ponti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Reddy</surname>
          </string-name>
          ,
          <article-title>Faithdial: A faithful benchmark for information-seeking dialogue, arXiv preprint</article-title>
          ,
          <source>arXiv:2204.10757</source>
          (
          <year>2022</year>
          ). URL: https://arxiv.org/abs/2204.10757.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Penha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hauf</surname>
          </string-name>
          ,
          <article-title>Challenges in the evaluation of conversational search systems</article-title>
          ., in: Converse@ KDD,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zamani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Trippas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dalton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Radlinski</surname>
          </string-name>
          , et al.,
          <source>Conversational information seeking, Foundations and Trends® in Information Retrieval</source>
          <volume>17</volume>
          (
          <year>2023</year>
          )
          <fpage>244</fpage>
          -
          <lpage>456</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lipani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Carterette</surname>
          </string-name>
          , E. Yilmaz,
          <article-title>How am i doing?: Evaluating conversational search systems ofline</article-title>
          ,
          <source>ACM Trans. Inf. Syst</source>
          .
          <volume>39</volume>
          (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1145/3451160.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hassan</surname>
          </string-name>
          <string-name>
            <surname>Awadallah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Ozertem</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Zitouni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Gurunath</given-names>
            <surname>Kulkarni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. Z.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <article-title>Automatic online evaluation of intelligent assistants</article-title>
          ,
          <source>in: Proceedings of the 24th International Conference on World Wide Web, WWW '15, International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE</source>
          ,
          <year>2015</year>
          . doi:
          <volume>10</volume>
          .1145/2736277.2741669.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J. I.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ahmadvand</surname>
          </string-name>
          , E. Agichtein,
          <article-title>Ofline and online satisfaction prediction in open-domain conversational systems</article-title>
          ,
          <source>in: Proceedings of the 28th ACM International Conference on Information and Knowledge Management</source>
          , CIKM '19,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .1145/3357384.3358047.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>