<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">An Evaluation Framework for Conversational Information Retrieval Using User Simulation ⋆</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Xiao</forename><surname>Fu</surname></persName>
							<affiliation key="aff0">
								<orgName type="institution">National Institute of Informatics (NII)</orgName>
								<address>
									<postCode>101-8430</postCode>
									<settlement>Tokyo</settlement>
									<country key="JP">Japan</country>
								</address>
							</affiliation>
							<affiliation key="aff1">
								<orgName type="institution">University College London (UCL)</orgName>
								<address>
									<addrLine>Gower Street</addrLine>
									<postCode>WC1E 6BT</postCode>
									<settlement>London</settlement>
									<country key="GB">UK</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Aldo</forename><surname>Lipani</surname></persName>
							<affiliation key="aff1">
								<orgName type="institution">University College London (UCL)</orgName>
								<address>
									<addrLine>Gower Street</addrLine>
									<postCode>WC1E 6BT</postCode>
									<settlement>London</settlement>
									<country key="GB">UK</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Noriko</forename><surname>Kando</surname></persName>
							<affiliation key="aff0">
								<orgName type="institution">National Institute of Informatics (NII)</orgName>
								<address>
									<postCode>101-8430</postCode>
									<settlement>Tokyo</settlement>
									<country key="JP">Japan</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">An Evaluation Framework for Conversational Information Retrieval Using User Simulation ⋆</title>
					</analytic>
					<monogr>
						<idno type="ISSN">1613-0073</idno>
					</monogr>
					<idno type="MD5">48F72E5C25918F16EF0A1E0D2EAA4E1B</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2025-04-23T16:47+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>Conversation Information Retrieval</term>
					<term>Evaluation</term>
					<term>User Simulation</term>
					<term>Large Language Models</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Recent advancements in the field of Conversational Information Retrieval (CIR) have increased the demand for more sophisticated modelling and evaluation approaches. This paper introduces a novel framework for user simulation in CIR, aimed at enhancing the modeling and evaluation of user interactions. Additionally, this study explores the potential integration of large language models (LLMs) within this domain. Furthermore, the paper anticipates future developments in CIR, particularly in the context of the widespread use of LLMs. The study emphasizes the necessity for robust evaluation paradigms that go beyond traditional methods to effectively measure the success of CIR systems.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Recent advancements in Machine Learning (ML), Natural Language Processing (NLP), and the proliferation of smart devices have significantly enhanced conversational AI. This progress has led to a variety of commercial conversational services that enable natural spoken interactions, thereby increasing the demand for more human-centric approaches in information retrieval (IR) <ref type="bibr" target="#b0">[1]</ref>. The objective of Conversational Information Retrieval (CIR) is to facilitate information seeking through multi-turn natural language dialogues between users and systems, a longstanding yet challenging area of research <ref type="bibr" target="#b1">[2]</ref>. Figure <ref type="figure" target="#fig_0">1</ref> illustrates a comparison between a web search in a traditional IR system and a dialogue in CIR. Traditional IR systems typically focus solely on the user's current query, treating each query as an independent event. These systems do not account for the influence of previous queries on the current search.</p><p>UM-CIR 2024: The 1st Workshop on User Modelling in Conversational Information Retrieval, December 12, 2024, Tokyo, Japan Envelope xiao.fu.20@ucl.ac.uk (X. Fu); aldo.lipani@ucl.ac.uk (A. Lipani); Noriko.Kando@nii.ac.jp (N. Kando) Orcid 0000-0003-4676-8608 (X. Fu); 0000-0002-3643-6493 (A. Lipani); 0000-0002-2133-0215 (N. <ref type="bibr">Kando)</ref> In contrast, CIR systems emphasize the importance of context in the retrieval process. Previous queries and responses influence the current response. The second part of Figure <ref type="figure" target="#fig_0">1</ref> shows a conversation between a user and a CIR system from the FaithDial dataset <ref type="bibr" target="#b2">[3]</ref>, where the interaction is more natural, and queries form sequences akin to conversations. For instance, the word they in the second turn of the conversation refers to the shops mentioned in the system's previous response.</p><p>This shift in communication style introduces several differences between traditional IR and CIR. The conversation-like query format in CIR allows for richer, more interactive responses. Firstly, CIR systems can provide more detailed responses compared to traditional IR, which typically returns documents directly. CIR systems can refine information needs and the search space through user or system revealment. Secondly, users have more strategic options in their interactions. If unsatisfied with a response, users can ask follow-up questions to refine their search or request additional information. This necessitates the introduction of ad-hoc search methods (such as session-based and task-based search) in CIR, enabling the system to refine the search space based on the conversation, querying related sections instead of the entire database. Thirdly, context is crucial in CIR for understanding user queries, unlike traditional IR where queries are treated independently. For example, the meaning of pronouns can change depending on their order in a conversation.</p><p>Currently, CIR systems are predominantly used for simple tasks, as they are not sufficiently effective for complex and exploratory information-seeking conversations. Nonetheless, advancements in key components, particularly in ML, are driving a trend within the IR community towards more conversational methodologies <ref type="bibr" target="#b3">[4,</ref><ref type="bibr" target="#b4">5]</ref>.</p><p>Despite the progress, several critical questions remain unresolved, presenting challenges to the development of CIR systems. Zamani et al. <ref type="bibr" target="#b4">[5]</ref> identifies four key directions in the CIR field with potential for significant advancements: </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Measuring Interaction Success and Evaluation</head><p>This paper primarily focuses on the first and last directions, which are both challenging and interconnected. The first direction involves modelling and producing conversational interactions, addressing the uncertainty of information needs between multiple agents through a mixed-initiative approach. Additionally, understanding long-term conversational interactions and addressing associated privacy and transparency concerns are critical topics in this direction.</p><p>The final direction pertains to the measurement and evaluation of CIR. Both academia and industry face limitations due to the absence of a robust definition of success in this field. As CIR continues to evolve, there is an urgent need for an evaluation paradigm that transcends the traditional Cranfield Paradigm, especially in frontier tasks such as personalized evaluation and transparency.</p><p>These two directions are highly interconnected. The definition of success relies on the proper modelling of conversational interactions, while precise measures support the modelling process. This paper introduces a new potential contribution to this field: introducing user simulation to CIR.</p><p>In this paper, we present a framework for the automated evaluation of CIR systems via user simulation. This paper also includes the recent advancements within this framework, while exploring the encountered opportunities and challenges. Section 3 details the two-stage user simulation prototype, which amalgamates psychological principles and ML to enhance explainability, and utilizes Large Language Models (LLMs) to sustain high performance. Section 4 focuses on the application of this prototype in assessing CIR systems. The methodology proposed aims to connect user simulation with well-established CIR conversation modelling approaches, such as effort and cost, and seeks to align the simulations closely with real user interactions through indirect assessment. Section 5 discusses the challenges and future perspectives in this field.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Related Works</head><p>Conversational search, a well-established field, continues to be a popular research topic due to its relevance for modern devices with small or no screens <ref type="bibr" target="#b0">[1]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.">Evaluation</head><p>Despite significant progress, the evaluation of CIR remains relatively underdeveloped <ref type="bibr" target="#b5">[6,</ref><ref type="bibr" target="#b6">7,</ref><ref type="bibr" target="#b7">8]</ref>. While CIR extends functionalities from traditional IR systems <ref type="bibr" target="#b8">[9]</ref>, many studies still rely on conventional metrics such as Mean Average Precision (MAP), Normalized Discounted Cumulative Gain (nDCG), and Mean Reciprocal Rank (MRR) <ref type="bibr" target="#b6">[7,</ref><ref type="bibr" target="#b9">10,</ref><ref type="bibr" target="#b10">11]</ref>. At the heart of these metrics is the concept of relevance, where the documents retrieved are fitting to the topics (keywords) sought by the user <ref type="bibr" target="#b11">[12]</ref>. Additionally, metrics from other domains like ROUGE and BLEU are also utilized <ref type="bibr" target="#b5">[6,</ref><ref type="bibr" target="#b12">13,</ref><ref type="bibr" target="#b13">14]</ref>. However, recent research indicates that without real user interaction, these metrics may not accurately reflect user satisfaction <ref type="bibr" target="#b14">[15,</ref><ref type="bibr" target="#b15">16]</ref>.</p><p>User satisfaction, a highly abstract and subjective measure, pertains to the overall experience and interaction with a search system <ref type="bibr" target="#b9">[10,</ref><ref type="bibr" target="#b7">8]</ref>. It is defined as the fulfillment users achieve in pursuing their goals <ref type="bibr" target="#b16">[17]</ref>. Extensive research has been conducted to understand this measure <ref type="bibr" target="#b9">[10,</ref><ref type="bibr" target="#b14">15,</ref><ref type="bibr" target="#b17">18,</ref><ref type="bibr" target="#b18">19,</ref><ref type="bibr" target="#b19">20,</ref><ref type="bibr" target="#b20">21,</ref><ref type="bibr" target="#b21">22]</ref>. Studies like Yilmaz et al. <ref type="bibr" target="#b17">[18]</ref> offer various metrics to reflect user satisfaction, considering factors such as effort. Some studies further break down user satisfaction into query-level satisfactions <ref type="bibr" target="#b18">[19,</ref><ref type="bibr" target="#b19">20]</ref>, although others, such as Järvelin et al. <ref type="bibr" target="#b20">[21]</ref>, argue that summing query-level satisfaction misses the contextual journey where user queries are interconnected. Unlike traditional search systems, conversational search systems allow users to ask follow-up questions to refine their answers <ref type="bibr" target="#b21">[22]</ref>.</p><p>Despite widespread adoption, measuring user satisfaction remains an open question. Many studies rely on real-time user participation to gather feedback, providing fresh and realistic insights but requiring significant resources and participant incentives <ref type="bibr" target="#b9">[10,</ref><ref type="bibr" target="#b14">15]</ref>. Alternatively, satisfaction prediction proxies using deep learning models offer a computational approach, overcoming temporal and spatial constraints but demanding high-quality computational resources and datasets <ref type="bibr" target="#b7">[8]</ref>.</p><p>Emerging trends in user simulation offer promising solutions to these challenges <ref type="bibr" target="#b4">[5,</ref><ref type="bibr" target="#b1">2]</ref>, though prior studies are not yet comprehensive. Gao et al. <ref type="bibr" target="#b22">[23]</ref> noted that earlier simulators relied on randomly generated scenarios to approximate users' states of mind. While such models prove beneficial in specific domains like recommender systems <ref type="bibr" target="#b23">[24]</ref>, issues of explainability and scrutability remain unresolved. Our proposed framework addresses these issues by incorporating a user-profile-based personalized simulation approach, facilitating easier alignment with actual user behaviours and providing a scrutable means to control the simulation by modifying textual user profiles.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.">User Simulation</head><p>Azzopardi et al. <ref type="bibr" target="#b24">[25]</ref> define simulation as the imitation of the operation of real-world phenomena. Simulations enable detailed experimental design and control tailored to specific research questions. These high-level controls allow for experiments with user simulations to be conducted with several advantages <ref type="bibr" target="#b24">[25]</ref>.</p><p>Firstly, what-if experiments can be performed by setting up different scenarios <ref type="bibr" target="#b25">[26,</ref><ref type="bibr" target="#b26">27]</ref>. Secondly, user simulators ensure the repeatability of experimental results. Additionally, user simulations can achieve these benefits at a low cost <ref type="bibr" target="#b27">[28]</ref>.</p><p>In the IR community, user simulation methods are primarily divided into cognitive and statistical approaches <ref type="bibr" target="#b28">[29]</ref>. Cognitive approaches were among the first used in this field. Belkin <ref type="bibr" target="#b29">[30]</ref> described users, information resources, and IR models, characterizing users by their objectives, problems, and knowledge. Subsequent studies expanded on this foundation <ref type="bibr" target="#b30">[31,</ref><ref type="bibr" target="#b31">32,</ref><ref type="bibr" target="#b32">33]</ref>.</p><p>In contrast, statistical approaches focus on analyzing user behaviours and satisfaction <ref type="bibr" target="#b28">[29,</ref><ref type="bibr" target="#b33">34,</ref><ref type="bibr" target="#b34">35,</ref><ref type="bibr" target="#b35">36,</ref><ref type="bibr" target="#b36">37]</ref>. These approaches underpin early user simulators based on statistical models <ref type="bibr" target="#b37">[38,</ref><ref type="bibr" target="#b38">39]</ref>. Although these simulators heavily relied on corpora, they faced limitations such as the diversity of user intent <ref type="bibr" target="#b28">[29]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Table 1</head><p>Performance of models trained with priming factors in predicting users' actions, while AUC refers to the area under P-R curves. Adopted form Fu and Lipani <ref type="bibr" target="#b45">[46]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Action</head><p>Dataset Precision Recall F1 AUC All 0.844 0.574 0.684 0.835 Topi <ref type="bibr" target="#b46">[47]</ref> 0.977 0.951 0.963 0.995 TREC 1  1.000 0.625 0.769 0.800 Cran <ref type="bibr" target="#b27">[28]</ref> 0.795 0.875 0.833 0.883 QReCC <ref type="bibr" target="#b47">[48]</ref> 0.953 0.997 0.974 0.990 ORC <ref type="bibr" target="#b48">[49]</ref> 0 Agenda-based user simulations are popular due to their realistic responses and straightforward dialogue strategies <ref type="bibr" target="#b39">[40,</ref><ref type="bibr" target="#b40">41,</ref><ref type="bibr" target="#b41">42]</ref>. The latest trend involves employing deep learning models, including adversarial generative approaches <ref type="bibr" target="#b42">[43]</ref>, reinforcement learning <ref type="bibr" target="#b43">[44]</ref>, and inverse reinforcement learning to abstract knowledge from data <ref type="bibr" target="#b44">[45]</ref>.</p><p>Evaluating CIR with user simulations is becoming a key trend, enabling efficient and cost-effective evaluation at various levels of CIR <ref type="bibr" target="#b1">[2]</ref>. Our proposed prototype integrates benefits from the aforementioned approaches. Initially, it simulates user behaviour guided by statistical signals derived from the context, enhancing explainability. Subsequently, the prototype employs deep learning models, utilizing textual user profiles to produce realistic and diverse responses within controlled parameters, thereby enabling further exploration of the target system. The details of the user simulation prototype are in the next section.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Modelling and Simulating Users</head><p>In this section, we introduce an two-stage prototype for constructing a robust user simulation for CIR. The main structure of the prototype is divided into two parts to control simulated users: the Action Predictor and the Response Generator, as illustrated in Fig 2.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Modeling Actions</head><p>The Action Predictor controls the simulated users' actions during the current conversation. This concept originates from the psychology community's notion of priming, which refers to the unconscious influence of past experiences on current performance or behaviour <ref type="bibr" target="#b49">[50,</ref><ref type="bibr" target="#b50">51]</ref>. Many studies have established models to explain this mechanism <ref type="bibr" target="#b51">[52,</ref><ref type="bibr" target="#b49">50]</ref>. For instance, Tulving et al. <ref type="bibr" target="#b52">[53]</ref> conducted an experiment where participants viewed a list of 96 words and later completed graphemic word fragments both one hour and seven days after studying the list. The results demonstrated a significant influence of the word list on subsequent tests.</p><p>As shown in Table <ref type="table">1</ref>, results from Fu and Lipani <ref type="bibr" target="#b45">[46]</ref> demonstrate the potential of predicting users' next actions. The benefits of using lexical and textual patterns to predict users' next actions include low cost and ease of interpretation, which are valuable for evaluation.</p><p>In the current prototype, three user actions are modelled based on the dataset available. Stopping is defined as the action where users opt to end a conversation, typically indicating the conclusion of the exchange. This action is critical for evaluating effects such as the principle of least effort <ref type="bibr" target="#b53">[54]</ref>. Sessions of conversation often include Following up on queries that build upon previous interactions, acknowledging missing contexts and references to earlier discussed topics <ref type="bibr" target="#b48">[49]</ref>. As noted by Stede and Schlangen <ref type="bibr" target="#b54">[55]</ref>, an inquisitive user engaged in an ongoing dialogue may express interest in further related subjects as a response to the information provided. Switching topics is commonly seen in information-seeking dialogues, especially when using search systems for data acquisition <ref type="bibr" target="#b55">[56]</ref>.</p><p>As output from the Action Predictor, a general action as described above will be predicted, with the details elaborated in the Response Generator.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Personalized Responses</head><p>The Response Generator will generate realistic and diversified responses based on the conversation, user profiles and the predicted actions from the Action Predictor LLMs such as GPTs and Llama, renowned for their sophisticated natural language processing capabilities, present unique opportunities to enhance Conversational Systems through mechanisms like pre-training, fine-tuning, and prompting <ref type="bibr" target="#b56">[57]</ref>. The ability of LLMs to mimic diverse demographic characteristics offers a novel approach to simulating user behaviour and preferences <ref type="bibr" target="#b57">[58]</ref>. High-quality user simulations, which closely mirror real user behaviour distributions, can significantly advance CRS development, currently dependent on real data for training, with its inherent constraints and disadvantages.</p><p>Ramos et al. <ref type="bibr" target="#b58">[59]</ref> offers a valuable method for generating user profiles from the Amazon dataset, introducing a more compact style of personalization into the user simulator. Table <ref type="table">2</ref> demonstrates the agreement between simulated and real users' responses to the same items in the Amazon dataset. Since the simulated users are based on LLMs and prompts with user profiles, the results suggest the potential of generating personalized responses based on textual user profiles.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Table 2</head><p>Agreements (Cohen's 𝐾 𝑐 , Randolph's 𝐾 𝑟 and Krippendorff's 𝛼) between simulated users and actual users in response.Cohen's index considers the marginal distribution of categories, Randolph's assumes a uniform distribution, and Krippendorff's offers a broader approach to assessing agreement. Fair agreements for each metric (&gt;0.2) are marked as bold. Three simulation methods are evaluated: Real UP, where users are simulated based on their profiles; Rand UP, involving users simulated with random profiles; and Rand Sc, where scores are randomly generated based on the dataset's historical distribution.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Setting</head><p>𝐾 𝑐 𝐾 𝑟 𝛼 Real UP 0.20 0.34 0.20 Rand UP 0.14 0.26 0.14 Rand Sc -0.03 0.14 -0.03</p><p>In this simulation, responses are tailored based on user profiles, which consist of concise text that summarizes the attributes of users in a few succinct sentences. This can include motivations for task-oriented CIR systems. Furthermore, the Response Generator is tasked with handling clarification questions posed by the CIR system. At the conclusion of this phase, the user simulator is equipped to interact with CIR systems. The subsequent section proposes a linkage as the remaining component of this framework, specifically addressing the evaluation of the CIR system using this user simulator, given the absence of a direct indicator from the user simulation on the quality of the target CIR system.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Evaluation of CIR</head><p>This section discusses potential methods for applying evaluation tasks based on the user simulation prototype established in Section 3.</p><p>As introduced in Section 2, the target of evaluation originates from modelling users and conversations. According to Yilmaz et al. <ref type="bibr" target="#b17">[18]</ref>, user satisfaction can be reflected by the effort exerted. Effort also forms the foundation of modelling user actions, particularly stopping behaviours.</p><p>In the IR community, several studies have been conducted to depict stopping behaviours <ref type="bibr" target="#b59">[60,</ref><ref type="bibr" target="#b60">61,</ref><ref type="bibr" target="#b61">62,</ref><ref type="bibr" target="#b62">63]</ref>. These studies aim to quantify the feeling of having "enough. " For instance, users may decide to stop a conversation when they feel frustrated or satisfied.</p><p>A previous study by Fu and Lipani <ref type="bibr" target="#b45">[46]</ref> provided a reliable method for predicting stopping behaviours. The subsequent step is to explore the relationship between each stopping point and user satisfaction. The final part of the evaluation focuses on assessing the user simulation. To effectively perform evaluation for CIR, the user simulation must exhibit not only a diversity of reasonable responses but also a high alignment with real users, particularly in reflecting satisfaction or frustration.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.">Evaluating User Simulation</head><p>Following the study by Fu et al. <ref type="bibr" target="#b27">[28]</ref>, where direct and indirect assessments in CIR show substantial agreement, the alignment between user simulation and real users can be evaluated. Real users will review the conversations between the user simulation and the system to determine if the simulated user appears satisfied.</p><p>Figure <ref type="figure" target="#fig_2">3</ref> illustrates the operational flow within the evaluation component of the framework. In this component, a classifier, integrated with the user simulator, is utilized to predict user satisfaction. This classifier is aligned with annotations from real users. Feedback from these real annotators is employed to train both the user simulator and the classifier. The user simulator aims to accurately mimic real user behaviour through the Actions Predictor and the Response Generator. Similarly, the classifier is trained to align its judgments with those of real users regarding the same conversation.</p><p>During the development of this framework, numerous emerging trends were observed, particularly in LLMs. These observations are presented in the following section.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Towards Future</head><p>This represents a significant transformation since 2021, as the term Large Language Models (LLMs) has gained popularity. This shift has introduced both opportunities and challenges for contemporary research.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.1.">CIR and RAG</head><p>The integration of LLMs is not limited to the IR community; the ML community is also embracing IR techniques. A prominent development in this space is Retrieval-Augmented Generation (RAG), which enhances LLMs in domain-specific or knowledge-intensive tasks <ref type="bibr" target="#b63">[64]</ref>.</p><p>RAG involves multiple retrieval processes to enrich the context, going beyond traditional single retrieval methods. For instance, Self-RAG <ref type="bibr" target="#b64">[65]</ref> refines the RAG framework by enabling LLMs to actively determine the optimal moments and content for retrieval, thus improving the efficiency and relevance of the sourced information. The key aspect here is not merely multiple retrievals, but the reliance on the judgment of LLMs, indicating that LLMs can further participate in the processing with minimal human intervention.</p><p>For the evaluation of RAG, integrating typical LLMs alone may not suffice. A realistic inquiry is how agents based on LLMs can assess the responses from RAG systems that incorporate retrieved documents. One feasible approach is using RAG to evaluate itself. Here, the crucial aspects include not only the quality of the conversation and the retrieval process but also how effectively the documents are presented.</p><p>Moreover, in the specialized domain of CIR, where the systems are relatively light, LLMs can still serve as experts. The following section provides an example.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2.">Should We Ask LLMs First?</head><p>The TREC Interactive Knowledge Assistance Track (iKAT) builds upon the foundational work of the TREC Conversational Assistance Track (CAsT) <ref type="bibr" target="#b65">[66]</ref>, with a key difference being the addition of personal context for each user in the dataset. The primary task remains similar to TREC CAsT-retrieving and ranking documents from the corpus at each turn of the given conversations.</p><p>The best performance in iKAT 2023 introduced a novel approach <ref type="bibr" target="#b66">[67]</ref>. In this approach, the LLM generates an initial answer to the user's query based on the context of the conversation and the user profile. This answer is derived through reasoning over the context and the user's profile, but it is not grounded in the documents within the collection. Subsequently, the LLM generates a set of five queries to achieve this answer.</p><p>The controversial aspect of this approach is using the LLM-generated answer as the target without initial retrieval, followed by employing an IR system to achieve it. This process assumes that LLMs' answers are sufficiently accurate. Alternatively, it suggests that the documents have likely been exposed to the LLMs, raising concerns of potential data leakage.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.3.">How Can We Go Beyond Our Knowledge Borders?</head><p>Training LLMs from scratch is a challenging task for most research groups due to high costs and the lack of storage and computational resources. The most common practice involves fine-tuning a public base version of LLMs and accessing them via APIs. Top LLMs are trained on a substantial portion of internet text documents, making it nearly impossible to prevent data leakage once any public data is used in a study.</p><p>Furthermore, the widespread use of LLMs will inevitably introduce LLM-generated text back into the internet, posing a significant challenge that has already raised considerable concerns within the community.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.">Conclusion</head><p>This paper proposes a framework for evaluating CIR systems automatically using a user simulator prototype. The framework comprises two main components:</p><p>1. A prototype of user simulation that leverages the advancements from both psychology and ML fields to conduct realistic and scrutable simulations targeted at CIR systems. 2. A component that employs sophisticated conversation modelling concepts from the IR community to provide reasonable feedback aimed at predicting user satisfaction alongside the user simulation prototype.</p><p>In addition to the framework, this study also presents emerging trends observed with the promising development of LLMs. The advent of systems such as RAG introduces both opportunities and challenges. It raises concerns that the current practices within the community utilizing LLMs may lead to increased data leakages.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 :</head><label>1</label><figDesc>Figure 1: Comparative Example: Web Search (Left) vs. CIR Dialogue (Right)</figDesc><graphic coords="1,115.88,453.70,361.03,152.03" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>Figure 2 :</head><label>2</label><figDesc>Figure 2: The basic structure of the user simulation.</figDesc><graphic coords="4,115.88,480.05,361.04,169.26" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head>Figure 3 :</head><label>3</label><figDesc>Figure 3: The basic flow of the evaluation with the user simulator.</figDesc><graphic coords="6,115.88,389.41,361.03,230.70" type="bitmap" /></figure>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">Recent advances in conversational information retrieval</title>
		<author>
			<persName><forename type="first">J</forename><surname>Gao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Xiong</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Bennett</surname></persName>
		</author>
		<idno type="DOI">10.1145/3397271.3401418</idno>
		<idno>doi:10.1145/3397271.3401418</idno>
		<ptr target="https://doi.org/10.1145/3397271.3401418" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR &apos;20</title>
				<meeting>the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR &apos;20<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2020">2020</date>
			<biblScope unit="page" from="2421" to="2424" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<monogr>
		<title level="m" type="main">Neural approaches to conversational information retrieval</title>
		<author>
			<persName><forename type="first">J</forename><surname>Gao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Xiong</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Bennett</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Craswell</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2023">2023</date>
			<publisher>Springer Nature</publisher>
			<biblScope unit="volume">44</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<monogr>
		<author>
			<persName><forename type="first">N</forename><surname>Dziri</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Kamalloo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Milton</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Zaiane</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Yu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Ponti</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Reddy</surname></persName>
		</author>
		<idno type="arXiv">arXiv:2204.10757</idno>
		<ptr target="https://arxiv.org/abs/2204.10757" />
		<title level="m">Faithdial: A faithful benchmark for information-seeking dialogue</title>
				<imprint>
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
	<note type="report_type">arXiv preprint</note>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Challenges in the evaluation of conversational search systems</title>
		<author>
			<persName><forename type="first">G</forename><surname>Penha</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Hauff</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Converse@ KDD</title>
				<imprint>
			<date type="published" when="2020">2020</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">Conversational information seeking</title>
		<author>
			<persName><forename type="first">H</forename><surname>Zamani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">R</forename><surname>Trippas</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Dalton</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Radlinski</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Foundations and Trends® in Information Retrieval</title>
		<imprint>
			<biblScope unit="volume">17</biblScope>
			<biblScope unit="page" from="244" to="456" />
			<date type="published" when="2023">2023</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">How am i doing?: Evaluating conversational search systems offline</title>
		<author>
			<persName><forename type="first">A</forename><surname>Lipani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Carterette</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Yilmaz</surname></persName>
		</author>
		<idno type="DOI">10.1145/3451160</idno>
	</analytic>
	<monogr>
		<title level="j">ACM Trans. Inf. Syst</title>
		<imprint>
			<biblScope unit="volume">39</biblScope>
			<date type="published" when="2021">2021</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Automatic online evaluation of intelligent assistants</title>
		<author>
			<persName><forename type="first">J</forename><surname>Jiang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Hassan Awadallah</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Jones</surname></persName>
		</author>
		<author>
			<persName><forename type="first">U</forename><surname>Ozertem</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Zitouni</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Kulkarni</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><forename type="middle">Z</forename><surname>Khan</surname></persName>
		</author>
		<idno type="DOI">10.1145/2736277.2741669</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 24th International Conference on World Wide Web, WWW &apos;15, International World Wide Web Conferences Steering Committee</title>
				<meeting>the 24th International Conference on World Wide Web, WWW &apos;15, International World Wide Web Conferences Steering Committee<address><addrLine>Canton of Geneva, CHE</addrLine></address></meeting>
		<imprint>
			<publisher>Republic</publisher>
			<date type="published" when="2015">2015</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">Offline and online satisfaction prediction in open-domain conversational systems</title>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">I</forename><surname>Choi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Ahmadvand</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Agichtein</surname></persName>
		</author>
		<idno type="DOI">10.1145/3357384.3358047</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 28th ACM International Conference on Information and Knowledge Management, CIKM &apos;19</title>
				<meeting>the 28th ACM International Conference on Information and Knowledge Management, CIKM &apos;19<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2019">2019</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<analytic>
		<author>
			<persName><forename type="first">A</forename><surname>Anand</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Cavedon</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Joho</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Sanderson</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Stein</surname></persName>
		</author>
		<idno type="DOI">10.4230/DagRep.9.11.34</idno>
	</analytic>
	<monogr>
		<title level="m">Conversational Search (Dagstuhl Seminar 19461)</title>
				<imprint>
			<date type="published" when="2020">2020</date>
			<biblScope unit="volume">9</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">Predicting user satisfaction with intelligent assistants</title>
		<author>
			<persName><forename type="first">J</forename><surname>Kiseleva</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Williams</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Hassan Awadallah</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><forename type="middle">C</forename><surname>Crook</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Zitouni</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Anastasakos</surname></persName>
		</author>
		<idno type="DOI">10.1145/2911451.2911521</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR &apos;16</title>
				<meeting>the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR &apos;16<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2016">2016</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<monogr>
		<author>
			<persName><forename type="first">J</forename><surname>Dalton</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Xiong</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Callan</surname></persName>
		</author>
		<title level="m">Trec cast 2019: The conversational assistance track overview</title>
				<imprint>
			<date type="published" when="2020">2020</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Relevance reconsidered</title>
		<author>
			<persName><forename type="first">T</forename><surname>Saracevic</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the second conference on conceptions of library and information science (CoLIS 2)</title>
				<meeting>the second conference on conceptions of library and information science (CoLIS 2)</meeting>
		<imprint>
			<date type="published" when="1996">1996</date>
			<biblScope unit="page" from="201" to="218" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b12">
	<analytic>
		<title level="a" type="main">Bleu: A method for automatic evaluation of machine translation</title>
		<author>
			<persName><forename type="first">K</forename><surname>Papineni</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Roukos</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Ward</surname></persName>
		</author>
		<author>
			<persName><forename type="first">W.-J</forename><surname>Zhu</surname></persName>
		</author>
		<idno type="DOI">10.3115/1073083.1073135</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL &apos;02, Association for Computational Linguistics</title>
				<meeting>the 40th Annual Meeting on Association for Computational Linguistics, ACL &apos;02, Association for Computational Linguistics<address><addrLine>USA</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2002">2002</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b13">
	<analytic>
		<title level="a" type="main">A structured review of the validity of BLEU</title>
		<author>
			<persName><forename type="first">E</forename><surname>Reiter</surname></persName>
		</author>
		<idno type="DOI">10.1162/coli_a_00322</idno>
	</analytic>
	<monogr>
		<title level="j">Computational Linguistics</title>
		<imprint>
			<biblScope unit="volume">44</biblScope>
			<date type="published" when="2018">2018</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b14">
	<analytic>
		<title level="a" type="main">Understanding user satisfaction with intelligent assistants</title>
		<author>
			<persName><forename type="first">J</forename><surname>Kiseleva</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Williams</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Jiang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Hassan Awadallah</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><forename type="middle">C</forename><surname>Crook</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Zitouni</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Anastasakos</surname></persName>
		</author>
		<idno type="DOI">10.1145/2854946.2854961</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2016 ACM on Conference on Human Information Interaction and Retrieval, CHIIR &apos;16</title>
				<meeting>the 2016 ACM on Conference on Human Information Interaction and Retrieval, CHIIR &apos;16<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2016">2016</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b15">
	<analytic>
		<title level="a" type="main">How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation</title>
		<author>
			<persName><forename type="first">C.-W</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Lowe</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Serban</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Noseworthy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Charlin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Pineau</surname></persName>
		</author>
		<idno type="DOI">10.18653/v1/D16-1230</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics</title>
				<meeting>the 2016 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics<address><addrLine>Austin, Texas</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2016">2016</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b16">
	<analytic>
		<title level="a" type="main">Methods for evaluating interactive information retrieval systems with users</title>
		<author>
			<persName><forename type="first">D</forename><surname>Kelly</surname></persName>
		</author>
		<idno type="DOI">10.1561/1500000012</idno>
	</analytic>
	<monogr>
		<title level="j">Foundations and Trends® in Information Retrieval</title>
		<imprint>
			<biblScope unit="volume">3</biblScope>
			<date type="published" when="2009">2009</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b17">
	<analytic>
		<title level="a" type="main">Relevance and effort: An analysis of document utility</title>
		<author>
			<persName><forename type="first">E</forename><surname>Yilmaz</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Verma</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Craswell</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Radlinski</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Bailey</surname></persName>
		</author>
		<idno type="DOI">10.1145/2661829.2661953</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management, CIKM &apos;14</title>
				<meeting>the 23rd ACM International Conference on Conference on Information and Knowledge Management, CIKM &apos;14<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2014">2014</date>
			<biblScope unit="page" from="91" to="100" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b18">
	<analytic>
		<title level="a" type="main">Modelling and detecting changes in user satisfaction</title>
		<author>
			<persName><forename type="first">J</forename><surname>Kiseleva</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Crestan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Brigo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Dittel</surname></persName>
		</author>
		<idno type="DOI">10.1145/2661829.2661960</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management, CIKM &apos;14</title>
				<meeting>the 23rd ACM International Conference on Conference on Information and Knowledge Management, CIKM &apos;14<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2014">2014</date>
			<biblScope unit="page" from="1449" to="1458" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b19">
	<analytic>
		<title level="a" type="main">Behavioral dynamics from the serp&apos;s perspective: What are failed serps and how to fix them?</title>
		<author>
			<persName><forename type="first">J</forename><surname>Kiseleva</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Kamps</surname></persName>
		</author>
		<author>
			<persName><forename type="first">V</forename><surname>Nikulin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Makarov</surname></persName>
		</author>
		<idno type="DOI">10.1145/2806416.2806483</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 24th ACM International on Conference on Information and Knowledge Management, CIKM &apos;15</title>
				<meeting>the 24th ACM International on Conference on Information and Knowledge Management, CIKM &apos;15<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2015">2015</date>
			<biblScope unit="page" from="1561" to="1570" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b20">
	<analytic>
		<title level="a" type="main">Discounted cumulated gain based evaluation of multiple-query ir sessions</title>
		<author>
			<persName><forename type="first">K</forename><surname>Järvelin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">L</forename><surname>Price</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><forename type="middle">M L</forename><surname>Delcambre</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">L</forename><surname>Nielsen</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Advances in Information Retrieval</title>
				<editor>
			<persName><forename type="first">C</forename><surname>Macdonald</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">I</forename><surname>Ounis</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">V</forename><surname>Plachouras</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">I</forename><surname>Ruthven</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">R</forename><forename type="middle">W</forename><surname>White</surname></persName>
		</editor>
		<meeting><address><addrLine>Berlin Heidelberg; Berlin, Heidelberg</addrLine></address></meeting>
		<imprint>
			<publisher>Springer</publisher>
			<date type="published" when="2008">2008</date>
			<biblScope unit="page" from="4" to="15" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b21">
	<analytic>
		<title level="a" type="main">The relationship between ir effectiveness measures and user satisfaction</title>
		<author>
			<persName><forename type="first">A</forename><surname>Al-Maskari</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Sanderson</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Clough</surname></persName>
		</author>
		<idno type="DOI">10.1145/1277741.1277902</idno>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR &apos;07</title>
				<meeting>the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR &apos;07<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2007">2007</date>
			<biblScope unit="page" from="773" to="774" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b22">
	<analytic>
		<title level="a" type="main">Neural approaches to conversational ai</title>
		<author>
			<persName><forename type="first">J</forename><surname>Gao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Galley</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Li</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">The 41st international ACM SIGIR conference on research &amp; development in information retrieval</title>
				<imprint>
			<date type="published" when="2018">2018</date>
			<biblScope unit="page" from="1371" to="1374" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b23">
	<analytic>
		<title level="a" type="main">Evaluating conversational recommender systems via user simulation</title>
		<author>
			<persName><forename type="first">S</forename><surname>Zhang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Balog</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 26th acm sigkdd international conference on knowledge discovery &amp; data mining</title>
				<meeting>the 26th acm sigkdd international conference on knowledge discovery &amp; data mining</meeting>
		<imprint>
			<date type="published" when="2020">2020</date>
			<biblScope unit="page" from="1512" to="1520" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b24">
	<analytic>
		<title level="a" type="main">Report on the sigir 2010 workshop on the simulation of interaction</title>
		<author>
			<persName><forename type="first">L</forename><surname>Azzopardi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Järvelin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Kamps</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">D</forename><surname>Smucker</surname></persName>
		</author>
		<idno type="DOI">10.1145/1924475.1924484</idno>
		<idno>doi:10.1145/1924475.1924484</idno>
		<ptr target="https://doi.org/10.1145/1924475.1924484" />
	</analytic>
	<monogr>
		<title level="j">SIGIR Forum</title>
		<imprint>
			<biblScope unit="volume">44</biblScope>
			<biblScope unit="page" from="35" to="47" />
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b25">
	<analytic>
		<title level="a" type="main">Software process simulation modeling: why? what? how?</title>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">I</forename><surname>Kellner</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">J</forename><surname>Madachy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">M</forename><surname>Raffo</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of Systems and Software</title>
		<imprint>
			<biblScope unit="volume">46</biblScope>
			<biblScope unit="page" from="91" to="105" />
			<date type="published" when="1999">1999</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b26">
	<analytic>
		<title level="a" type="main">Report on the 1st simulation for information retrieval workshop (sim4ir</title>
		<author>
			<persName><forename type="first">K</forename><surname>Balog</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Maxwell</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Thomas</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Zhang</surname></persName>
		</author>
		<idno type="DOI">10.1145/3527546.3527559</idno>
		<idno>doi:10.1145/3527546.3527559</idno>
		<ptr target="https://doi.org/10.1145/3527546.3527559" />
	</analytic>
	<monogr>
		<title level="m">sigir 2021</title>
				<imprint>
			<date type="published" when="2021">2021. 2022</date>
			<biblScope unit="volume">55</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b27">
	<analytic>
		<title level="a" type="main">Evaluating the cranfield paradigm for conversational search systems</title>
		<author>
			<persName><forename type="first">X</forename><surname>Fu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Yilmaz</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Lipani</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2022 ACM SIGIR International Conference on Theory of Information Retrieval</title>
				<meeting>the 2022 ACM SIGIR International Conference on Theory of Information Retrieval</meeting>
		<imprint>
			<date type="published" when="2022">2022</date>
			<biblScope unit="page" from="275" to="280" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b28">
	<monogr>
		<author>
			<persName><forename type="first">P</forename><surname>Erbacher</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Soulier</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Denoyer</surname></persName>
		</author>
		<idno type="arXiv">arXiv:2201.03435</idno>
		<title level="m">State of the art of user simulation approaches for conversational information retrieval</title>
				<imprint>
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
	<note type="report_type">arXiv preprint</note>
</biblStruct>

<biblStruct xml:id="b29">
	<analytic>
		<title level="a" type="main">Cognitive models and information transfer</title>
		<author>
			<persName><forename type="first">N</forename><surname>Belkin</surname></persName>
		</author>
		<idno type="DOI">10.1016/0143-6236(84)90070-X</idno>
		<ptr target="https://doi.org/10.1016/0143-6236(84)90070-X" />
	</analytic>
	<monogr>
		<title level="j">Social Science Information Studies</title>
		<imprint>
			<biblScope unit="volume">4</biblScope>
			<biblScope unit="page" from="111" to="129" />
			<date type="published" when="1984">1984</date>
		</imprint>
	</monogr>
	<note>special Issue Seminar on the Psychological Aspects of Information Searching</note>
</biblStruct>

<biblStruct xml:id="b30">
	<analytic>
		<title level="a" type="main">Inside the search process: Information seeking from the user&apos;s perspective</title>
		<author>
			<persName><forename type="first">C</forename><forename type="middle">C</forename><surname>Kuhlthau</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">J. Am. Soc. Inf. Sci</title>
		<imprint>
			<biblScope unit="volume">42</biblScope>
			<biblScope unit="page" from="361" to="371" />
			<date type="published" when="1991">1991</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b31">
	<monogr>
		<title level="m" type="main">The Turn: Integration of Information Seeking and Retrieval in Context</title>
		<author>
			<persName><forename type="first">P</forename><surname>Ingwersen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Järvelin</surname></persName>
		</author>
		<idno type="DOI">10.1007/1-4020-3851-8</idno>
		<imprint>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b32">
	<analytic>
		<title level="a" type="main">A behavioral approach to information retrieval system design</title>
		<author>
			<persName><forename type="first">D</forename><surname>Ellis</surname></persName>
		</author>
		<idno type="DOI">10.1108/eb026843</idno>
		<ptr target="https://doi.org/10.1108/eb026843.doi:10.1108/eb026843" />
	</analytic>
	<monogr>
		<title level="j">J. Doc</title>
		<imprint>
			<biblScope unit="volume">45</biblScope>
			<biblScope unit="page" from="171" to="212" />
			<date type="published" when="1989">1989</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b33">
	<analytic>
		<title level="a" type="main">An experimental comparison of click position-bias models</title>
		<author>
			<persName><forename type="first">N</forename><surname>Craswell</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Zoeter</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Taylor</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Ramsey</surname></persName>
		</author>
		<idno type="DOI">10.1145/1341531.1341545</idno>
		<idno>doi:10.1145/1341531.1341545</idno>
		<ptr target="http://doi.acm.org/10.1145/1341531.1341545" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the international conference on Web search and web data mining, WSDM &apos;08</title>
				<meeting>the international conference on Web search and web data mining, WSDM &apos;08<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>ACM</publisher>
			<date type="published" when="2008">2008</date>
			<biblScope unit="page" from="87" to="94" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b34">
	<monogr>
		<title level="m" type="main">A dynamic bayesian network click model for web search ranking</title>
		<author>
			<persName><forename type="first">O</forename><surname>Chapelle</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Zhang</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2009">2009</date>
			<publisher>WWW</publisher>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b35">
	<analytic>
		<title level="a" type="main">A user browsing model to predict search engine click data from past observations</title>
		<author>
			<persName><forename type="first">G</forename><forename type="middle">E</forename><surname>Dupret</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Piwowarski</surname></persName>
		</author>
		<idno type="DOI">10.1145/1390334.1390392</idno>
		<ptr target="http://portal.acm.org/citation.cfm?id=1390334.1390392.doi:10.1145/1390334.1390392" />
	</analytic>
	<monogr>
		<title level="m">SIGIR &apos;08: Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval</title>
				<meeting><address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>ACM</publisher>
			<date type="published" when="2008">2008</date>
			<biblScope unit="page" from="331" to="338" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b36">
	<monogr>
		<title level="m" type="main">Modeling clicks beyond the first result page</title>
		<author>
			<persName><forename type="first">A</forename><surname>Chuklin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Serdyukov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Rijke</surname></persName>
		</author>
		<idno type="DOI">10.1145/2505515.2507859</idno>
		<imprint>
			<date type="published" when="2013">2013</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b37">
	<analytic>
		<title level="a" type="main">User modeling for spoken dialogue system evaluation</title>
		<author>
			<persName><forename type="first">W</forename><surname>Eckert</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Levin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Pieraccini</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings</title>
				<imprint>
			<date type="published" when="1997">1997. 1997</date>
			<biblScope unit="page" from="80" to="87" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b38">
	<analytic>
		<title level="a" type="main">Probabilistic simulation of human-machine dialogues</title>
		<author>
			<persName><forename type="first">K</forename><surname>Scheffler</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">J</forename><surname>Young</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100</title>
				<imprint>
			<date type="published" when="2000">2000. 2000</date>
			<biblScope unit="volume">2</biblScope>
			<biblScope unit="page" from="I1217" to="I1220" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b39">
	<analytic>
		<title level="a" type="main">User modeling in spoken dialogue systems to generate flexible guidance</title>
		<author>
			<persName><forename type="first">K</forename><surname>Komatani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Ueno</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Kawahara</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Okuno</surname></persName>
		</author>
		<idno type="DOI">10.1007/s11257-004-5659-0</idno>
	</analytic>
	<monogr>
		<title level="j">User Modeling and User-Adapted Interaction</title>
		<imprint>
			<biblScope unit="volume">15</biblScope>
			<biblScope unit="page" from="169" to="183" />
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b40">
	<monogr>
		<title level="m" type="main">Agenda-based user simulation for bootstrapping a pomdp dialogue system</title>
		<author>
			<persName><forename type="first">J</forename><surname>Schatzmann</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Thomson</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Weilhammer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Ye</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Young</surname></persName>
		</author>
		<idno type="DOI">10.3115/1614108.1614146</idno>
		<imprint>
			<date type="published" when="2007">2007</date>
			<biblScope unit="page" from="149" to="152" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b41">
	<monogr>
		<title level="m" type="main">A user simulator for task-completion dialogues</title>
		<author>
			<persName><forename type="first">X</forename><surname>Li</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><forename type="middle">C</forename><surname>Lipton</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Dhingra</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Li</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Gao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y.-N</forename><surname>Chen</surname></persName>
		</author>
		<idno>CoRR abs/1612.05688</idno>
		<ptr target="http://dblp.uni-trier.de/db/journals/corr/corr1612.html#LiLDLGC16" />
		<imprint>
			<date type="published" when="2016">2016</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b42">
	<monogr>
		<title level="m" type="main">Toward simulating environments in reinforcement learning based recommendations</title>
		<author>
			<persName><forename type="first">X</forename><surname>Zhao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Xia</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><surname>Ding</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Yin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Tang</surname></persName>
		</author>
		<idno>ArXiv abs/1906.11462</idno>
		<imprint>
			<date type="published" when="2019">2019</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b43">
	<analytic>
		<title level="a" type="main">Generative adversarial user model for reinforcement learning based recommendation system</title>
		<author>
			<persName><forename type="first">X</forename><surname>Chen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Li</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Li</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Jiang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Qi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Song</surname></persName>
		</author>
		<ptr target="https://proceedings.mlr.press/v97/chen19f.html" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 36th International Conference on Machine Learning</title>
				<editor>
			<persName><forename type="first">K</forename><surname>Chaudhuri</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">R</forename><surname>Salakhutdinov</surname></persName>
		</editor>
		<meeting>the 36th International Conference on Machine Learning<address><addrLine>PMLR</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2019">2019</date>
			<biblScope unit="volume">97</biblScope>
			<biblScope unit="page" from="1052" to="1061" />
		</imprint>
	</monogr>
	<note>Proceedings of Machine Learning Research</note>
</biblStruct>

<biblStruct xml:id="b44">
	<analytic>
		<title level="a" type="main">User simulation in dialogue systems using inverse reinforcement learning</title>
		<author>
			<persName><forename type="first">S</forename><surname>Chandramohan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Geist</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Lefèvre</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Pietquin</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">INTERSPEECH</title>
				<imprint>
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b45">
	<monogr>
		<title level="m" type="main">Priming and actions: An analysis in conversational search systems</title>
		<author>
			<persName><forename type="first">X</forename><surname>Fu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Lipani</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2023">2023</date>
			<publisher>SIGIR</publisher>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b46">
	<monogr>
		<title level="m" type="main">Topi-OCQA: Open-domain conversational question answering with topic switching</title>
		<author>
			<persName><forename type="first">V</forename><surname>Adlakha</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Dhuliawala</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Suleman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Vries</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Reddy</surname></persName>
		</author>
		<idno type="DOI">10.1162/tacl_a_00471/2008126/tacl_a_00471.pdf</idno>
		<ptr target="https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl_a_00471/2008126/tacl_a_00471.pdf" />
		<imprint>
			<date type="published" when="2022">2022</date>
			<biblScope unit="volume">10</biblScope>
			<biblScope unit="page" from="468" to="483" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b47">
	<analytic>
		<title level="a" type="main">Open-domain question answering goes conversational via question rewriting</title>
		<author>
			<persName><forename type="first">R</forename><surname>Anantha</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Vakulenko</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><surname>Tu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Longpre</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Pulman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Chappidi</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</title>
				<meeting>the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</meeting>
		<imprint>
			<date type="published" when="2021">2021</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b48">
	<monogr>
		<title level="m" type="main">Open-Retrieval Conversational Question Answering</title>
		<author>
			<persName><forename type="first">C</forename><surname>Qu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Yang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Chen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Qiu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">W</forename><forename type="middle">B</forename><surname>Croft</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Iyyer</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2020">2020</date>
			<publisher>SIGIR</publisher>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b49">
	<analytic>
		<title level="a" type="main">Priming and the brain</title>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">L</forename><surname>Schacter</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">L</forename><surname>Buckner</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Neuron</title>
		<imprint>
			<biblScope unit="volume">20</biblScope>
			<biblScope unit="page" from="185" to="195" />
			<date type="published" when="1998">1998</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b50">
	<monogr>
		<title level="m" type="main">The mind in the middle: A practical guide to priming and automaticity research</title>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">A</forename><surname>Bargh</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><forename type="middle">L</forename><surname>Chartrand</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b51">
	<analytic>
		<title level="a" type="main">Content and process priming: A review</title>
		<author>
			<persName><forename type="first">C</forename><surname>Janiszewski</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">S</forename><surname>Wyer</surname><genName>Jr</genName></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of consumer psychology</title>
		<imprint>
			<biblScope unit="volume">24</biblScope>
			<biblScope unit="page" from="96" to="118" />
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b52">
	<analytic>
		<title level="a" type="main">Priming effects in word-fragment completion are independent of recognition memory</title>
		<author>
			<persName><forename type="first">E</forename><surname>Tulving</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">L</forename><surname>Schacter</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><forename type="middle">A</forename><surname>Stark</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of experimental psychology: learning, memory, and cognition</title>
		<imprint>
			<biblScope unit="volume">8</biblScope>
			<biblScope unit="page">336</biblScope>
			<date type="published" when="1982">1982</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b53">
	<monogr>
		<author>
			<persName><forename type="first">G</forename><forename type="middle">K</forename><surname>Zipf</surname></persName>
		</author>
		<title level="m">Human behavior and the principle of least effort: An introduction to human ecology</title>
				<imprint>
			<publisher>Ravenio Books</publisher>
			<date type="published" when="2016">2016</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b54">
	<monogr>
		<title level="m" type="main">Information-seeking chat : Dialogue management by topic structure</title>
		<author>
			<persName><forename type="first">M</forename><surname>Stede</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Schlangen</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b55">
	<analytic>
		<title level="a" type="main">Multitasking information seeking and searching processes</title>
		<author>
			<persName><forename type="first">A</forename><surname>Spink</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Özmutlu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Özmutlu</surname></persName>
		</author>
		<idno type="DOI">10.1002/asi.10124</idno>
	</analytic>
	<monogr>
		<title level="j">JASIST</title>
		<imprint>
			<biblScope unit="volume">53</biblScope>
			<biblScope unit="page" from="639" to="652" />
			<date type="published" when="2002">2002</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b56">
	<monogr>
		<author>
			<persName><forename type="first">W</forename><surname>Fan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><surname>Zhao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Li</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Mei</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Tang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Q</forename><surname>Li</surname></persName>
		</author>
		<idno type="arXiv">arXiv:2307.02046</idno>
		<title level="m">Recommender systems in the era of large language models (llms)</title>
				<imprint>
			<date type="published" when="2023">2023</date>
		</imprint>
	</monogr>
	<note type="report_type">arXiv preprint</note>
</biblStruct>

<biblStruct xml:id="b57">
	<analytic>
		<title level="a" type="main">Using large language models to simulate multiple humans and replicate human subject studies</title>
		<author>
			<persName><forename type="first">G</forename><forename type="middle">V</forename><surname>Aher</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">I</forename><surname>Arriaga</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><forename type="middle">T</forename><surname>Kalai</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">International Conference on Machine Learning</title>
				<meeting><address><addrLine>PMLR</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2023">2023</date>
			<biblScope unit="page" from="337" to="371" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b58">
	<monogr>
		<title level="m" type="main">Natural language user profiles for transparent and scrutable recommendations</title>
		<author>
			<persName><forename type="first">J</forename><surname>Ramos</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><forename type="middle">A</forename><surname>Rahmani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Fu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Lipani</surname></persName>
		</author>
		<idno type="arXiv">arXiv:2402.05810</idno>
		<imprint>
			<date type="published" when="2024">2024</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b59">
	<analytic>
		<title level="a" type="main">Searching and stopping: An analysis of stopping rules and strategies</title>
		<author>
			<persName><forename type="first">D</forename><surname>Maxwell</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Azzopardi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Järvelin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Keskustalo</surname></persName>
		</author>
		<idno type="DOI">10.1145/2806416.2806476</idno>
		<idno>doi:10.1145/ 2806416.2806476</idno>
		<ptr target="https://doi.org/10.1145/2806416.2806476" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 24th ACM International on Conference on Information and Knowledge Management, CIKM &apos;15</title>
				<meeting>the 24th ACM International on Conference on Information and Knowledge Management, CIKM &apos;15<address><addrLine>New York, NY, USA</addrLine></address></meeting>
		<imprint>
			<publisher>Association for Computing Machinery</publisher>
			<date type="published" when="2015">2015</date>
			<biblScope unit="page" from="313" to="322" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b60">
	<analytic>
		<title level="a" type="main">On selecting a measure of retrieval effectiveness part ii. implementation of the philosophy</title>
		<author>
			<persName><forename type="first">W</forename><forename type="middle">S</forename><surname>Cooper</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of the American Society for information Science</title>
		<imprint>
			<biblScope unit="volume">24</biblScope>
			<biblScope unit="page" from="413" to="424" />
			<date type="published" when="1973">1973</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b61">
	<analytic>
		<title level="a" type="main">Stopping rules and their effect on expected search length</title>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">H</forename><surname>Kraft</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Lee</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Information Processing &amp; Management</title>
		<imprint>
			<biblScope unit="volume">15</biblScope>
			<biblScope unit="page" from="47" to="58" />
			<date type="published" when="1979">1979</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b62">
	<monogr>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">R</forename><surname>Nickles</surname></persName>
		</author>
		<title level="m">Judgment-based and reasoning-based stopping rules in decision-making under uncertainty</title>
				<imprint>
			<date type="published" when="1995">1995</date>
		</imprint>
		<respStmt>
			<orgName>University of Minnesota</orgName>
		</respStmt>
	</monogr>
</biblStruct>

<biblStruct xml:id="b63">
	<monogr>
		<author>
			<persName><forename type="first">Y</forename><surname>Gao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Xiong</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Gao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Jia</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Pan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Bi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Dai</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Sun</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Wang</surname></persName>
		</author>
		<idno type="arXiv">arXiv:2312.10997</idno>
		<title level="m">Retrieval-augmented generation for large language models: A survey</title>
				<imprint>
			<date type="published" when="2023">2023</date>
		</imprint>
	</monogr>
	<note type="report_type">arXiv preprint</note>
</biblStruct>

<biblStruct xml:id="b64">
	<monogr>
		<author>
			<persName><forename type="first">A</forename><surname>Asai</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><surname>Wu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Sil</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Hajishirzi</surname></persName>
		</author>
		<idno type="arXiv">arXiv:2310.11511</idno>
		<title level="m">Self-rag: Learning to retrieve, generate, and critique through self-reflection</title>
				<imprint>
			<date type="published" when="2023">2023</date>
		</imprint>
	</monogr>
	<note type="report_type">arXiv preprint</note>
</biblStruct>

<biblStruct xml:id="b65">
	<monogr>
		<author>
			<persName><forename type="first">M</forename><surname>Aliannejadi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><surname>Abbasiantaeb</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Chatterjee</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Dalton</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Azzopardi</surname></persName>
		</author>
		<idno type="arXiv">arXiv:2401.01330</idno>
		<title level="m">Trec ikat 2023: The interactive knowledge assistance track overview</title>
				<imprint>
			<date type="published" when="2024">2024</date>
		</imprint>
	</monogr>
	<note type="report_type">arXiv preprint</note>
</biblStruct>

<biblStruct xml:id="b66">
	<monogr>
		<title level="m" type="main">Llm-based retrieval and generation pipelines for trec interactive knowledge assistance track (ikat)</title>
		<author>
			<persName><forename type="first">Z</forename><surname>Abbasiantaeb</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Meng</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Rau</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Krasakis</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><forename type="middle">A</forename><surname>Rahmani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Aliannejadi</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2023">2023. 2023</date>
			<publisher>TREC</publisher>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
