<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SEUPD@CLEF: Team QEVALS on Information Retrieval Adapted to the Temporal Evolution of Web Documents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Enrico D'Alberton</string-name>
          <email>enrico.dalberton@studenti.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saverio Fincato</string-name>
          <email>saverio.fincato@studenti.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vaidas Lenartavicius</string-name>
          <email>vaidas.lenartavicius@studenti.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Pallante</string-name>
          <email>P@10</email>
          <email>laura.pallante@studenti.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yijian Qiu</string-name>
          <email>yijian.qiu@studenti.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This report presents the work conducted by our team for LongEval-Retrieval Task 1 [1] in CLEF 2023 [2]. The primary objective of this task is to develop an information retrieval system that can efectively adapt to the temporal evolution of Web documents. Using the Longeval Websearch collection provided by the commercial search engine Qwant[3], our team has built a retrieval system that addresses the challenges posed by the changing nature of Web documents and user search preferences. This paper discusses our approach to the subtasks of short-term persistence and long-term persistence, as well as the evaluation of our retrieval system's performance.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;CLEF 2023</kwd>
        <kwd>LongEval-Retrieval</kwd>
        <kwd>Information Retrieval</kwd>
        <kwd>Temporal Evolution</kwd>
        <kwd>Search Engines</kwd>
        <kwd>Short-term Persistence</kwd>
        <kwd>Long-term Persistence</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Nowadays, the advent of the internet has led to an exponential growth of the information
available online and this brought a more challenging way to find relevant and accurate information.
This is the environment where search engines play a vital role that allows us to access the
correct information quickly and eficiently just with a few clicks. Information retrieval systems
are more crucial than ever, influencing several fields such as healthcare, business and mainly
education.</p>
      <p>Our objective is to investigate the subject of information retrieval in search engines and its
adaptability to changes over time. The necessity to create a retrieval system that can successfully
handle the dynamic nature of the web is what stimulates this research’s development.
Our team, QEVALS, is participating in this challenge as a student group project conducted in
the Search Engines course a.y. 2022/23 at the Computer Engineering master’s degree at the</p>
      <p>
        University of Padua. We’ll be working on the Longeval [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] Websearch collection, a sizable
collection of data made up of web pages, user interactions, and queries made available by Qwant,
a privacy-focused French search engine.
      </p>
      <p>This project helped develop our understanding of the information retrieval systems used in
the context of a web search engine. Specifically, we applied ourselves to the task of processing
queries and documents in order to get the best possible ranking results to provide users with
the most relevant information. The paper is organized as follows: Section 2 describes our
approach; Section 3 explains our experimental setup; Section 4 discusses our main findings;
ifnally, Section 5 draws some conclusions and outlooks for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <p>The development process was iterative and based heavily on discussions between the group
members regarding the subjects covered during the lectures. However, our first results were
significantly below the baseline set by the organizers of the LongEval task.</p>
      <sec id="sec-2-1">
        <title>2.1. Analyzer</title>
        <p>
          The initial implementations of the classes were very simple and we started adding features to
them day by day. The main purpose of an analyzer in an IR model is to pre-process the input
data in a way that reduces the complexity of the document representation, while still
retaining the relevant information necessary for accurate retrieval. The final models we submitted,
present a tokenizer implemented using LetterTokenizer class from Apache Lucene [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and
it simply divides text at non-letters. Subsequently, we applied a LowerCaseFilter and an
ElisionFilter, which, respectively, normalize token text to lowercase and remove elisions
from a token. This last filter has been used specifically for our case since the dataset provided
by LongEval is in French. We have also applied a ASCIIFoldingFilter to convert alphabetic,
numeric, and symbolic Unicode characters that are not in the first 127 ASCII characters (the
"Basic Latin" Unicode block) into their ASCII equivalents.
        </p>
        <p>One of the tools that had the biggest impact on our system’s performance was the
implementation of a StopFilter. As the name suggests, it is used to remove stopwords (very frequent
words that contain little information about the contents of a document or a query) from a stream
of text. While testing diferent stoplists and experimenting with query boosting some parts of
the queries, we discovered that for some stoplists, individually boosting query tokens that were
non-stopwords would improve system performance, even though the queries were subject to a
StopFilter later down the pipeline. We tried using the same stoplist for our query boosting
and our StopFilter, as well as diferent ones and found that the best combination varied on a
case-by-case basis. This particular type of query boosting was efective only on less-performing
stoplists, thus it isn’t included in the final system. These are the results compared:</p>
        <sec id="sec-2-1-1">
          <title>NoFilter</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>OwnFilter FST</title>
        </sec>
        <sec id="sec-2-1-3">
          <title>NoFilter</title>
        </sec>
        <sec id="sec-2-1-4">
          <title>OwnFilter FST</title>
        </sec>
        <sec id="sec-2-1-5">
          <title>NoFilter</title>
        </sec>
        <sec id="sec-2-1-6">
          <title>OwnFilter FST CFW 0.194</title>
          <p>
            • CountWordsFree [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]: CountWordsFree is a web-based service for content writers, web
developers, and professionals who provide search ranking optimization services such as
the list of stopwords we have used;
• Google Stopwords [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]: set of common words that are filtered out or ignored by Google’s
search engine algorithms when processing search queries;
• Stopwords ISO [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]: collection of stop words for various languages that are commonly
used in natural language processing (NLP) tasks;
• Ranks.nl [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]: Dutch website that provides various online marketing services and tools;
• Savoy Stopwords [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]: it is a standard library that provides a collection of robust,
highperformance libraries for mathematics, statistics, data processing, streams, and more and
includes many of the utilities you would expect from a standard library;
• NoStop: it is simply an empty list, to evaluate if the filter was enhancing the performances
or not.
          </p>
          <p>As we can see from the obtained values, the Google stopwords list is the one that performs
best, without stoplist based query boosting. It is worth noting that it is the shortest list that
was tested.
2.1.1. Tested but unused filters
The systems submitted by our group are the results of numerous trials and runs, and
unfortunately many of our attempted approaches proved themselves to be inefective or unfeasible.
In our first prototypes, a LengthFilter was used to remove the words with a number of
characters below and over certain thresholds. Initially it incremented the model’s scores, but
upon further refinements of other parts of the system it started having a negative impact on
performance.</p>
          <p>A few diferent configurations of a ShingleFilter were also attempted to create tokens from
overlapping sequences of n words (token n-grams) from a token stream. We found they weren’t
suited to the LongEval task.</p>
          <p>
            Multiple attempts were made to integrate an NLP (Natural Language Processing) library into
our system. It could be used for many tasks, such as tokenization, sentence segmentation,
part-of-speech tagging and named entity extraction. Specifically, we focused on two diferent
libraries: OpenNLP [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] and CoreNLP [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. We tried running many configurations of both
libraries, however for a dataset as large as LongEval’s they were requiring computational power
well beyond what we had access to, therefore we were forced to discard the idea.
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Indexer</title>
        <p>The Indexer is a class responsible for creating an index of the terms in a collection of documents.
The purpose of the indexer is to speed up the retrieval of relevant documents in response to user
queries. It creates an internal data structure, the ’index’, that allows quicker access to data. An
indexer reads through the text data, identifies important and searchable entities, and constructs
an index of those entities and their location in the data.</p>
        <p>Our custom indexer, called LongEvalIndexer, is an update of the initial indexer design
user for the Tipster collection discussed during the lectures. The latest version of it, recursively
visits each file in the specified document directory and extracts the contained text data through
the Parser class. Each document is parsed and javascript or PHP code is removed through
the TextFilter class. Subsequently, each document is stored in the index and divided into the
following fields:
• IDField: this field stores the unique identifier associated with each document;
• BodyField: the body of the document is indexed and stored in this field after being
processed by the parser. The field type is defined using a FieldType object, which is
configured to store the term vectors, positions, and ofsets;
• LinkField: the link corresponding to the document is extracted from the URLs file
provided in the LongEval dataset and is processed to remove specific characters in order
to obtain the words contained in the link. It is then added to the document using a custom
LinkField, which is configured to store the term vectors, positions, and ofsets.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Searcher</title>
        <p>The Searcher is a class used to represent the user or the query that seeks information from a
collection of documents, so its main function is to express the information need of the user in
the form of a query. This is the last class to be used in the information retrieval process since to
perform research on the documents they have to be indexed first. The final version of this class
takes as input a set of queries from the set provided by LongEval and it builds a BoostQuery
object from each query. The BoostQuery class allows to give a boost to the wrapped query.
Boost values that are less than one will give less importance to this query compared to other
ones while values that are greater than one will give more importance to the scores returned
by this query. Combined with this class, we exploited the MultiFieldQueryParser class,
which is used to associate queries and documents with multiple fields, also defining the weight
they will have in the search. In our case, the fields we researched were "body" and "link". Using
diferent weights, we noticed that the weights given to these fields were crucial for the model’s
performance.</p>
        <p>
          Other approaches for the implementation of this class were attempted but did not result in
performance gains. We tested two types of query expansion techniques, one based on the top
relevant documents for each query and another one developed with the help of gpt-3.5-turbo
model by OpenAI [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The functioning of the first one was based on an algorithm that found
the most relevant words in the top ten retrieved documents for each query and then it added
these words to the main query. After that, a search on the collection was repeated with the
expanded query. Unfortunately, this technique proved inefective, as it significantly lowered
the accuracy of the system. This due to the low accuracy rate that the system can achieve even
in the best-ranked dopuments. This led to adding more noise words than relevant and thus
worsening the search, forcing us to abandon this approach.
        </p>
        <p>The second query expansion we tested made use of an NLP model developed by OpenAI. The
idea was to give the model a single query as input and ask it to return an expanded one. The first
problem we encountered was linked to the number of calls we could do, by default OpenAI limits
users to 3 calls per minute, but each run was about 700 queries. In order to avoid this bottleneck,
we formed a prompt to expand 100 queries in each call. Unfortunately, the model was not
particularly accurate and sometimes it diverged too much, expanding the queries incorrectly.
For this reason, we had to abandon this idea too, but we believe that with some refinements it
could be a good improvement for a future model.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Similarity</title>
        <p>Similarity function is a fundamental component used to measure the similarity, or relatedness,
between queries and documents. The function returns a score that reflects how well the
document matches the query. During the development process the BM25 similarity model with
default parameters was used as a baseline. The systems our team is submitting for the LongEval
task have various configurations of similarity models in order to test how they perform with
separate but similar datasets and study any trends that emerge. The models being tested are
described in the Results 4 section.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Setup</title>
      <sec id="sec-3-1">
        <title>3.1. Data description</title>
        <p>The data used in the project are the evaluation collection provided for task 1,
"LongEvalRetrieval", consisting of collections of web documents and queries provided by Qwant, a search
engine. Both documents and queries were collected in French and then automatically translated
into English. This scheme represents the collection process: The data can be described as
follows:
3.1.1. Documents
The collection includes relevant documents that are selected to be retrieved for each query and
potentially irrelevant documents randomly sampled from the Qwant index to better represent
the nature of a Web test collection.</p>
        <p>• Train data: consists of 1,570,734 web pages, acquired during June 2022. They can be
downloaded from the Lindat/Clarin website.
• Test data: consist of 1,593,376 documents and 882 queries, collected over July 2022, for
the short-term persistence sub-task and 1,081,334 documents and 923 queries, collected
over September 2022, for the long-term persistence sub-task. Both can be downloaded
from the Lindat/Clarin website.
3.1.2. Topics
3.1.3. Queries
The queries are extracted from Qwant’s search logs, based on a set of selected topics, but exact
details regarding them were not made available to us.</p>
        <p>The queries are extracted from Qwant’s search logs, based on a set of selected topics. The query
set was created in French and was automatically translated into English.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Evaluation Measures</title>
        <p>
          The evaluation tools used during development are trec_eval [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], an executable that through
qrels allows us to obtain measures of precision, accuracy, etc., and Luke [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], a Lucene GUI
that allows us to look inside the index, to check the health and consistency of the indexes.
The main evaluation measures used to check the system during the various stages of development
are:
• nDCG or Normalized Discounted Cumulated Gain: it measures the efectiveness of a
retrieval system by considering the relevance of documents and their positions in the
ranked list. In our case, we used an evaluation depth of 5 (nDCG@5). To compute it
is necessary to normalize the score by dividing nDCG by the iDCG (Ideal Discounted
Cumulated Gain). These are the formulas to compute them:
@ = ∑︁ 
=1 (1, ())
        </p>
        <p>(1)
where  and  are the precision and recall at the ℎ threshold.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Repository</title>
        <p>The git repository containing the source code of the system is available in the repository at the
link seupd2223-qevals (https://bitbucket.org/upd-dei-stud-prj/seupd2223-qevals).
3.4. Hardware components
• Saverio’s PC:
- CPU: Intel® Core™ i5-8600K 3.60GHz × 6
- GPU: NVIDIA GeForce GTX 1060 6GB
- RAM: 16 GB DDR4
- SO: Ubuntu 22.04.2 LTS
• Enrico’s PC (laptop)
- CPU: Intel Core i7-8550U 1.8 GHz Base, 4.0 GHz Turbo
- GPU: NVIDIA GeForce MX250 2GB GDDR5
- RAM: 16 GiB DDR4
- SO: Pop!_OS 22.04 LTS</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Discussion</title>
      <p>
        The numerous runs that were executed on the LongEval training dataset provided us with
good baseline analyzer and searcher configurations that we could use to compare how diferent
similarity models performed on the LongEval task. Specifically, we learned that:
• Better results were obtained working on the dataset containing documents in the original
language compared to the English-translated one.
• In the case of French queries and documents the best-performing tokenization is obtained
through Lucene’s LetterTokenizer.
• Stop Lists can improve performance, but most available lists tend to over-filter the data
to the point of performing worse than using no filter at all.
• In case of stop list over-filtering it is possible to regain some performance by query
boosting tokens in the query that are specifically not in a stop list, even though the stop
list tokens get filtered later in the pipeline.
• The best performing Stop Lists had at most 200 words. Specifically, we chose the Google
stop list [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] as it lightly edged out the Ranks stop list [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
• The use of stemmers had little efect, but seemed to reduce overall performance. We
tested LightFrenchStemmer and MinimalFrenchStemmer, available on Apache Lucene,
and chose to proceed without using either.
• Adding more filters (other than the LowercaseFilter, the ElisionFilter and the
ASCIIFoldingFilter) such as the LengthFilter, in an attempt to exclude less useful terms, led to worse
performance.
• The use of word n-gram tokens was inefective, possibly because such short queries are
often not semantically coherent sentences, so they would not match the contents of the
longer documents.
• Our implementations of Query Expansion did not perform well, though we are uncertain
whether the LongEval dataset is ill-suited to this approach or our implementations had
inherent issues.
• Given the very short format of the queries adding the link field of the document to the
index led to notable performance improvements.
      </p>
      <p>Using the configuration described above five diferent similarity models were tested:
• QEVALS_BM25DFLT: BM25 similarity model with default parameters (k1 = 1.2, b =
0.75).
• QEVALS_BM25CSTM: BM25 similarity model with the best-performing parameters that
were found (k1= 1.2, b = 0.9).
• QEVALS_LMDirichlet: Language model based similarity with Bayesian smoothing
using Dirichlet priors.
• QEVALS_DFR: Probabilistic model that measures the divergence from randomness. The
three components chosen were: Geometric approximation of Bose-Einstein, Laplace’s
law of succession and Uniform distribution of term frequency.
• QEVALS_IB: Information-based model. The three components chosen were: Log-logistic
probabilistic distribution, Total Term Frequency Lambda and Uniform distribution of
term frequency.</p>
      <sec id="sec-4-1">
        <title>4.1. Training set</title>
        <p>Following is the summary table reporting the MAP and nDCG values obtained in the training
set for each run:
nDCG</p>
        <p>As reported, the values of each run are similar, in fact the variances for both sets of measures
are very low. This is due to the fact that the changes across the models are not so relevant,
indeed the main structure is almost the same.</p>
        <p>Here instead, there is the Interpolated Precision vs Recall graph for the five tested similarity
models:</p>
        <p>Interpolated Precision vs Recall</p>
        <p>QEVALS_BM25DFLT
QEVALS_BM25CSTM
QEVALS_LMDirichlet</p>
        <p>QEVALS_DFR
QEVALS_IB
0.4
0.35</p>
        <p>0.3
0.25</p>
        <p>0.2
0.15</p>
        <p>0.1
n
o
i
s
i
c
e
r
P
5 · 10− 2
0
0</p>
        <p>0.5</p>
        <p>Recall
0.1
0.2
0.3
0.4
0.6
0.7
0.8
0.9
1</p>
        <p>The graph shows that the performances of the chosen similarity models are quite close to
each other for the LongEval training dataset. In particular, BM25 performs best overall, and
changes to the default parameters can result in further (though limited) gains in performance.
The language model is very slightly better than BM25 at low recall, but drops of more steeply
at around 30% Recall, while the statistical models have comparable performance to BM25 at
most points, but lag behind at low recall.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Test Set</title>
        <p>This section aims to provide a comprehensive analysis of the performance of the five models
submitted for evaluation. Various metrics are used to evaluate the efectiveness and eficiency
of each model. These metrics include average precision(MAP), precision(P@10), normalized
discounted cumulated gain (nDCG) and recall (Recall_1000).</p>
        <p>First, we present three tables, each corresponding to one of the test sets. These tables compare
the performance of each run based on the metrics listed above.</p>
        <p>Furthermore, we reported the boxplots calculated from the average precision (AP) respectively
for short and long term. These graphical representations are intended to provide a detailed
overview of the variability and distortion of the system’s accuracy values.</p>
        <p>Next, there are the comparative graphs of the MAP values obtained with the Tukey’s HSD of
the runs for short and long term. This will ofer a comparison of the mean performance of the
systems, highlighting their respective strengths and weaknesses.</p>
        <p>Subsequently, we reported the ANOVA tables (one-way) respectively for long and short term.
These tables show how the sum of squares are distributed according to source of variation, and
hence the mean sum of squares.</p>
        <p>Finally, we have plotted the topics performance at 5 for short and long term.
0.1329
0.1296</p>
        <p>Heldout</p>
        <p>Short
nDCG
nDCG</p>
        <p>In Tables 6, 7, 8 the observed values of the diferent "runs" of the system are similar, without
significant deviations, and we can observe that the MAP values in both tables remain around 20%.
It is interesting to note that the QEVALS_IB run performed well in the heldout set, obtaining
the best results (map: 0.2108, p@10: 0.1329, nDCG: 0.3597), indeed it did not scored as expected
in the Short and Long Test datasets. The values obtained from run QEVALS_BM25DFLT in the
short and long test set seem slightly better than the others.</p>
        <p>From the boxplots in Figure 3 we can see that the accuracy values obtained from the runs are
very similar, also we see that QEVALS_LMDirichlet and QEVALS_BM25CSTM seem to perform
well in one test set and less in the other, instead QEVALS_DFR and QEVALS_BM25DFLT seem
to maintain performance in both.</p>
        <p>From the Figure 4, even if there are not relevant diferences in performances, as already
confirmed from the previous analysis, we can highlight that the best run is QEVALS_BM25DFLT.</p>
        <p>From the analysis of the two-way ANOVA, visible in Table 9 and Table 10, it is confirmed
how there is no substantial diference between the runs. A low F-value statistic suggests that
the variation between the groups is relatively small compared to the variation within them; in
addition, a high F-value probability indicates that the probability that the observed diferences
are due to chance is very high. This further confirms that the diferences between runs tested
are not statistically significant.</p>
        <p>Similarly to the training dataset, the diferent runs on the short-term and long-term datasets
manifested significant consistency, with similar MAP, nDCG, and Recall values. This indicates
a relative stability of system performance and shows the capacity of adaptation to the temporal
evolution of the documents.</p>
        <p>As expected, the overall performance of the runs showed no significant variation between
our training baseline and the short-term and long-term document sets. Figure 5, in particular,
shows that our system’s performance is virtually identical for the entire range of topics of the
short-term and long-term datasets. The capability to have consistent performance regardless
of the contents of a dataset can be considered an advantage of a system such as ours, which
does not use past data to train a model for future applications. One notable aspect that we can
infer from the test data is that the fine-tuning of the parameters of the BM25 similarity model is
quite dependent on the available data. In fact, for both short-term and long-term dataset test
runs, the default BM25 model performed slightly better than our fine-tuned one.
However, in general, the overall performances don’t change drastically. Indeed, the values of
the measurements difer at most ∼ 2%. We can conclude that for a task such as LongEval the
similarity model choice alone has little impact on the overall performance of the system.</p>
        <p>Finally, QEVALS_BM25DFLT (the default configuration of the BM25 similarity model) proved
to be the best performing run, achieving the highest MAP, nDCG and Recall values among the
diferent runs evaluated. Our conclusion is that this system is the most suited to the task out of
the ones we tested.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Future Work</title>
      <p>While our systems do not perform much better than the baseline provided by LongEvals, the
process to arrive at the systems we are submitting ofered us an opportunity to acquire experience
and knowledge in dealing with information retrieval systems. Many of our approaches did not
perform as hoped, however thanks to our numerous trials we are much better equipped to try
new approaches to information retrieval in the future. Moreover, we are aware of many possible
enhancements to the project.</p>
      <sec id="sec-5-1">
        <title>More studies on the datasets</title>
        <p>Our first challenge was the language. At a baseline, performance is better on the dataset in the
original language, so our team decided to focus more on the French dataset early on. Our initial
attempts on the English dataset seemed less promising, but it’s possible that we would have
had more success implementing the more complex systems for it.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.1. Usage of NLP</title>
        <p>The use of an NLP library such as OpenNLP or CoreNLP would have helped us take a big step
forward. We learned that using Natural Language Processing technology to implement the
whole model is probably unfeasible since it would be too expensive from a computational point
of view, but we believe that a part of part-of-speech tagging or named entity extraction would
have been helpful. Therefore, if we had a chance to run the project on more powerful computers,
and suficient time to build eficient parallelization algorithms, we believe that the model could
gain better results and performances</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.2. Query expansion techniques</title>
        <p>Another important aspect we were not able to implement properly was a query expansion
technique, as discussed in the Methodology 4 section. There are many algorithms that could be
used and some of them are already present in some libraries. The main problem was that our
dataset was in French and we have not been able to find a good dictionary in order to expand
the queries. So, we think that a query expansion technique with a proper dictionary could be
a good addition to the project. Moreover, the two algorithms we tried to implement could be
improved. In fact, the first one, which expands the query using the top retrieved documents’
words, may already be in a working state, but in order for it to work well, the rest of the system
must provide precise enough results at low recall values. This is due to the fact that a better
base model leads to more relevant documents and so, to more relevant words with which to
expand the query. Due to low precision at low recall we were adding noise to the initial query
that, in the end, was performing worse than the original one.</p>
        <p>The query expansion technique which was using the gpt-3.5-turbo model, was not a
success either, however, it must be noted that compared to our other approaches it is the one
we worked on the least because of time limitations and OpenAI’s API access quota limitations
for non-paying users. The problem our team ran into was the unpredictability of the answers:
the model was asked to expand each query to five words, but in some instances it expanded
queries to up to fifteen words, adding useless noise. For such a system to work, some additional
experience in prompt formulation for information retrieval purposes would likely bring much
better results.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>CLEF</given-names>
            ,
            <surname>Longeval</surname>
          </string-name>
          ,
          <year>2023</year>
          . URL: https://clef-longeval.github.io/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          ,
          <year>Clef 2023</year>
          ,
          <year>2022</year>
          . URL: http://clef2023.clef-initiative.eu/index.php.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Qwant</surname>
          </string-name>
          , About qwant,
          <year>2011</year>
          . URL: https://about.qwant.com/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Alkhalifa</surname>
          </string-name>
          , I. Bilal,
          <string-name>
            <given-names>H.</given-names>
            <surname>Borkakoty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Camacho-Collados</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Deveaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>El-Ebshihy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Espinosa-Anke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gonzalez-Saez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuscakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          , E. Kochkina,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liakata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Loureiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. T.</given-names>
            <surname>Madabushi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mulhem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Popel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Servan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zubiaga</surname>
          </string-name>
          ,
          <article-title>Overview of the clef-2023 longeval lab on longitudinal evaluation of model performance, in: Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Fourteenth International Conference of the CLEF Association (CLEF 2023), Lecture Notes in Computer Science (LNCS)</source>
          , Springer, Thessaliniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Apache</surname>
          </string-name>
          , Apache lucene,
          <year>2000</year>
          . URL: https://lucene.apache.org/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6] CountWordsFree, Countwordsfree french stopwords,
          <year>2023</year>
          . URL: https://countwordsfree. com/stopwords/french.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Google</surname>
            <given-names>french stopwords</given-names>
          </string-name>
          ,
          <year>2023</year>
          . URL: https://meta.wikimedia.org/wiki/Stop_word_list/ google_stop_word_list#French.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Stopwords-Iso</surname>
          </string-name>
          ,
          <article-title>Stopwords-iso/stopwords-fr: French stopwords collection</article-title>
          ,
          <year>2023</year>
          . URL: https: //github.com/stopwords-iso/stopwords-fr.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Ranks</surname>
          </string-name>
          .nl french stopwords,
          <year>2023</year>
          . URL: https://www.ranks.nl/stopwords/french.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Stdlib-Js</surname>
          </string-name>
          ,
          <article-title>Stdlib-js/datasets-savoy-stopwords-fr: A list of french stop words</article-title>
          .,
          <year>2023</year>
          . URL: https://github.com/stdlib-js/
          <article-title>datasets-savoy-stopwords-fr.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Apache</surname>
          </string-name>
          , Opennlp,
          <year>2004</year>
          . URL: https://opennlp.apache.org/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. N.</given-names>
            <surname>Group</surname>
          </string-name>
          , corenlp,
          <year>2010</year>
          . URL: https://stanfordnlp.github.io/CoreNLP/.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>OpenAI</surname>
          </string-name>
          , Models - openai,
          <year>2023</year>
          . URL: https://platform.openai.com/docs/models/gpt-3-5.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14] NIST, trec_eval,
          <year>2023</year>
          . URL: https://trec.nist.gov/trec_eval/.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Apache</surname>
          </string-name>
          , Luke,
          <year>2009</year>
          . URL: https://github.com/DmitryKey/luke.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>