<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Diverse Semantics Representation is King</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Martin Geletka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vojtěch Kalivoda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michal Štefánik</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marek Toma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Petr Sojka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bill Gates Petr Sojka</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Informatics, Masaryk University</institution>
          ,
          <addr-line>Botanická 68a, 602 00 Brno</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>We report on the systems that the Math Information Retrieval group at Masaryk University (mirmu) and the team of Faculty of Informatics students (msm) prepared for task 1 (find answers) of the arqmath lab at the clef conference. To study the efects of diferent system settings and hyperparameters, we have prototyped several diverse math-aware information retrieval (mir) systems: both “old” inverted index-based ones and new neural ones. By ensembling the results of the “weak” individual systems into committees, we report on entailments, benefits, and drawbacks of system ensembling. We evaluated the proposed individual systems and ensembles, considering their diversity, hyperparameters, and representations used, and classified their approaches. Our prototypes have helped to understand the challenging problems of question-answering in the stem domain: the key lies in the proper representation of document semantics. Our reproducible evaluation Python library PV211-utils allows to reproduce and further advance mir re-search.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Information retrieval</kwd>
        <kwd>question answering</kwd>
        <kwd>math representations</kwd>
        <kwd>math-aware information retrieval</kwd>
        <kwd>word embeddings</kwd>
        <kwd>ensembling</kwd>
        <kwd>voting</kwd>
        <kwd>reranking</kwd>
        <kwd>data fusion</kwd>
        <kwd>diversity</kwd>
        <kwd>transformers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Math Information Retrieval (mir) and math-aware representation of meaning of scientific
documents have been researched at mir laboratory at Masaryk University for decades, as nicely
summarized by Novotný [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in his dissertation. As in the previous year [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], we formed two
teams, mirmu and msm. Under the mirmu team, we submitted five diferent versions of the deep
neural information system, which tries to overcome the performance of tf-idf likes systems.
Under the msm group, we submitted diferent versions of the student information systems and
their ensemble with the best variant from the mirmu submission. Finally, we report that an
ensemble of all fine-tuned individual systems’ by reciprocal rank fusion performed best.
      </p>
      <p>
        Our arqmath reports [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] showed promising directions stemming from the enormous
capacity of neural languages models, their ensembling [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], their diferent training sets,
hyperparametrization, input preprocessing, and math tokenization.
      </p>
      <p>
        Our systems were mainly developed as part of the Information Retrieval course PV211 and
will allow reproducible research using Python package PV211-utils [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>In Section 2 we describe resources and methods used to train and develop our systems.
Section 3 describes systems and strategies used to prepare our runs. We report and evaluate
our results in Section 4 on page 6. We are summing up our conclusions with Section 5,
drawing possible research plans based on computed metrics and availability of collected systems,
ensembling techniques, and ground truth datasets.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Datasets and Methods</title>
      <p>This section describes the math representations ingested by our information retrieval systems,
the corpora used for training the models that power our systems, and the relevance judgments
we used for parameter optimization, model selection, and performance estimation.</p>
      <sec id="sec-2-1">
        <title>2.1. Math Input Representations</title>
        <p>We used the most straightforward math representation for all our submitted systems: LATEX.</p>
        <p>In one submitted run of mirmu group, we studied the efect of the L ATEX encoding of math
compared to the sole text to see how the presence of the math representation afects the resulting
score.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Datasets and Methods</title>
        <p>
          We described our datasets, ensembling methods, and evaluation measures in detail in our
previous reports [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and [2, Section 2]. We also used datasets from arqmath 2021 and arqmath
2022 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Systems Description</title>
      <p>The following sections describe nine systems students have developed as part of their studies.
Their diversity brings diferent ways to represent the meaning of math content and how to pick
and rerank the answers for given topics. Table 1 on page 7 summarizes ten submitted runs by
both MU teams.</p>
      <sec id="sec-3-1">
        <title>3.1. Retriever + ReRanker System</title>
        <p>
          Our systems submitted under the mirmu group consist of the following parts applied sequentially
after one another, inspired by RE3QA architecture [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>• Indexer – to assign each document a dense vector representation;
• Retriever – to compute the dense vector representation of the input query and to compute
the cosine similarity between query and each document representation and sorting all
documents by this similarity;
• ReRanker – to rerank top- most relevant documents from Retrieving part. We achieved
the best results with tiered reranking, taking multiple non-overlapping slices from top-
results and reranks each slice separately. This part is computationally expensive; therefore,
we cannot do it on the whole dataset in practice.</p>
        <p>
          For implementing our system, we used the Sentence Transformer library, which is the Python
framework for state-of-the-art sentence, text, and image embeddings. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
        </p>
        <sec id="sec-3-1-1">
          <title>3.1.1. Implementation and Hyper-parameters</title>
          <p>
            For the implementation of the Base model, we used the BiEncoder model with the pre-trained
all-MiniLM-L12-v2 model [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]. For the ReRanking phase, we fine-tuned the CrossEncoder
model with the pre-trained roberta-large model [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ]. Identically, we experimented with
math-specific CrossEncoder model, MathBERTa 1 [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ], that extends generic RoBERTa with
math-specific tokenization and fine-tunes the extended model using ArXiv collection.
          </p>
          <p>We used Sentence Transformers library2 for fine-tuning of all our models. The architecture
of both models and their diferences are depicted in Figure 1. The Base model fine-tuned only
the ReRanker model, as the fine-tuning of the generic Retriever pre-trained for a retrieval on a
vast and heterogeneous datasets did not bring measurable benefits of quality. We submitted
ifne-tuned version of the ReRanker as the alternative run.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.1.2. ReRanker fine-tuning</title>
          <p>The ReRanker was fine-tuned using the BCE with logits loss and the AdamW optimizer. The
training set consisted of relevance judgments, a train+validation set from ARQMath 2021,
extended by additional samples. The additional positive samples, with the label equal to 1, were
pairs of a question and an answer from the same thread, where the answer received at least
50 upvotes. The additional negative samples, with the label equal to 0, were the most similar
documents given by Retriever but with no overlapping Math Stack Exchange tags. Because
ifnding similar documents for all questions would not be feasible, we decided to generate ten
negative samples for ten random questions at the start of each epoch.</p>
          <p>The model was trained on all training queries and the same number of additional samples
every epoch. The ratio of positive and negative samples was 0.5. The training process was
stopped as there was no decrease in loss over the test set from ARQMath 2021.</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>3.1.3. Description of Individual Runs</title>
          <p>
            For the mirmu team, we submitted diferent versions of the described system to see the efect of
the diferent components of the system. The variation we used are:
• finetuning / not finetuning of the Retriever model
• using ReRanker model pretrained on the Math texts (RoBERTa vs MathBERTa3) [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]
1https://huggingface.co/witiko/mathberta
2https://www.sbert.net/docs/pretrained_models.html
3https://huggingface.co/witiko/mathberta
          </p>
          <p>
            • using diferent representation of the input (text vs text + L ATEX)
We submitted the Base model as the primary as we couldn’t outperform its performance. The
other variants we submitted as the alternative runs are:
• Base – model described in the previous section;
• Trained only on text – system using only the text representation of documents and
queries;
• MathBERTa ReRanker – system using MathBERTa model as the ReRanker;
• Retriever fine-tuned – Retriever model adjusted by the relevance judgements collected
over ARQMath2020 and ARQMath 2021, using Negatives Ranking loss [
            <xref ref-type="bibr" rid="ref13 ref5">13, 5</xref>
            ];
• MathBERTa ReRanker + Retriever fine-tuned – system with MathBERTa ReRanker
and fine-tuned Retriever model.
          </p>
          <p>
            To quantify the efect of the individual changes in the mirmu system, only one parameter was
changed at time for given run, and all other hyperparameters were left the same.
3.2. tf-idf
As one of the students’ baseline systems, we used the tf-idf model implementation available in
the Gensim library [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ].
          </p>
          <p>In the preprocessing phase, we removed extreme values (below 8 absolute term frequency
and higher than 0.7 relative document frequency), where we chose the hyperparameters
experimentally on train subset. Then we removed punctuation, repeating whitespaces. Lastly, we
tokenized the document by splitting on whitespace and stemmed individual tokens using the
snowball stemmer available in snowball_py 4 library.</p>
          <p>
            We used smart (System for the Mechanical Analysis and Retrieval of Text)    tf-idf
weighting variant [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. We used logarithm for term frequency weighting, no document weighting,
and cosine document length normalization.
          </p>
          <p>
            4https://github.com/shibukawa/snowball_py
3.3. BM25
BM25+ is an improvement over BM25 introduced by Lv and Zhai [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]. Together with other
alternatives, such as BM25-L, BM25-adapt, and BM25-T, this improvement surpasses the basic
BM25 algorithm on trec collections. [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] BM25+ estimates the relevance of a document  for a
query  by formula (1).
          </p>
          <p>BM25+(, ) = ∑︁ log
∈
︂(  + 1 )︂
df</p>
          <p>⎛
· ⎝</p>
          <p>(1 + 1) · tf,
1 · ︁( (1 − ) +  avg
︁( d )︁ )︁
+ tfd</p>
          <p>⎞
+  ⎠,
(1)
where 1, , and  are hyperparameters,  is the number of documents in the collection, df is
the number of documents containing the term , tf, is frequency of term  in document , 
the length of document  in words, and avg is the expected length of a document in words.</p>
          <p>We represented our answers as a concatenation of its body and title, body, and tags of its
parent question. In the next preprocessing stage, we firstly removed punctuation, repeating
whitespaces, and English stopwords. Then we transformed the text into lowercase and stemmed
it using Porter stemmer. Lastly, we tokenized the text by splitting it on whitespace.</p>
          <p>
            We used the implementation of BM25+ in the rank_bm25 Python library [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] with its default
hyperparameters.
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.4. BM25 + tf-idf ensemble</title>
        <p>As the example of the most straightforward possible ensemble system of two unsupervised
systems, we used the ensembling of tf-idf and BM25 models. We constructed the ensemble
as the simple sum of the given scores from the individual systems. The BM25 is configured
as described in Section 3.3. The tf-idf system uses the same preprocessing as BM25 system
and tf-idf implementation available in the Gensim library with smart    tf-idf weighting
variant, which corresponds to logarithmic term frequency weighting, zero-corrected idf, and
no document normalization.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.5. Compubert</title>
      </sec>
      <sec id="sec-3-4">
        <title>3.6. Ensemble Systems</title>
        <p>
          We also submitted the Compubert model prepared and submitted by the mirmu team last year.
The model and hyper-parameters can be found in last year’s report [2, Section 3.4].
This section describes the used ensemble systems. With ensembling and voting techniques, we
can combine the strengths of diferent systems to produce more accurate results. Historically,
there is a long tradition of boosting, [
          <xref ref-type="bibr" rid="ref19">19, 20</xref>
          ], ensembling [21], data fusion [22] and voting
approaches [23, 24] in the information retrieval research.
        </p>
        <p>
          As we believe that our systems could agree on a small portion of the most relevant documents,
reflecting diferent ‘points of view’ on the search problem. Depending on dozens of parameters,
each individual system will miss the great majority of relevant documents. With ensembling and
voting techniques, we can combine the strengths of diferent systems to produce more accurate
results. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. All our ensemble algorithms are agnostic to the scoring functions of individual
systems and only use the ranks of the results.
        </p>
        <p>
          The ibc is ensemble technique we introduced in arqmath 2020 in our paper [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The ensemble
combines SERPs from the individual systems by Median Inverse Rank, which is equal to
(1000 −  )/1000, where  is Median Rank of individual systems. For detail explanation of
the ibc ensemble algorithm we refer the reader to our arqmath 2020 paper. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
We used reciprocal rank fusion (rrf) [25] to construct the ensemble model from the previously
described systems. The ensemble, given ranks from all individual systems, sorts the documents
by Formula (2).
        </p>
        <p>rrf+() = ∑︁
∈</p>
        <p>1
 + ()
(2)
where  is set of rankings and () is ranking of document . The hyper-parameter 
parameterizes the ensemble. We used default value of  = 60, suggested in [25].</p>
        <p>As the rbc model we refer to trained regression model, which predicts the gain of train
judgements from the ranks of the individual.</p>
        <p>For the performance estimation of rbc, we produced a result list by taking the 1,000 answers
with the highest predicted gain for each topic in the test subset.
3.6.1. IBC
3.6.2. RRF
3.6.3. RBC
3.6.4. WIBC
wibc is weighted variant of the ibc algorithm. Instead of electing the candidate with the highest
median rating, wibc elects the candidate with the highest weighted median rating. Instead of
breaking ties by selecting a random rating out of a uniform distribution of all ratings, we select
a random rating out of a weighted uniform distribution.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation and Results</title>
      <p>To compare the submitted systems, we evaluate their performance on the topics from previous
arqmath competitions.</p>
      <sec id="sec-4-1">
        <title>4.1. Submitted runs</title>
        <p>
          For each system, we report the resulted scores as in Table 2 in similar form as in the overview
arqmath 2022 paper by Mansouri et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>Performance drop in mirmu 2 run compared to other runs indicate that training and adjusting
the system on math is a must.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Runs with enhanced systems and ensembles</title>
        <p>We further fine-tuned our systems benefiting from more ground truth data and experience we
got by previous evaluations.</p>
        <p>• MathBERTa ReRanker – MathBerta ReRanker fine-tuned on altered preprocessing;
• BM25-based system – BM25-based system msm 3 described in Section 3.3, where we
optimized hyperparameters 1, , and  using grid search. The values we found to yield
the best ndcg′ on our training set are 1 = 1.8,  = 0.75, and  = 1.
• Improved Base – Base system with reduced preprocessing, more precisely trained
ReRanker and improved tiered reranking with slices on indexes: 3, 7, 12, 16, 20, 50, 100,
125.
• Retriever only – A system with removed reranking phase. Document ranking is based
only on cosine similarity on embeddings obtained from Retriever.</p>
        <p>We report the results of extended experiments with ensembles in Table 3 on the next page.
rrf ensembles deliver best results by far margin. The more diverse systems one combines, the
more the quality metric as ndcg′ monotonously grows. As the math information systems have
to cope with really complex problems, it is really hard to build one complex system that is
capable of learning all the complex stuf: disambiguation of overloaded math symbols, structured
ambiguous notation, deduction and long causal dependencies.
ens 1: rrf 60 of 4 fine systems ⁓0⁓.4⁓9⁓3
ens 2: ibc of all 0.401
ens 3: ibc of all 4 msm 0.324
ens 4: ibc of all 5 mirmu 0.468
ens 5: ibc of all 4 msm +mirmu 1 0.354
ens 6: ibc of msm 4 and mirmu 1 0.459
ens 7: rrf 60 of all 0.480
ens 8: rrf 180 of all 0.486
ens 9: rrf 60 of all 4 msm 0.328
ens 10: rrf 60 of all 5 mirmu 0.465
ens 11: rrf 60 of all 4 msm and mirmu 1 0.422
ens 12: rrf 60 of msm 4 + mirmu 1 0.465
ens 13: rbc of all 0.476
ens 14: rbc of all 4 msm 0.312
ens 15: rbc of all mirmu 0.468
ens 16: rbc of all 4 msm and mirmu 1 0.475
ens 17: rbc of msm 4 and mirmu 1 0.474
ens 18: wibc of all 0.466
ens 19: wibc of all 4 msm 0.332
ens 20: wibc of all mirmu 0.466
ens 21: wibc of all 4 msm and mirmu 1 0.488
ens 22: wibc of msm 4 and mirmu 1 0.466</p>
        <p>The Figure 2 shows the dependence of rrf quality on the parameter . Instead of suggested
 = 60 from the original paper, the best performance is with  = 180. We hypotetize that the
more diverse the primary systems are, the higher optimal parameter  scores.</p>
        <p>It remains to be studied to which extend the performance depends on participating individual
systems. Also, how the performance change with diferent hyperparameters and initial setting
and choices of pre-trained models, diversity and initial setup of individual systems.</p>
        <p>
          For reproducibility, we are going to publish our notebooks, ensemble implementations and
models in our lab and course repository [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>“You must cultivate activities that you love. You must discover work that you do, not for its
utility, but for itself, whether it succeeds or not, whether you are praised for it or not, whether
you are loved and rewarded for it or not, whether people know about it and are grateful to you
for it or not.” Anthony de Mello</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion &amp; Future Work</title>
      <p>We have developed nine mir systems with as diverse approaches and variants as possible. We
have evaluated them on available arqmath data from last three years. We have studied the
ways how they could be ensembled to gain better performance. We have reported our findings:
a) math-aware representation with deep models started to outperform flat token-level based
systems b) ensembling done with expertise and insight about merits of individual systems
matter.</p>
      <p>In the future, we plan to further enhance our neural models and ensembling algorithms by
several means:
• evaluation of diferent ensembling strategies based on individual systems’ types and
hyperparameter settings of neural systems;
• evaluation of initial setting and hyperparameters of neural Retriever and ReRanker
systems;
• deep systems’ adaptation to math specifics by using Adapt r library [26];
• study robustness and out-of-domain performance of mir systems.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We thank all PV211 course students and former members of mir group for their contributions.
We thank the two anonymous reviewers for their insightful comments. We extend our gratitude
to the arqmath 2022 organizers for keeping the research of math information retrieval aflame.</p>
      <p>This work has been partly supported by the Ministry of Education of CR within the
LINDATCLARIAH-CZ project LM2018101.
[20] Q. Wu, C. J. C. Burges, K. M. Svore, J. Gao, Adapting boosting for information retrieval
measures, Information Retrieval 13 (2010) 254–270. doi:10.1007/s10791-009-9112-1.
[21] Y. Wang, I.-C. Choi, H. Liu, Generalized Ensemble Model for Document Ranking in
Information Retrieval, Computer Science and Information Systems 14 (2017) 123––151.
doi:10.2298/csis160229042w.
[22] R. Nuray, F. Can, Automatic Ranking of Information Retrieval Systems Using Data Fusion,
Information Processing and Management 42 (2006) 595–614. doi:10.1016/j.ipm.2005.
03.023.
[23] M. Mosbah, B. Boucheham, Majority Voting Re-ranking Algorithm for Content
BasedImage Retrieval, in: E. Garoufallou, R. J. Hartley, P. Gaitanou (Eds.), Metadata and Semantics
Research, Springer International Publishing, Cham, 2015, pp. 121–131.
[24] A. T. Albaham, N. Salim, Quality Biased Thread Retrieval Using the Voting Model, in:
Proceedings of the 18th Australasian Document Computing Symposium, ADCS ’13, ACM,
New York, NY, USA, 2013, pp. 97–100. doi:10.1145/2537734.2537752.
[25] G. V. Cormack, C. L. A. Clarke, S. Buettcher, Reciprocal Rank Fusion Outperforms
Condorcet and Individual Rank Learning Methods, in: Proceedings of the 32nd International
ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’09,
ACM, New York, NY, USA, 2009, pp. 758–759. doi:10.1145/1571941.1572114.
[26] M. Štefánik, V. Novotný, N. Groverová, P. Sojka, Adaptr: Objective-Centric Adaptation
Framework for Language Models, in: Proceedings of the 60th Annual Meeting of the
Association for Computational Linguistics: System Demonstrations, ACL, Dublin, Ireland,
2022, pp. 261–269. URL: https://aclanthology.org/2022.acl-demo.26.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>V.</given-names>
            <surname>Novotný</surname>
          </string-name>
          ,
          <source>Interpretable Representations for Fast and Accurate Retrieval of Mathematical Information [online], Dissertation</source>
          , Masaryk University, Faculty of Informatics, Brno,
          <year>2022</year>
          [cit. 2022-
          <volume>05</volume>
          -26]. URL: https://is.muni.cz/th/o4thd/Revidovana_verze_po_obhajobe_ disertace.pdf, supervisor: Petr Sojka.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>V.</given-names>
            <surname>Novotný</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Štefánik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lupták</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Geletka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zelina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sojka</surname>
          </string-name>
          ,
          <source>Ensembling Ten Math Information Retrieval Systems: MIRMU and MSM at ARQMath</source>
          <year>2021</year>
          , in:
          <source>Proceedings of the Working Notes of CLEF 2021 - Conference and Labs of the Evaluation Forum</source>
          , volume
          <volume>2696</volume>
          ,
          <string-name>
            <surname>CEUR-WS</surname>
          </string-name>
          , Bucharest, Romania,
          <year>2021</year>
          , pp.
          <fpage>82</fpage>
          -
          <lpage>106</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2936</volume>
          / paper-06.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Reusch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Thiele</surname>
          </string-name>
          , W. Lehner,
          <source>TU_DBS in the ARQMath Lab</source>
          <year>2021</year>
          ,
          <article-title>CLEF</article-title>
          , in
          <source>: Proceedings of the Working Notes of CLEF 2021 - Conference and Labs of the Evaluation Forum</source>
          , volume
          <volume>2696</volume>
          ,
          <string-name>
            <surname>CEUR-WS</surname>
          </string-name>
          , Bucharest, Romania,
          <year>2021</year>
          , pp.
          <fpage>107</fpage>
          -
          <lpage>124</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2936</volume>
          / paper-07.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rohatgi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Giles</surname>
          </string-name>
          , Ranked List Fusion and
          <article-title>Re-ranking with Pre-trained Transformers for ARQMath Lab</article-title>
          , in
          <source>: Proceedings of the Working Notes of CLEF 2021 - Conference and Labs of the Evaluation Forum</source>
          , volume
          <volume>2696</volume>
          ,
          <string-name>
            <surname>CEUR-WS</surname>
          </string-name>
          , Bucharest, Romania,
          <year>2021</year>
          , pp.
          <fpage>125</fpage>
          -
          <lpage>132</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2936</volume>
          /paper-08.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Novotný</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sojka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Štefánik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lupták</surname>
          </string-name>
          , Three is Better than One, in: CEUR Workshop Proceedings: ARQMath task at CLEF conference, volume
          <volume>2696</volume>
          ,
          <string-name>
            <surname>CEUR-WS</surname>
          </string-name>
          , Thessaloniki, Greece,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2696</volume>
          /paper_235.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Štefánik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Novotný</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Geletka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kalivoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Toma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lupták</surname>
          </string-name>
          , P. Sojka,
          <source>PV211 Utils</source>
          ,
          <year>2022</year>
          . URL: https://github.com/MIR-MU/pv211-utils/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Mansouri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Novotný</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zanibbi</surname>
          </string-name>
          ,
          <source>Overview of ARQMath3</source>
          (
          <year>2022</year>
          )
          <article-title>: Third CLEF lab on Answer Retrieval for Questions on Math (Working Notes Version)</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          , M. Potthast (Eds.), Working Notes of CLEF 2022 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-</article-title>
          <string-name>
            <surname>WS</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          , Retrieve, Read, Rerank:
          <article-title>Towards End-to-End MultiDocument Reading Comprehension</article-title>
          , in
          <source>: Proceedings of the 57th Annual Meeting of the ACL</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2285</fpage>
          -
          <lpage>2295</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          -1221.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          , Sentence-BERT:
          <article-title>Sentence Embeddings using Siamese BERTNetworks</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in NLP and the 9th International Joint Conference on NLP (EMNLP-IJCNLP)</source>
          , ACL,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China,
          <year>2019</year>
          , pp.
          <fpage>3982</fpage>
          -
          <lpage>3992</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1410.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <article-title>MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers</article-title>
          ,
          <source>in: Findings of the ACL: ACLIJCNLP</source>
          <year>2021</year>
          , ACL,
          <year>2021</year>
          , pp.
          <fpage>2140</fpage>
          -
          <lpage>2151</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .findings-acl.
          <volume>188</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .findings-acl.
          <volume>188</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , V. Stoyanov,
          <article-title>RoBERTa: A Robustly Optimized BERT Pretraining Approach</article-title>
          , ArXiv abs/
          <year>1907</year>
          .11692 (
          <year>2019</year>
          ). URL: https://openreview.net/forum?id=
          <fpage>SyxS0T4tvS</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>V.</given-names>
            <surname>Novotný</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Štefánik</surname>
          </string-name>
          , Combining Sparse and
          <article-title>Dense Information Retrieval</article-title>
          ,
          <source>in: Proceedings of the Working Notes of CLEF</source>
          <year>2022</year>
          ,
          <article-title>CEUR-</article-title>
          <string-name>
            <surname>WS</surname>
          </string-name>
          ,
          <year>2022</year>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Henderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Al-Rfou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Strope</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Sung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lukács</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Miklos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kurzweil</surname>
          </string-name>
          ,
          <article-title>Eficient Natural Language Response Suggestion for Smart Reply</article-title>
          ,
          <source>ArXiv abs/1705</source>
          .00652 (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .48550/ARXIV.1705.00652.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Řehůřek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sojka</surname>
          </string-name>
          ,
          <article-title>Software Framework for Topic Modelling with Large Corpora</article-title>
          ,
          <source>in: Proceedings of LREC 2010 workshop New Challenges for NLP Frameworks</source>
          ,
          <string-name>
            <surname>ELRA</surname>
          </string-name>
          , Valletta, Malta,
          <year>2010</year>
          , pp.
          <fpage>45</fpage>
          -
          <lpage>50</lpage>
          . doi:
          <volume>10</volume>
          .13140/2.1.2393.
          <year>1847</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Buckley</surname>
          </string-name>
          ,
          <article-title>Term-weighting approaches in automatic text retrieval</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>24</volume>
          (
          <year>1988</year>
          )
          <fpage>513</fpage>
          -
          <lpage>523</lpage>
          . doi:
          <volume>10</volume>
          .1016/
          <fpage>0306</fpage>
          -
          <lpage>4573</lpage>
          (
          <issue>88</issue>
          )
          <fpage>90021</fpage>
          -
          <lpage>0</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Log-Logistic</surname>
          </string-name>
          Model
          <article-title>-Based Interpretation of TF Normalization of BM25</article-title>
          , in: R.
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          de Vries,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zaragoza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. B.</given-names>
            <surname>Cambazoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Murdock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lempel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Silvestri</surname>
          </string-name>
          (Eds.),
          <source>Advances in Information Retrieval</source>
          , Springer, Berlin, Heidelberg,
          <year>2012</year>
          , pp.
          <fpage>244</fpage>
          -
          <lpage>255</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>642</fpage>
          -28997-2_
          <fpage>21</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Trotman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Puurula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Burgess</surname>
          </string-name>
          ,
          <article-title>Improvements to BM25 and Language Models Examined</article-title>
          ,
          <source>in: Proceedings of the 2014 Australasian Document Computing Symposium, ADCS '14</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2014</year>
          , pp.
          <fpage>58</fpage>
          -
          <lpage>65</lpage>
          . doi:
          <volume>10</volume>
          .1145/2682862.2682863.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Brown</surname>
          </string-name>
          , S. Jain,
          <string-name>
            <given-names>V.</given-names>
            <surname>Novotný</surname>
          </string-name>
          , nlp4whp, dorianbrown/rank_bm25:,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .5281/ zenodo.6106156.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gulin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Kuralenok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pavlov</surname>
          </string-name>
          ,
          <article-title>Winning The Transfer Learning Track of Yahoo!'s Learning To Rank Challenge with YetiRank</article-title>
          , in: O.
          <string-name>
            <surname>Chapelle</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , T.-Y. Liu (Eds.),
          <source>Proceedings of the Learning to Rank Challenge</source>
          , volume
          <volume>14</volume>
          <source>of Proceedings of Machine Learning Research</source>
          , PMLR, Haifa, Israel,
          <year>2011</year>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>76</lpage>
          . URL: http://proceedings.mlr.press/ v14/gulin11a.html.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>