<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Fight Against COVID-19 Misinformation via Clustering-Based Subset Selection Fusion Methods</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yidong Huang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qiuyu Xu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shengli Wu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christopher Nugent</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrian Moore</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science, Jiangsu University</institution>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computing, Ulster University</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The worldwide COVID-19 pandemic has brought about a lot of changes in people's life. It also emerges as a new challenge to information search services. This is because up to now our understanding about the virus is still limited, and there is a lot of misinformation online. In such a situation, how to provide useful and correct information to the public is not straightforward. Responsibility of search engines is crucial because many people make decisions based on the information available to them. In this piece of work, we try to improve retrieval quality via the data fusion technique. Especially, a clustering-based approach is proposed for selecting a subset of systems from all available ones for finding relevant, credible, and correct documents. Experimented with a group of runs submitted to the 2020 TREC Health Misinformation Track, we demonstrate that data fusion is a very beneficial approach for this task, whether measured by some traditional metrics such as MAP or some task specific metrics such as CAM. When choosing 17 runs, which is one third of all component retrieval systems available, the linear combination method is better than the best component retrieval system by 31.42% in MAP and 21.72% in CAM. The proposed methods are also better than the state-of-the-art subset selection method by a clear margin.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data Fusion</kwd>
        <kwd>Information Retrieval</kwd>
        <kwd>Health Misinformation</kwd>
        <kwd>Credibility</kwd>
        <kwd>COVID-19</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Since the initial cases were discovered at the end of 2019, within two years COVID-19 has
been spreading globally to almost all major countries and territories, with over two hundred
million confirmed cases and over four million deaths so far. Such a unprecedented pandemic
has impacted people’s life significantly. For many, it is very valuable to get useful and correct
information about the virus. However, this may not be as straightforward as it looks, because
there are still a lot of things we do not know about the virus and considerable misinformation
exists on the web [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and disseminates on social media [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. The consequences of such
infodemic is very harmful to the society and has negative impact on our fight against the
pandemic. Therefore, it is necessary to understand this phenomenon and develop some measures
to fight against it.
      </p>
      <p>
        Some research on this issue has been conducted so far. A few of them focus on observation
and analysis while some others focus on misinformation detection. For example, [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] analysed
a Singapore-based COVID-19 Telegram group with more than 10,000 participants. There are
a few observations and one of them is that authority-identified misinformation is rare. Both
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] analysed misinformation on Chinese Sina Weibo, while [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] did it on Twitter. In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
machine learning techniques including decision tree and convolutional neural network based
models were used to classify COVID-19 related information and misinformation.
      </p>
      <p>
        In 2020, TREC (Text REtrieval Conference) 1 held two COVID-19 related tracks: COVID [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
and Health Misinformation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In this piece of work, we focus on the ad-hoc retrieval task
in the Health Misinformation Track. 51 runs were submitted to this task by eight research
groups. 50 queries were used for this task. Two baseline runs, BM25 desc and BM25 title, were
submitted by the UWaterlooMDS group on behalf of the track organizers. Those submitted
runs are available on TREC’s web site. They provide us a very good opportunity to investigate
how system-level data fusion can improve retrieval quality in this task.
      </p>
      <p>
        Data fusion has been widely used in information retrieval for diferent tasks [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref14">11, 12, 13, 14</xref>
        ].
Fusion performance is afected by many factors including fusion methods, each of the component
retrieval systems (results) involved, the number of component systems in total, evaluation
metrics, among others. In this study, we investigate how to improve retrieval performance in
this task by using the data fusion technology. More specifically, our research question is: given
a large collection of retrieval systems, how can we choose a subset of them for efective and
eficient fusion? This has rarely been investigated before. To our knowledge, [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] is the only one
that addressed this issue. Because there are many diferent retrieval models, many components
such as name recognition, phrases, semantic relations of concepts, diferent techniques for
document credibility, and many others, it is possible to build/collect relatively a large number
of component retrieval systems for fusion. However, the eficiency of a fusion-based system
decreases when more component retrieval systems are involved. For such a retrieval system,
both performance and eficiency need to be considered. It is an important problem that deserves
research. In this paper, we propose a clustering-based method to deal with this problem. First we
apply K-means to divide all the systems into a given number of clusters, then one representative
is chosen from each cluster to form a group for fusion. In this way, both system performance
and diversity among systems can be considered at the same time. It is able to obtain better
fusion performance than those selection methods that only consider system performance only
as in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Experimented with all 51 runs submitted to the 2020 TREC Health Misinformation
Track, our results show that the method is very efective.
      </p>
      <p>The rest of this paper is organized as follows: related work is discussed in Section 2. The
proposed clustering-based data fusion method is detailed in Section 3. Section 4 presents
experimental settings and results of the proposed method and some other baseline methods.
Section 5 presents some more analytical results on the clustering method and the clusters
generated on the 2020 TREC Health Misinformation data set. Section 6 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>In this paper, we investigate how to apply data fusion to COVID-19 related health information
retrieval with considerable misinformation inside the collection. Therefore, we review some
previous work on COVID-19 misinformation detection and credible information retrieval. After</p>
      <sec id="sec-2-1">
        <title>1Its web site is located at https://trec.nist.gov/</title>
        <p>that, we review some data fusion methods and some of its application in medical information
retrieval.</p>
        <sec id="sec-2-1-1">
          <title>2.1. Misinformation Detection &amp; Credible Information Retrieval</title>
          <p>
            Since confirmed COVID-19 cases first occurred at the end of 2019 and began to spread around the
world afterwards, a lot of rumours, misinformation, and disinformation turn up on social media
and the Web, and circulate in certain communities. How to detect misinformation becomes
a key issue in medical information retrieval. Various machine learning techniques have been
used to detect misinformation. In [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ], both decision tree classifiers and convolutional neural
networks were used to classify COVID-19 related information and misinformation. [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] applied
the Elaboration Likelihood Model with four types of features: linguistic, topical, sentimental,
and behavioural features. It was found that behavioural features are more informative than
linguistic features for their detection. [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] proposed a deep learning network that could leverage
both visual and textual information. In their semantic and task level attention model, three
branches were defined to extract features of diferent types. An ensemble method was also used
for the detection. Some more work were presented in [
            <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
            ] among others.
          </p>
          <p>
            Because the Web is an open environment, documents on the Web may be in a variety of
quality. Web documents’ credibility has been a research issue for the last two decades [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ]. In
this article, we employ the term credibility with the meaning it has in [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ], where it is described
as a general concept that encompasses trustworthiness, expertise, quality, and reliability. Such
a term has been adopted in computer science for many years [
            <xref ref-type="bibr" rid="ref21 ref22 ref23 ref24">21, 22, 23, 24</xref>
            ], and it has special
importance in the Information Retrieval/Web search community [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ].
          </p>
          <p>
            For medical retrieval systems, document credibility is also an important and challenging issue
[
            <xref ref-type="bibr" rid="ref25">25</xref>
            ]. To retrieve documents that are both relevant and credible, usually a two-stage process is
taken. First documents are retrieved by only considering their relevance to the query. Then the
documents are re-ranked by considering both relevance and credibility. Some traditional models
such as BM25 can be used for relevance-concerned retrieval, while credibility of documents can
be predicted by some machine learning methods [
            <xref ref-type="bibr" rid="ref26 ref27 ref28">26, 27, 28</xref>
            ].
          </p>
        </sec>
        <sec id="sec-2-1-2">
          <title>2.2. Data Fusion</title>
          <p>
            Data fusion methods can be divided into two categories: supervised and unsupervised methods.
CombSum [
            <xref ref-type="bibr" rid="ref29">29</xref>
            ], CombMNZ [
            <xref ref-type="bibr" rid="ref29">29</xref>
            ], and the Reciprocal Rank [
            <xref ref-type="bibr" rid="ref30">30</xref>
            ] are typical unsupervised methods,
while linear combination [
            <xref ref-type="bibr" rid="ref31">31</xref>
            ] is a typical supervised method. Unsupervised methods are easy
to use, while supervised methods are suitable for various situations in which unsupervised
methods do not perform well.
          </p>
          <p>
            Data fusion methods have been applied to various tasks in information retrieval [
            <xref ref-type="bibr" rid="ref11 ref12 ref13 ref14">11, 12, 13, 14</xref>
            ].
It is also popular for medical retrieval tasks [
            <xref ref-type="bibr" rid="ref32 ref33 ref34">32, 33, 34</xref>
            ]. Some form of data fusion techniques
are also used in those runs submitted to the 2020 TREC Health Misinformation Track, which
we use for the experiments. For example, both runs, CiTIUSCrdRelAdh and CiTIUSSimRelAdh,
submitted by the CiTIUS group [
            <xref ref-type="bibr" rid="ref35">35</xref>
            ], used Borda Count, to combine two types of rankings:
relevance and reliability (credibility &amp; correctness). For the h2oloo group [
            <xref ref-type="bibr" rid="ref36">36</xref>
            ], query expansion
and two types of machine learning technologies were used for re-ranking. All eight runs
submitted were various combinations of them and the BM25 baseline run, in which equal or
simple inequal weights were used. Similar situation exists in some other submissions.
          </p>
          <p>
            Usually, the number of component retrieval systems involved is a good indicator of the
complexity of a fusion-based system. With equal final performance, it is preferable to have
fewer component retrieval systems involved. [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ] investigated how to choose a subset from
a large group of retrieval systems for better fusion performance, although those retrieval
systems/results are not for medical retrieval tasks. A DCG-like (Discounted Cumulative Gain,
a commonly used metric in information retrieval evaluation) measure was defined for the
selection purpose. One limitation of this research is: it only considered performance of those
candidate systems, but not diversity of those systems (results) chosen. As a matter of fact, both
component system performance and dissimilarity among component systems (results) afect
fusion performance significantly.
          </p>
          <p>
            In this piece of work, we investigate how to achieve the best possible results by using the data
fusion technology for this misinformation retrieval task. We focus on the problem of subset
selection for efective fusion. The task is: for a group of  retrieval systems, how to select
 ( &lt;  ) of them to obtain the best fusion performance? This task is the same as that in
[
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. However, we propose a clustering-based method for this task, which is diferent from [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ].
Both theoretical analysis and empirical investigation demonstrate that our proposed method
is more efective than the one proposed in [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. Besides, [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ] used four data sets from CLEF
(Cross-Language Evaluation Forum) 2 for their empirical investigation.
          </p>
          <p>
            This piece of work is also diferent from those applied data fusion methods for medical
retrieval [
            <xref ref-type="bibr" rid="ref11 ref12 ref13 ref14 ref35 ref36">11, 12, 13, 14, 35, 36</xref>
            ]. All of them empirically investigated the efectiveness of a few
typical data fusion methods for the chosen task. Choosing a subset from a large group of
candidate systems is not a research task in those studies.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Subset Selection for Fusion</title>
      <p>
        For a group of information retrieval systems, how to select a subset for best possible fusion
efectiveness is a challenging task. For example, if we have 50 retrieval systems and try to select
10 of them for better fusion performance, then the number of possible combinations is huge.
As a matter of fact, the exact number to this question is 50*49*...*41, or 37,276,043,023,296,000.
Therefore, it may not be possible to test all of them. Instead of doing an exhaustive search to try
to find the best possible solution, to develop and use some heuristic methods is more realistic.
3.1. Top_J
In this vein, [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] defined a DCG-like measure, which is referred to as J-measure later in this
paper. It is defined as
      </p>
      <p>|| ()
 () = ∑︁ (1 − || ) * ()
=1
(1)</p>
      <sec id="sec-3-1">
        <title>2Its web site is located at https://www.clef-campaign.org/</title>
        <p>
          where  is a ranked list of documents for a given query, || is the number of documents in ,
1, 2, ..., || are documents in , and ()=1 if  is relevant to the query and ()=0
otherwise. For a group of resulting lists, J values can be used to evaluate and select component
results (and corresponding retrieval systems) for fusion. This selection method is referred to as
Top_J in this paper. Top_J is reasonable because it is found that better component systems/results
usually lead to better fusion performance [
          <xref ref-type="bibr" rid="ref37 ref38">37, 38</xref>
          ].
        </p>
        <sec id="sec-3-1-1">
          <title>3.2. Clustering-Based Subset Selection for Fusion</title>
          <p>
            Previous research [
            <xref ref-type="bibr" rid="ref38">38</xref>
            ] found that performance of component systems/results is not the only
factor that afects fusion performance. Diversity of the component retrieval systems/results
is also a factor that afects fusion performance significantly, but it is not considered in Top_J.
To incorporate diversity to the selection process, we propose clustering-based methods. There
are two major steps involved. First all component systems/results are set into clusters by
considering their similarity. Consequently, we can expect that the systems/results in the same
cluster are similar and the systems/results not in the same cluster are very diferent. The second
step is to choose a group of retrieval systems for fusion. In this step, we can take top performers
from diferent clusters, thus both performance of component systems (good performers in a
cluster) and diversity in the selected systems (chosen from diferent clusters) can be considered
in tandem.
          </p>
          <p>Now let us see how to perform the clustering method for those retrieval systems. We assume
that the properties of a retrieval system is fully reflected by the results it retrieves. For two
retrieval systems, we can observe the similarity/dissimilarity of the two ranked lists of results
they generate for the same query. Scoring is used in this work and we can define the Euclidean
distance to measure the dissimilarity of two resulting lists.</p>
          <p>||
(1, 2) = ∑︁ √︀(1() − 2())2
=1
(2)
where 1 and 2 are retrieved result lists from two retrieval systems for the same collection 
and same query , || is the number of documents in , 1() is the score that  obtains in
1, and 2() is the score that  obtains in 2. For all the documents in  that do not appear
in 1 (or 2), we need to define a default score (e.g., zero) for them. (1, 2) denotes the
distance between 1 and 2, which is a good indicator of the dissimilarity between 1 and 2.
Although not used here, ranking information is an alternative for the same purpose.</p>
          <p>For our investigation, K-means is a good option for clustering relatively a small number of
retrieval systems (e.g., the data set of Health Misinformation Track in TREC 2000 comprises 51
runs) and the Euclidean distance between them is well-defined for clustering. Most clustering
methods such as K-means requires a pre-defined value as the number of clusters. That value
needs to be considered carefully. When the number of clusters are very small, it is possible
that quite diferent results have to go to the same cluster. Therefore, such a situation should
be avoided even we just need a small number of component results for fusion. On the other
hand, if too many clusters are generated, then each cluster will become very small. Considering
that there are 51 runs in the data set used for the experiment, we decide to generate 17 clusters.
Thus each cluster has three result lists on average. It would give us some flexibility for the
selection of candidates. For the time being, we take a simple selection method: first we select
the best performer  (in MAP, or Mean Average Precision) in all the clusters. Then we removed
the cluster to which  belongs. For the remaining clusters repeat the above process until we get
enough result lists. In this way both performance of component result lists and their diversity
can be considered at the same time. This method is referred to as C1 later in this paper.</p>
          <p>
            The quality of clusters generated by K-means is determined by the initial  points, which
are chosen randomly. In order to improve the quality of clustering, we use a variant of K-means
presented in [
            <xref ref-type="bibr" rid="ref39">39</xref>
            ]. Its main idea is to generate  solutions by K-means. Then the best is chosen
from those  candidates. It is a little more complicated than standard K-means but usually
produce clusters in better quality. It is referred to as C2 later in this paper.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Settings and Results</title>
      <p>In this section we present the setting and results of the experiment carried out to validate the
proposed methods. Especially, the data set used is the ad-hoc task of the Health Misinformation
Track in TREC 2020, we would demonstrate the applicability of the proposed methods to this
special information seeking task.</p>
      <sec id="sec-4-1">
        <title>4.1. Experimental Settings</title>
        <p>
          In November 2020, TREC held a Health Information Track [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The track used the documents
found in the CommonCrawl News crawl from January 1, 2020 to April 30, 2020. The crawl
contains news articles from web sites all over the world.
        </p>
        <p>The topics (queries) for this track focused on the consumer health search domain relevant
to COVID-19. Fifty topics with a fixed structure were provided. All include number, title,
description, answer, evidence, and narrative. Fig. 1 gives an example of the topic used. The
title field has the form of a pair of treatment and disease. The description is formulated as a
question, which contains treatment, efect, and disease. The answer corresponds to the medical
consensus at the time of topic creation. Finally, the remaining fields were not intended to be
used by the retrieval systems, but only by human assessors to produce relevance judgment
document “qrels”.</p>
        <p>It set two tasks: total recall and ad-hoc retrieval. In this study, we use all 51 runs submitted
to the ad-hoc retrieval task by eight research groups. Their statistics are summarized in Table 1.</p>
        <p>Apart from C1 and C2, two baseline methods, Top_J and Top_MAP, are also tested. The
common ground of Top_J and Top_MAP is that both of them only consider performance of
component systems but not diversity of the selected systems. However, slightly diferent from
Top_J, Top_MAP chooses retrieval systems based on their MAP values.</p>
        <p>
          Two measures, MAP (Mean Average Precision) and CAM (Convex Aggregation Measure), are
used for retrieval results evaluation. MAP is a classical measure commonly used for efectiveness
evaluation of retrieval results, while CAM considers multiple aspects of a retrieved result list
[
          <xref ref-type="bibr" rid="ref40">40</xref>
          ]. It is defined as
 () = ()/3 + ()/3 + ()/3
(3)
where  is a ranked list of documents with multi-aspect labels, , , and  denote
respectively any valid relevance, correctness, and credibility evaluation measures. In this study,
we follow the instantiation of TREC by using nDCG for each individual aspect. That is to
calculate  as standard nDCG with respect to relevance,  as standard nDCG with
respect to correctness labels, and  as standard nDCG with respect to credibility.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Experimental Methodology and Results</title>
        <p>When the resulting lists are chosen from those clusters, we use CombSum, CombMNZ, and
linear combination to fuse them.</p>
        <p>
          For the same document collection  and a group of retrieval systems  for (1 ≤  ≤ ). All
retrieval systems  (1 ≤  ≤ ) search  for a given query  and each of them provides a
ranked list of documents  = &lt; 1, 2, ...,  &gt;. Assume that a relevance score ( ) is
associated with each of the retrieved documents in the list. CombSum [
          <xref ref-type="bibr" rid="ref29 ref41">29, 41</xref>
          ] uses the following
equation
        </p>
        <p>(5)

() = ∑︁ () (4)</p>
        <p>=1
to calculate scores for every document . Here () is the score that  assigns to . If  does
not appear in any , then a default score (e.g., 0) must be assigned to it. After that, every
document  obtains a global score () and all the documents can be ranked according to the
global scores they obtain.</p>
        <p>
          CombMNZ [
          <xref ref-type="bibr" rid="ref29 ref41">29, 41</xref>
          ] uses the equation

() =  * ∑︁ ()
=1
to calculate scores. Here  is the number of results in which document  appears.
        </p>
        <p>
          The linear combination method [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] uses the equation below
        </p>
        <p>() = ∑︁  * () (6)</p>
        <p>
          =1
to calculate scores.  is the weight assigned to system . Obviously, the linear combination
is a general form of CombSum. If all the weights  are equals to 1, then the linear combination
is the same as CombSum. Note that how to assign weights to diferent retrieval systems is an
important issue. We use multiple linear regression to train weights [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] for it.
        </p>
        <p>Let a training data set comprises a collection of  documents (), a group of  queries (),
and a group of  information retrieval systems (). For each query , all information retrieval
systems  (1 ≤  ≤ ) provide their estimated relevance scores to all the documents in the
collection. Therefore, we have (1 ,  , ) for  = (1, 2, ..., ),  = (1, 2, ..., ). Here 
2,..., 
stands for the score assigned by retrieval system  to document  for query ;  is the
judged relevance score of  for query . If binary relevance judgment is used, then it is 1 for
relevant documents and 0 otherwise.</p>
        <p>= {;  = (1, 2, ..., ),  = (1, 2, ..., )} can be estimated by a linear combination of
scores from all component systems. Consider the following quantity
ℱ = ∑︁ ∑︁ [ − ( ˆ0 +  ˆ11  +  ˆ22  + ... +  ˆ)]2</p>
        <p>=1 =1
when ℱ reaches its minimum, the estimation is the most accurate.  0,  1,  2,..., and  , the
multiple linear regression coeficients, are numerical constants that can be determined from
observed data.</p>
        <p>
          In the least squares sense the coeficients obtained by multiple linear regression can bring us
the optimum fusion results by the linear combination method, since they can be used to make
the most accurate estimation of the relevance scores of all the documents to all the queries as a
whole [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ].   can be used as weights for retrieval systems  (1 ≤  ≤ ) for fusion.
        </p>
        <p>
          Score normalization is a necessary step for fusing all the result lists. For any of the component
result lists, the retrieved documents are assigned scores using 1/(rank()+60), where rank() is
the ranking position of document . It is proposed in [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ].
        </p>
        <p>All the queries are divided into two groups: odd-numbered and even-numbered, then two-fold
cross-validation 3 is applied. There are uncertainty involved in C1 and C2. We run both of them
3N-fold cross-validation is a commonly used methodology in machine learning for training and testing a model
on the same dataset.
0.52
0.5
50 times. The results presented in this section are the average of them.</p>
        <p>In almost all the cases, CombMNZ is slightly worse than CombSum. Therefore, in the
following we do not present CombMNZ’s performance. Figs 2-5 present the performance of
CombSum and linear combination.</p>
        <p>From these figures, we can see that C1 and C2 are better than Top_J and Top_MAP in most
cases, whether CombSum or linear combination is used for fusion, and whether MAP or CAM
is used for evaluation. Very often C1 and C2 are close. It shows that using a simple or a more
sophisticated clustering method does not change fusion performance very much.</p>
        <p>In all 51 runs submitted, the best performer is h2oloo.m5, with a MAP of 0.3832 and a CAM
of 0.5883. Both C1 and C2 outperform it throughout from fusing 2 to 17 component systems.
Obviously, h2oloo.m5 is always a participant in both C1 and C2. If we consider the situation of
fusing 17 retrieval systems, then C1+CombSum achieves 0.4638 in MAP, and 0.7011 in CAM,
which are better than h2oloo.m5 by 21.03% and 19.17%, respectively; C2+linear combination
achieves 0.5036 in MAP, and 0.7161 in CAM, which are better than h2oloo.m5 by 31.42% and
21.72%, respectively. It is also noticeable that when fusing three to six systems, Top_MAP
achieves the best performance with CombSum. On the other hand, the advantage of C1 and
C2 is more prominent with linear combination throughout all diferent number of component
retrieval systems.</p>
        <p>The following Table 2 shows pairwise comparison results of all four methods on average
of 16 groups of fusion (2-17 resulting lists). For example, the figure at column “C1:C2” and
row “CombSum/MAP” means that 1.20% is the improvement rate of subset section methods C1
over C2 using CombSum for fusion and measured by MAP. A figure in bold indicates that the
diference between the two methods is significant at the .05 level (paired samples t test). From
Table 2 we can see that C1 and C2 are very close and better than the other two. C1 performs
better than C2 when fused with CombSum, while C2 performs better than C1 when linear
combination is used for fusion. Top_MAP is in the third place, while Top_J is the worst. The
diference between either C1 or C2 and Top_J is always over 5%, while the diference in other
situations is less than 5% apart from one case: C2 vs. Top_MAP fused by linear combination
and measured by MAP.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Clustering &amp; Subset Selection Analysis</title>
      <p>In this section, we present some further observations and some analysis about clustering-based
methods.</p>
      <sec id="sec-5-1">
        <title>5.1. Clustering Analysis</title>
        <p>First let us look at clustering. Fig. 6 shows a clustering example of K-means with five clusters:
all the resulting lists in a cluster are shown in the same colour, the distance between any two
resulting lists represents their dissimilarity, and the size of the font represents the performance
of the resulting list in MAP.</p>
        <p>It seems that K-means does reasonably well in this example. However, we may observe
that performance varies considerably across diferent clusters. As a matter of fact, the best in
ifve clusters are h2oloo_5 (0.3832), KU_10 (0.3640), KU_3 (0.3122), RSL_4 (0.1913), and NLM_8
(0.1111), respectively. Three of them are much higher than the other two. Such an observation
may be a positive evidence that generating more clusters is a good approach. If more clusters
are generated, we can avoid picking some really bad ones. In this example, if we only choose
three, then all selected runs are above 0.3 in MAP.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Subset Selection Analysis</title>
        <p>In Section 4, we evaluated and compared four subset selection methods. Now for all the selected
lists by each method, we calculate their average MAP and average pairwise distance between all
the selected lists. See Table 3 for the detailed information. Distance values reflect the diversity of
the chosen lists. From Table 3, we can see that Top_J and especially Top_MAP only choose those
lists with top MAP values. When more lists are chosen, the average performance of resulting
lists in all four methods decrease. However, the decrease in C1 and C2 is much more quickly
than it in Top_MAP and Top_J. On the other hand, higher distance values appear in all the cases
for both C1 and C2 while that values are always lower for Top_MAP and Top_J. This give us a
clear view of the four selection methods on two important aspects: performance and diversity.
Top_MAP and Top_J only concern performance, they always choose top performers, but with
less diversity, especially when a larger number of runs are selected. On the other hand, C1 &amp;
C2 have a balanced view about those two aspects. Compared with their counterparts Top_MAP
and Top_J, more often they choose those runs with smaller MAP values but larger distance
values on average. If we define a new measure =0.5*MAP/Max_MAP+0.5*Dist/Max_Dist,
where Max_MAP (0.377) and Max_Dist (4.629) are the maximal values observed, respectively,
then we can find that in most cases C1 and C2 have large  values than Top_MAP and
Top_J do in almost all the cases except two. It can explain why C1 &amp; C2 are more efective than
Top_MAP and Top_J in most cases and on average.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>In this paper, we have presented clustering-based methods for selecting a subset of component
retrieval systems from all available ones to achieve good fusion performance. Experiments
carried out with the Health Misinformation data set in TREC 2020 show that the proposed
methods are very good. When fusing up to 17 retrieval systems, the proposed methods are
better than the best component retrieval system by 20% to 30%, and they are also better than
the state-of-the-art subset selection method by a clear margin. One major characteristic of
the proposed methods is they take both performance of component systems and dissimilarity
among them into consideration at the same time. Such results demonstrate that data fusion is a
good approach for this Health Misinformation task.</p>
      <p>In our future work, we plan to further investigate the relationship between component system
performance and dissimilarity among component results. If a more precise relationship can be
set up for them, then it is possible to find more eficient and efective system selection methods
for fusion. Another direction is to design an unsupervised version of such methods. At present,
generating a usable training dataset can be very costly because relevance judgment by human
referees is required for those retrieved documents. If some automatic performance estimation
methods can be applied instead, then its usability can be improved.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Marcoux</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Agarwal</given-names>
            <surname>Nitin</surname>
          </string-name>
          ,
          <article-title>Narrative trends of covid-19 misinformation</article-title>
          .,
          <source>in: Text2Story@ ECIR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>77</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Chandrasekaran</given-names>
            <surname>Ranganathan</surname>
          </string-name>
          , Mehta Vikalp, Valkunde Tejali, Moustakas Evangelos, Topics, trends, and
          <article-title>sentiments of tweets about the covid-19 pandemic: Temporal infoveillance study</article-title>
          ,
          <source>Journal of medical Internet research 22</source>
          (
          <year>2020</year>
          )
          <article-title>e22624</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Abdelminaam</given-names>
            <surname>Diaa</surname>
          </string-name>
          <string-name>
            <given-names>Salama</given-names>
            , Ismail Fatma Helmy, Taha Mohamed, Taha Ahmed,
            <surname>Houssein Essam</surname>
          </string-name>
          <string-name>
            <given-names>H</given-names>
            ,
            <surname>Nabil</surname>
          </string-name>
          <string-name>
            <surname>Ayman</surname>
          </string-name>
          ,
          <article-title>Coaid-deep: An optimized intelligent framework for automated detecting covid-19 misleading information on twitter</article-title>
          ,
          <source>IEEE Access 9</source>
          (
          <year>2021</year>
          )
          <fpage>27840</fpage>
          -
          <lpage>27867</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Ng</given-names>
            <surname>Lynnette Hui Xian</surname>
          </string-name>
          , Loke Jia Yuan,
          <article-title>Analyzing public opinion and misinformation in a covid-19 telegram group chat</article-title>
          ,
          <source>IEEE Internet Computing</source>
          <volume>25</volume>
          (
          <year>2020</year>
          )
          <fpage>84</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Leng</given-names>
            <surname>Yan</surname>
          </string-name>
          , Zhai Yujia, Sun Shaojing, Wu Yifei, Selzer
          <string-name>
            <surname>Jordan</surname>
          </string-name>
          , Strover Sharon, Zhang Hezhao, Chen Anfan, Ding Ying,
          <article-title>Misinformation during the covid-19 outbreak in china: Cultural, social and political entanglements</article-title>
          ,
          <source>IEEE Transactions on Big Data</source>
          <volume>7</volume>
          (
          <year>2021</year>
          )
          <fpage>69</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Zhou</surname>
            <given-names>Cheng</given-names>
          </string-name>
          , Xiu Haoxin,
          <string-name>
            <surname>Wang</surname>
            <given-names>Yuqiu</given-names>
          </string-name>
          ,
          <article-title>Yu Xinyao, Characterizing the dissemination of misinformation on social media in health emergencies: An empirical study based on covid-19,</article-title>
          <source>Information Processing &amp; Management</source>
          <volume>58</volume>
          (
          <year>2021</year>
          )
          <fpage>102554</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Shahi</given-names>
            <surname>Gautam</surname>
          </string-name>
          <string-name>
            <given-names>Kishore</given-names>
            , Dirkson Anne,
            <surname>Majchrzak Tim</surname>
          </string-name>
          <string-name>
            <surname>A</surname>
          </string-name>
          ,
          <article-title>An exploratory study of covid-19 misinformation on twitter</article-title>
          ,
          <source>Online social networks and media 22</source>
          (
          <year>2021</year>
          )
          <fpage>100104</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Choudrie</given-names>
            <surname>Jyoti</surname>
          </string-name>
          , Banerjee Snehasish, Kotecha Ketan, Walambe Rahee, Karende Hema, Ameta Juhi,
          <article-title>Machine learning techniques and older adults processing of online information and misinformation: a covid 19 study</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>119</volume>
          (
          <year>2021</year>
          )
          <fpage>106716</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Roberts</given-names>
            <surname>Kirk</surname>
          </string-name>
          , Alam Tasmeer, Bedrick Steven,
          <string-name>
            <surname>Demner-Fushman</surname>
            <given-names>Dina</given-names>
          </string-name>
          , Lo Kyle, Soborof Ian, Voorhees Ellen, Wang Lucy Lu,
          <string-name>
            <surname>Hersh William</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <article-title>Trec-covid: rationale and structure of an information retrieval shared task for covid-</article-title>
          19
          <source>, Journal of the American Medical Informatics Association</source>
          <volume>27</volume>
          (
          <year>2020</year>
          )
          <fpage>1431</fpage>
          -
          <lpage>1436</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Charles</surname>
            <given-names>L.A.</given-names>
          </string-name>
          <string-name>
            <surname>Clarke</surname>
            , Maria Maistro, Saira Rizvi,
            <given-names>Mark D</given-names>
          </string-name>
          .
          <article-title>Smuckerand Guido Zuccon, Overview of the TREC 2020 health misinformation track</article-title>
          ,
          <source>in: Proceedings of the TwentyNinth Text REtrieval Conference</source>
          , TREC
          <year>2020</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Budíková</surname>
            <given-names>Petra</given-names>
          </string-name>
          , Batko Michal, Zezula Pavel,
          <article-title>Fusion strategies for large-scale multi-modal image retrieval, in: Transactions on Large-Scale Data-</article-title>
          and
          <string-name>
            <surname>Knowledge-Centered Systems</surname>
            <given-names>XXXIII</given-names>
          </string-name>
          , Springer,
          <year>2017</year>
          , pp.
          <fpage>146</fpage>
          -
          <lpage>184</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Kato</surname>
            <given-names>Sosuke</given-names>
          </string-name>
          , Shimizu Toru, Fujita Sumio, Sakai Tetsuya,
          <article-title>Unsupervised answer retrieval with data fusion for community question answering</article-title>
          ,
          <source>in: Asia Information Retrieval Symposium</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Roostaee</surname>
            <given-names>Meysam</given-names>
          </string-name>
          , Sadreddini Mohammad Hadi, Fakhrahmad Seyed Mostafa,
          <article-title>An efective approach to candidate retrieval for cross-language plagiarism detection: A fusion of conceptual and keyword-based schemes</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>57</volume>
          (
          <year>2020</year>
          )
          <fpage>102150</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Smeaton Alan</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Connor Edel</surname>
          </string-name>
          , Regan Fiona,
          <article-title>Multimedia information retrieval and environmental monitoring: Shared perspectives on data fusion</article-title>
          ,
          <source>Ecological informatics 23</source>
          (
          <year>2014</year>
          )
          <fpage>118</fpage>
          -
          <lpage>125</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Juárez-González</surname>
            <given-names>Antonio</given-names>
          </string-name>
          , Montes-y-Gómez
          <string-name>
            <surname>Manuel</surname>
          </string-name>
          ,
          <string-name>
            <surname>Villaseñor-Pineda</surname>
            <given-names>Luis</given-names>
          </string-name>
          , PintoAvendaño David,
          <string-name>
            <surname>Pérez-Coutiño</surname>
            <given-names>Manuel</given-names>
          </string-name>
          ,
          <article-title>Selecting the n-top retrieval result lists for an efective data fusion</article-title>
          ,
          <source>in: International Conference on Intelligent Text Processing and Computational Linguistics</source>
          , Springer,
          <year>2010</year>
          , pp.
          <fpage>580</fpage>
          -
          <lpage>589</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Zhao</surname>
            <given-names>Yuehua</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Da</surname>
            <given-names>Jingwei</given-names>
          </string-name>
          , Yan Jiaqi,
          <article-title>Detecting health misinformation in online health communities: Incorporating behavioral features into machine learning based approaches</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>58</volume>
          (
          <year>2021</year>
          )
          <fpage>102390</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Wang</surname>
            <given-names>Zuhui</given-names>
          </string-name>
          , Yin Zhaozheng, Argyris Young Anna,
          <article-title>Detecting medical misinformation on social media using multimodal deep learning</article-title>
          ,
          <source>IEEE Journal of Biomedical and Health Informatics</source>
          <volume>25</volume>
          (
          <year>2020</year>
          )
          <fpage>2193</fpage>
          -
          <lpage>2203</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Zhang</surname>
            <given-names>Qiang</given-names>
          </string-name>
          , Cook Jonathan, Yilmaz Emine,
          <article-title>Detecting and forecasting misinformation via temporal and geometric propagation patterns</article-title>
          .,
          <source>in: ECIR (2)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>455</fpage>
          -
          <lpage>462</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Lee</surname>
            <given-names>Nayeon</given-names>
          </string-name>
          , Li Belinda
          <string-name>
            <given-names>Z</given-names>
            ,
            <surname>Wang</surname>
          </string-name>
          <string-name>
            <surname>Sinong</surname>
          </string-name>
          , Fung Pascale, Ma Hao, Yih Wen-tau, Khabsa Madian,
          <article-title>On unifying misinformation detection</article-title>
          ,
          <source>arXiv preprint arXiv:2104.05243</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Ginsca Alexandru</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popescu</surname>
            <given-names>Adrian</given-names>
          </string-name>
          , Lupu Mihai, Credibility in information retrieval,
          <source>Foundations and Trends in Information Retrieval</source>
          <volume>9</volume>
          (
          <year>2015</year>
          )
          <fpage>355</fpage>
          -
          <lpage>475</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Fogg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tseng</surname>
          </string-name>
          ,
          <article-title>The elements of computer credibility</article-title>
          , in: M. G. Williams,
          <string-name>
            <given-names>M. W.</given-names>
            <surname>Altom</surname>
          </string-name>
          (Eds.),
          <source>Proceeding of the CHI '99 Conference on Human Factors in Computing Systems: The CHI is the Limit</source>
          , Pittsburgh, PA, USA, May
          <volume>15</volume>
          -20,
          <year>1999</year>
          , ACM,
          <year>1999</year>
          , pp.
          <fpage>80</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Tseng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Fogg</surname>
          </string-name>
          , Credibility and computing technology,
          <source>Commun. ACM</source>
          <volume>42</volume>
          (
          <year>1999</year>
          )
          <fpage>39</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>M. J. Metzger</surname>
            ,
            <given-names>A. J.</given-names>
          </string-name>
          <string-name>
            <surname>Flanagin</surname>
          </string-name>
          ,
          <article-title>Information in online environments: the use of cognitive heuristics</article-title>
          ,
          <source>Journal of Pragmatics</source>
          <volume>59</volume>
          (
          <year>2013</year>
          )
          <fpage>210</fpage>
          -
          <lpage>220</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          , G. Pasi,
          <article-title>Credibility in social media: opinions, news, and health information - a survey, WIREs Data Mining Knowl</article-title>
          .
          <source>Discov</source>
          .
          <volume>7</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Song</surname>
            <given-names>Shijie</given-names>
          </string-name>
          , Zhang Yan, Yu Bei,
          <article-title>Interventions to support consumer evaluation of online health information credibility: A scoping review</article-title>
          ,
          <source>International Journal of Medical Informatics</source>
          <volume>145</volume>
          (
          <year>2021</year>
          )
          <fpage>104321</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Wu</surname>
            <given-names>Shu</given-names>
          </string-name>
          , Liu Qiang, Liu Yong, Wang Liang,
          <string-name>
            <surname>Tan</surname>
            <given-names>Tieniu</given-names>
          </string-name>
          ,
          <article-title>Information credibility evaluation on social media</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>30</volume>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alrubaian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Al-Qurishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. M.</given-names>
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alamri</surname>
          </string-name>
          ,
          <article-title>A credibility analysis system for assessing information on twitter</article-title>
          ,
          <source>IEEE Transactions on Dependable and Secure Computing</source>
          <volume>15</volume>
          (
          <year>2016</year>
          )
          <fpage>661</fpage>
          -
          <lpage>674</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Wu</surname>
            <given-names>Lianwei</given-names>
          </string-name>
          , Rao Yuan, Nazir Ambreen, Jin Haolin,
          <article-title>Discovering diferential features: Adversarial learning for information credibility evaluation</article-title>
          ,
          <source>Information Sciences 516</source>
          (
          <year>2020</year>
          )
          <fpage>453</fpage>
          -
          <lpage>473</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Fox</given-names>
            <surname>Edward</surname>
          </string-name>
          <string-name>
            <given-names>A</given-names>
            ,
            <surname>Koushik</surname>
          </string-name>
          <string-name>
            <surname>M Prabhakar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Shaw</given-names>
            <surname>Joseph</surname>
          </string-name>
          , Modlin Russell,
          <string-name>
            <given-names>Rao</given-names>
            <surname>Durgesh</surname>
          </string-name>
          , et al.,
          <article-title>Combining evidence from multiple searches</article-title>
          ,
          <source>in: The first text retrieval conference (TREC1)</source>
          ,
          <year>1993</year>
          , pp.
          <fpage>319</fpage>
          -
          <lpage>328</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Cormack Gordon</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clarke Charles</surname>
            <given-names>LA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buettcher</surname>
            <given-names>Stefan</given-names>
          </string-name>
          ,
          <article-title>Reciprocal rank fusion outperforms condorcet and individual rank learning methods</article-title>
          ,
          <source>in: Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>758</fpage>
          -
          <lpage>759</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Wu</surname>
            <given-names>Shengli</given-names>
          </string-name>
          ,
          <article-title>Linear combination of component results in information retrieval</article-title>
          ,
          <source>Data &amp; Knowledge Engineering</source>
          <volume>71</volume>
          (
          <year>2012</year>
          )
          <fpage>114</fpage>
          -
          <lpage>126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Clipa</surname>
            <given-names>Teofan</given-names>
          </string-name>
          ,
          <article-title>Di Nunzio Giorgio Maria, A study on ranking fusion approaches for the retrieval of medical publications</article-title>
          ,
          <source>Information</source>
          <volume>11</volume>
          (
          <year>2020</year>
          )
          <fpage>103</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>de Herrera Alba G Seco</surname>
          </string-name>
          ,
          <article-title>Schaer Roger</article-title>
          , Markonis Dimitrios, Müller Henning,
          <article-title>Comparing fusion techniques for the imageclef 2013 medical case retrieval task</article-title>
          ,
          <source>Computerized Medical Imaging and Graphics</source>
          <volume>39</volume>
          (
          <year>2015</year>
          )
          <fpage>46</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Mourão</surname>
            <given-names>André</given-names>
          </string-name>
          , Martins Flávio, Magalhaes Joao,
          <article-title>Multimodal medical information retrieval with unsupervised rank fusion</article-title>
          ,
          <source>Computerized Medical Imaging and Graphics</source>
          <volume>39</volume>
          (
          <year>2015</year>
          )
          <fpage>35</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Fernández-Pichel</surname>
            <given-names>Marcos</given-names>
          </string-name>
          , Losada David E,
          <string-name>
            <surname>Pichel Juan</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elsweiler</surname>
            <given-names>David</given-names>
          </string-name>
          ,
          <source>Citius at the trec 2020 health misinformation track</source>
          <volume>1266</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Pradeep</surname>
            <given-names>Ronak</given-names>
          </string-name>
          , Ma Xueguang, Zhang Xinyu, Cui Hang, Xu Ruizhou, Nogueira Rodrigo,
          <string-name>
            <surname>Lin</surname>
            <given-names>Jimmy</given-names>
          </string-name>
          , H2oloo at trec 2020:
          <article-title>When all you got is a hammer... deep learning, health misinformation, and precision medicine</article-title>
          ,
          <source>Corpus</source>
          <volume>5</volume>
          (
          <year>2020</year>
          )
          <article-title>d2</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Leonard</surname>
            <given-names>David</given-names>
          </string-name>
          , Lillis David, Zhang Lusheng, Toolan Fergus,
          <string-name>
            <surname>Collier Rem</surname>
            <given-names>W</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunnion</surname>
            <given-names>John,</given-names>
          </string-name>
          <article-title>Applying machine learning diversity metrics to data fusion in information retrieval</article-title>
          ,
          <source>in: European Conference on Information Retrieval</source>
          , Springer,
          <year>2011</year>
          , pp.
          <fpage>695</fpage>
          -
          <lpage>698</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <surname>Wu</surname>
            <given-names>Shengli</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McClean Sally</surname>
          </string-name>
          ,
          <article-title>Performance prediction of data fusion for information retrieval</article-title>
          ,
          <source>Information processing &amp; management 42</source>
          (
          <year>2006</year>
          )
          <fpage>899</fpage>
          -
          <lpage>915</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>Bradley</surname>
            <given-names>Paul S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fayyad Usama</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <article-title>Refining initial points for k-means clustering</article-title>
          ., in: ICML, volume
          <volume>98</volume>
          ,
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <year>1998</year>
          , pp.
          <fpage>91</fpage>
          -
          <lpage>99</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <surname>Mustafa</surname>
            <given-names>Abualsaud</given-names>
          </string-name>
          , Christina Lioma, Maria Maistro,
          <string-name>
            <given-names>Mark D.</given-names>
            <surname>Smucker</surname>
          </string-name>
          , Guido Zuccon,
          <article-title>Overview of the trec 2019 decision track</article-title>
          ,
          <source>in: TREC</source>
          , volume
          <volume>1250</volume>
          ,
          <string-name>
            <surname>Special</surname>
            <given-names>Publication</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>Fox</given-names>
            <surname>Edward</surname>
          </string-name>
          <string-name>
            <given-names>A</given-names>
            ,
            <surname>Shaw Joseph</surname>
          </string-name>
          <string-name>
            <surname>A</surname>
          </string-name>
          ,
          <article-title>Combination of multiple searches</article-title>
          ,
          <source>NIST special publication SP</source>
          <volume>243</volume>
          (
          <year>1994</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>