<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>M. G. Constantin); mihai.dogariu@upb.ro
(M. Dogariu); bogdan.ionescu@upb.ro (B. Ionescu)
 https://lstefan.aimultimedialab.ro/ (L. Ştefan)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Overview of ImageCLEFfusion 2023 Task - Testing Ensembling Methods in Diverse Scenarios</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Liviu-Daniel Ştefan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihai Gabriel Constantin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihai Dogariu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bogdan Ionescu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AI Multimedia Lab, Politehnica University of Bucharest</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This paper presents a comprehensive overview of the second edition of the ImageCLEFfusion task, held in 2023. The primary goal of this endeavor is to facilitate the advancement of late fusion or ensembling methodologies, which possess the capability to leverage prediction outcomes derived from pre-computed inducers to generate superior and enhanced prediction outputs. The present iteration of this task encompasses three distinct challenges: the continuation of the previous year's regression challenge utilizing media interestingness data, where performance is measured via the mAP at 10 metric; the continuation of the retrieval challenge involving image search result diversification data, where performance is measured via the F1-score and Cluster Recall at 20; and the addition of a new multi-label classification task focused on concepts detection in medical data, where performance is measured via the F1-score. Participants were provided with a predetermined set of pre-computed inducers and were strictly prohibited from incorporating external inducers during the competition. This ensured a fair and standardized playing field for all participants. A total of 23 runs were received and the analysis of the proposed methods shows diversity among them ranging from machine learning approaches that join the inducer predictions to ensemble schemes that learn the results of other methods.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Late fusion</kwd>
        <kwd>Ensembling</kwd>
        <kwd>Fusion benchmarking</kwd>
        <kwd>Visual interestingness prediction</kwd>
        <kwd>Image search results diversification</kwd>
        <kwd>Caption detection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The fusion task, part of ImageCLEF [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], was first proposed in 2022 [ 3] comprising of two
subtasks: a regression challenge utilizing media interestingness data (ImageCLEFfusion-int) and
a retrieval challenge involving image search result diversification data (ImageCLEFfusion-div).
In 2023 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], both subtasks, ImageCLEFfusion-int and ImageCLEFfusion-div, were running again
with the addition of a multi-label classification task focused on concepts detection in medical
data (ImageCLEFfusion-cap). These type of tasks typically exhibit inferior performance in
endto-end systems when juxtaposed with conventional computer vision tasks. This phenomenon is
frequently ascribed to their intrinsic subjectivity and multi-modality, compounded by challenges
associated with establishing dependable ground-truth annotations [4, 5]. To address these
limitations, researchers have turned to late fusion or ensembling systems as a primary approach
to enhance model performance. These systems involve the integration of multiple individual
prediction systems, referred to as inducers, through fusion schemes.
      </p>
      <p>Given these factors, the participants in this task are faced with several challenges that
necessitate exploration. These challenges include diversity, which pertains to a collection of
classifiers that generate varying predictions for the same instance; voting mechanism, which
governs the utilization of individual outputs from the base models during prediction; dependency,
which refers to the influence of a base model on the construction of the subsequent model in the
fusion chain; cardinality, which denotes the number of individual base models composing the
ensemble—a delicate balance must be struck, as incorporating too many models may diminish
diversity within the fusion; and finally, the learning mode of the base models, which represents
the characteristic that enables the classifiers to efectively adapt to new, previously unseen data
while retaining previously acquired knowledge.</p>
      <p>This paper presents an overview of the 2023 ImageCLEFfusion task including the data creation
in Section 2, the evaluation methodology in Section 3, and the task and participation in Section
4. The results are described in Section 5, followed by conclusion in Sections 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data description</title>
      <p>The ImageCLEFfusion framework encompasses three distinct tasks, each utilizing diferent
datasets and associated challenges:</p>
      <p>ImageCLEFfusion–int: This task focuses on the Interestingness10k dataset [5].
Specifically, it utilizes image-based prediction data derived from the 2017 MediaEval Predicting
Media Interestingness task [6]. The task provides prediction outputs from 29 systems that
participated in the benchmarking task. To facilitate the training and evaluation of fusion
systems, the available data is divided into 1,877 samples for training and 558 samples for
testing.</p>
      <p>ImageCLEFfusion–div: This task relies on the Retrieving Diverse Social Images
dataset [7], specifically targeting the DIV150Multi challenge [ 8]. The task provides
retrieval outputs from 56 systems, which are further divided into 60 queries for the
training data and 63 queries for the testing data.</p>
      <p>ImageCLEFfusion–cap: This task is derived from the ImageCLEF Medical Caption
Task [9]. It involves the extraction of multi-label outputs from 84 inducers. The data used
for this task consists of 6,101 images for the development set and 1,500 images for the
testing set.</p>
      <p>For the training sets, we provide a comprehensive package comprising ground truth data,
inducer prediction outputs, detailed inducer performance metrics, and the requisite scripts for
metric computation. Conversely, the testing sets solely include the inducer prediction outputs.
The characteristics of the datasets used in these tasks are presented in Table 1. Participants have
the freedom to generate their own validation sets by partitioning the training set according
to their specific needs. However, to ensure a fair and reasonable selection of proposed fusion
methods, participants are limited to a maximum of 10 runs for each of the three tasks.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Evaluation Methodology</title>
      <p>Participants were required to devise late fusion learning strategies based on the outputs of the
inducers associated with the media samples for each of the subtasks. The evaluation of the
participants’ submissions was conducted using the Mean Average Precision at 10 (mAP@10) metric
for the ImageCLEFfusion–int task, F1 at 20 (F1@20) and Cluster Recall at 20 (Cluster Recall@20)
metrics for the ImageCLEFfusion-div task, and the F1 metric for the ImageCLEFfusion–cap task.
The aforementioned metrics align with the evaluation measures employed for the individual
datasets pertaining to each of the three tasks. Participants were encouraged to submit their
solutions for all three tasks.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Participation</title>
      <p>A total of 12 teams completed their registration for ImageCLEFfusion, demonstrating a strong
interest in the competition. Among these teams, two successfully submitted their runs and
completed the competition by submitting detailed working notes describing their methods.
In terms of the interestingness task, both teams collectively submitted 13 runs, while one
team submitted a total of 10 runs for the diversification task. No runs were recorded for the
ImageCLEFfusion–cap task. For a comprehensive overview of the participating teams, please
refer to Table 2.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>Institutions
Computer Science Department, Morgan State University,
Baltimore, Maryland, US
Sri Sivasubramaniya Nadar College of Engineering,
Chennai, Tamil Nadu, India
Runs
int
10
3</p>
      <p>Runs Runs
div cap
10
0
0
0</p>
      <p>Paper
yes
yes</p>
      <sec id="sec-5-1">
        <title>5.1. ImageCLEFfusion-int task</title>
        <p>A total of 13 runs were submitted by two teams for the ImageCLEFfusion-int task. The highest
achieved performance in terms of MAP@10 value was 0.1331, indicating a significant
improvement of 40.65% compared to the baseline value of 0.0946. Despite the reduced number of
participants compared to the previous year, the participating team achieved a performance
that surpassed the majority of the participants in the previous year, but still under the
state-ofthe-art result of the last year achieved by [12]. The results for the participating teams for the
ImageCLEFfusion-int task are presented in Tables 3.</p>
        <p>SSN CSE-ML: The SSN CSE-ML team’s most successful run attained a mAP@10 score of
0.1331, establishing itself as the highest-scoring submission among the participating teams
in this year’s competition for this particular subtask. Balasundaram et al. [11] utilized an
ensemble learning model approach based on a Voting Classifier that leverages XGBoost,
decision trees, and K-nearest neighbors algorithms, and uses grid search to find the best
hyperparameters for each classifier, and the optimal voting scheme and weights for the
Voting Classifier.</p>
        <p>CS_Morgan: The best performing run from the CS_Morgan team achieved a mAP@10 of
0.1287. For this approach, Emon and Rahman [10] utilized an ensemble of decision
trees trained sequentially, with each tree using the predictions from the previous tree
to calculate residual errors. A shrinkage technique is applied to reduce the ensemble’s
impact after each tree’s prediction. The ensemble’s final predictions are obtained by
averaging the regression results. Additionally, the predictions undergo scaling through
min-max normalization.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. ImageCLEFfusion-div task</title>
        <p>A single team submitted a total of 10 runs for the ImageCLEFfusion-div task. The highest
performance achieved by the team resulted in an F1@20 score of 0.5708, indicating a 7.4%
improvement compared to the baseline value of 0.5313. Additionally, for the secondary metric
CR@20, the corresponding system exhibited an improvement of 8.45%. The results for the
participating teams in the ImageCLEFfusion-div task can be found in Tables 4.</p>
        <p>SSN CSE-ML: The best performing run from the SSN CSE-ML team achieved an F1@20
score of 0.5708 and an CR@20 score of 0.449 using the same model construction as for
the ImageCLEFfusion-int task, i.e., bulding an ensemble model based on three classifier
models (XGBoost, decision tree, and K-nearest neighbors), and finally creating a Voting
Classifier based on the best combination of voting scheme and weights obtained through
a grid search.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>The second edition of the ImageCLEFfusion task garnered submissions from a total of two
teams. The participants were presented with three tasks: the continuation of the previous
year’s regression challenge, which involved media interestingness data; the continuation of the
retrieval challenge, which focused on image search result diversification data; and the addition
of a new multi-label classification task centered around concepts detection in medical data.
In total, the teams submitted 23 runs, with 13 runs for media interestingness and 10 runs for
diversification. Unfortunately, no runs were recorded for the concept detection task. Despite the
reduced number of participants compared to the previous year, with only two teams submitting
runs for two out of the three presented tasks, the participating teams achieved commendable
performance that surpassed the majority of the participants from the previous year. However,
their performance still fell short of the state-of-the-art result achieved in the previous year.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Acknowledgments</title>
      <p>This work is supported under the H2020 AI4Media “A European Excellence Centre for Media,
Society and Democracy” project, contract #951911.
media, and recommender systems applications, in: Experimental IR Meets
Multilinguality, Multimodality, and Interaction, Proceedings of the 14th International Conference of
the CLEF Association (CLEF 2023), Springer Lecture Notes in Computer Science (LNCS),
Thessaloniki, Greece, 2023.
[3] L.-D. Ştefan, M. G. Constantin, M. Dogariu, B. Ionescu, Overview of imageclefusion 2022
task-ensembling methods for media interestingness prediction and result diversification,
in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS. org, Bologna,
Italy, 2022.
[4] M. G. Constantin, L. D. Stefan, B. Ionescu, C.-H. Demarty, M. Sjoberg, M. Schedl, G. Gravier,
Afect in multimedia: benchmarking violent scenes detection, IEEE Transactions on
Afective Computing (2020).
[5] M. G. Constantin, L.-D. Ştefan, B. Ionescu, N. Q. Duong, C.-H. Demarty, M. Sjöberg, Visual
interestingness prediction: a benchmark framework and literature review, International
Journal of Computer Vision 129 (2021) 1526–1550.
[6] C.-H. Demarty, M. Sjöberg, B. Ionescu, T.-T. Do, M. Gygli, N. Duong, Mediaeval 2017
predicting media interestingness task, in: MediaEval workshop, 2017.
[7] B. Ionescu, M. Rohm, B. Boteanu, A. L. Gînscă, M. Lupu, H. Müller, Benchmarking image
retrieval diversification techniques for social media, IEEE Transactions on Multimedia 23
(2020) 677–691.
[8] B. Ionescu, A. L. Gînscă, B. Boteanu, M. Lupu, A. Popescu, H. Müller, Div150multi: a social
image retrieval result diversification dataset with multi-topic queries, in: Proceedings of
the 7th international conference on multimedia systems, 2016, pp. 1–6.
[9] J. Rückert, A. Ben Abacha, A. García Seco de Herrera, L. Bloch, R. Brüngel, A.
IdrissiYaghir, H. Schäfer, H. Müller, C. M. Friedrich, Overview of ImageCLEFmedical 2022 –
Caption Prediction and Concept Detection, in: CLEF2022 Working Notes, CEUR Workshop
Proceedings, CEUR-WS.org, Bologna, Italy, 2022.
[10] I. S. Emon, M. Rahman, Media interestingness prediction in imageclefusion 2023 with dense
architecture-based ensemble &amp; scaled gradient boosting regressor model, in: Experimental
IR Meets Multilinguality, Multimodality, and Interaction, CEUR Workshop Proceedings,
CEUR-WS.org, Thessaloniki, Greece, 2023.
[11] B. Prabavathy, G. G. Sai, N. Kishore, M. Olirva, A. M. Vaibhav, N. S. Murali, P. S. Harshith,
Eficient fusion techniques for result diversification and image interestingness tasks, in:
Experimental IR Meets Multilinguality, Multimodality, and Interaction, CEUR Workshop
Proceedings, CEUR-WS.org, Thessaloniki, Greece, 2023.
[12] M. G. Constantin, L.-D. Ştefan, M. Dogariu, B. Ionescu, Ai multimedia lab at imageclefusion
2022: Deepfusion methods for ensembling in diverse scenarios, in: CLEF2022 Working
Notes, CEUR Workshop Proceedings, CEUR-WS. org, Bologna, Italy, 2022.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Péteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bloch</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Brüngel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Idrissi-Yaghir</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schäfer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kozlovski</surname>
            ,
            <given-names>Y. D.</given-names>
          </string-name>
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Kovalev</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-D. Ştefan</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Deshayes-Chossart</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2022: Multimedia Retrieval in Medical, Social Media and Nature Applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 13th International Conference of the CLEF Association (CLEF</source>
          <year>2022</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.-M. Drăgulinescu</surname>
          </string-name>
          , W. wai
          <string-name>
            <surname>Yim</surname>
            ,
            <given-names>A. B.</given-names>
          </string-name>
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Snider</surname>
            , G. Adams,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Yetisgen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bloch</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Brüngel</surname>
            , A. IdrissiYaghir, H. Schäfer,
            <given-names>S. A.</given-names>
          </string-name>
          <string-name>
            <surname>Hicks</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Thambawita</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Storås</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Papachrysos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Schöler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>A.-G.</given-names>
          </string-name>
          <string-name>
            <surname>Andrei</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radzhabov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Coman</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Kovalev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Stan</surname>
            , G. Ioannidis,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Manguinhas</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-D. Ştefan</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Deshayes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
          </string-name>
          , Overview of ImageCLEF 2023:
          <article-title>Multimedia retrieval in medical, social</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>