<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Methods for Text-Image-Rematching using Pair-wise Similarity and Canonical Similarity Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kani Abdul</string-name>
          <email>kani.abdul@campus.tu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kiran Kiran</string-name>
          <email>k.kiran@campus.tu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Max Rudat</string-name>
          <email>rudat@campus.tu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexandros Vasileiou</string-name>
          <email>a.vasileiou@campus.tu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Lommatzsch</string-name>
          <email>andreas.lommatzsch@campus.tu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Technische Universität Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Matching images to text plays an important role in cross-media retrieval and research has proven this to be an underestimated challenge. This problem is addressed by the MediaEval 2021 NewsImages Challenge with the goal to gain more insights into the real-world relationship of news articles and images. We develop models for re-establishing the connection of a news article to its corresponding image using datasets of a German news publisher (“task 1”). Our approaches follow the idea of pairwise similarity learning and are optimized by algorithmic hill climbing. Additionally, we employ Canonical Correlation Analysis as an approach using joint embedding learning. The evaluation shows that our approaches produce good results for the underlying image-text rematching task, yet require further optimization to yield stable prediction performance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Multimedia content is accompanying our everyday life. News
articles are one form of multimedia that are characterized by textual
content accompanied by imagery. The assumption that a simple
relationship underlies this connection has frequently turned out
to be oversimplified in research [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The MediaEval 2021
NewsImages task aims at addressing this challenge by investigating the
real-world relationship of news and images. The challenge provides
a dataset consisting of three batches training data and one batch
for the evaluation. The performance of the participants’ algorithms
is evaluated on a test batch. The evaluation metrics Recall@ and
Mean Reciprocal Rank are used [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The challenge of image-text retrieval has been addressed broadly
in research around multimedia analysis [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Deep image-text
matching serves as one frequently used approach for this scenario. Zhang
and Lu [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] classify the main approaches based on deep learning
into two categories: pairwise similarity learning and joint
embedding learning. For pairwise similarity learning, the main idea is to
learn a similarity network for predicting the score of image-text
pairs [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. As for the other category of joint embedding learning, a
joint latent space is defined in that the vectors of texts and images
can be compared directly. The typically used learning methods
belonging to this category are canonical correlation analysis (CCA)
and bi-directional ranking loss [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>Based on the existing methods, we develop two text-image
matching strategies optimized for the specific requirements of the
NewsImages task. In Sec. 2 we explain the preprocessing and the steps of our
strategies in detail. Sec. 3 presented the evaluation results. Finally,
the overall findings are discussed in Sec. 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>We develop two approaches addressing the text-image rematching
task. This section discusses the steps of our approaches.</p>
      <p>
        Data Preprocessing. We preprocess the provided data for
eficiently computing similarity scores. Firstly, we translate the image
labels (computed by VGG-19 trained on ImageNet) from English to
German using Google Translate [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We decided to translate the
image labels instead of the article snippets due to the smaller volume
to translate. In addition, we enhanced the dataset by extracting
the item category (e.g. ‘koeln’, ‘panorama’ ‘wirtschaft’ ‘politik’)
from the article URL. In the next step, we normalize the dataset
by removing stop words, punctuation marks, spaces, special
characters, and digits. After removing the above tokens, we employ
part-of-speech (POS) tagging to identify the nouns in the dataset.
Finally, we perform Morphological Processing (“lemmatization”) for
creating the dataset.
      </p>
      <p>
        In addition to the standard preprocessing, we integrate
Open-deWordNet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] enabling us to consider synonyms when computing
the similarity score. Moreover, we implement an outlier detection
and removal strategy based on the Z-score for the terms derived
from the textual description. If a word’s z-score is larger than 3.0 (3
standard deviations away from the mean), then it is considered an
outlier and gets removed from the dataset. Adding these two derives
ifelds to our dataset, enables us to research, whether additional
preprocessing improves the performance.
      </p>
      <p>
        Pairwise similarity learning &amp; algorithmic hill climbing. Our first
approach follows the idea of pairwise similarity learning. The
objective is to compute a similarity score for each image-article pair
to be used to construct 1-to-1 matches. We implement this using
spaCy similarity from the natural language processing (nlp)
module SpaCy [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. SpaCy ofers two methods to find the similarity
between words: one based on context-sensitive tensors and another
one based on word vectors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We utilize the later method and
generate a similarity matrix containing the similarity scores for
each image-text pair. Our pre-processed dataset gives diferent
options for setting the input for computing the similarity scores. For
example, we tested to include only the words within the article
text or only the words within the article title. Analogously for the
images, we have data from 10 diferent labeler configurations, each
generating at least 2 image labels with diferent label probabilities.
      </p>
      <p>
        For computing the best parameter configuration, we make use
of algorithmic hill climbing. Starting with an (arbitrary) initial
conifguration, the solution is incrementally adapted [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We implement
this by first initializing the parameter (e.g. set labelprobability to 0.0,
number of considered labels per image to nine and the remaining
parameters to False). Then we go iteratively through the parameter
space and compute for each parameter the (locally) optimal value.
For the optimization of each parameter, we randomly select 1,000
samples.
      </p>
      <p>The parameter configurations are evaluated using matrices
containing pairwise similarity scores with the columns representing
the image IDs, the rows the article IDs and the correct matches
being located along the diagonal of the matrix. The score is computed
using the variables Row counter, Column counter, and Total counter.</p>
      <p>For each generated parameter configuration, the performance is
evaluated as follows: if the similarity score for the diagonal value
is higher than the other scores within its row and column, then
the index of the total counter is increased by 1. If this condition is
not met, then it is checked whether the diagonal value is higher
than the other scores within its row and the row counter index
increased by 1 dependently. If both conditions are not met, the
column counter index is increased by 1. Once the values of all
the 3 counting variables are set, the performance is calculated by
comparing the values of the row and column counters: if the row
counter is larger than the column counter, then the row counter
is divided by the number of pairs (n) and returned. Otherwise the
column value divided by the number of pairs (n) is returned as the
ifnal score and evaluated by our hill climbing algorithm.</p>
      <p>Canonical Correlation Analysis. As a second approach, we apply
the sklearn Canonical Correlation Analysis. On the preprocessed
dataset, the spaCy implementation of Word2Vec is used for
computing a vector representation (having 300 dimensions) of the dataset.
We use a random sample of size 1,500 data points, utilizing the
article text and image labels columns. Then we split this set into
train set (2/3) and validation set (1/3). Initial tests of CCA showed a
poor performance; that is why, we adapted the method. We applied
Kernel PCA (kPCA) to transform the data through a Radial Basis
Function (RBF) expansion, limited to utmost 700 generated data
dimensions. The kPCA transformation is applied on both the train
and test data to keep compatible and comparable dimensions. Then,
we train a CCA instance on the train data set and evaluate the model
on both the train and validation set. The evaluation is performed
based on the predicted vector for each of the article text vectors.
Since the CCA-predicted vectors are most likely not corresponding
to actual word2vec image label transformations, we compare each
CCA prediction to all of the word2vec image label vectors using
the cosine similarity measure, thus constructing a similarity matrix.
This allows us to make 1-1 article texts to image label mappings.
3</p>
    </sec>
    <sec id="sec-3">
      <title>EVALUATION</title>
      <p>The evaluation (by the task organizers) shows that out pairwise
similarity-based approach reaches  @100 = 0.21, the
CCAbased approach reaches  @100 = 0.24. Even though CCA
outperforms the pairwise similarity-based approach in general, the
recall score for k=5 and k=10 is higher with hill-climbing,
indicating that low-level semantics are found more efectively using the
straightforward method of pairwise similarity learning. The better
evaluation scores for higher  observed for CCA show that CCA
performs better with regard to high-level semantic similarity.</p>
      <p>We analyze the parameter settings for maximizing the
performance for the pairwise similarity learning approach. We find
that the used configuration considers for each image the 8 labels
with the highest score (no label probability score has been applied).
This indicates that a detailed image description is crucial for the
text-image rematching task. Furthermore, we find that considering
the title in addition to the article snippet does not improve the
performance. Moreover, our analysis shows that the replacement
of words with the first word of their synsets as well as the removal
of outliers and duplicates are not activated for the final evaluation.</p>
      <p>Analyzing the parameters used by CCA we find, that considering
lemmatized article texts and image labels does not yield the optimal
results. Using Kernel PCA with a RBF kernel, the R2 score, which
indicates how well the regression model fits the observed data, was
greatly increased for the train set (achieving values up to 99.8%
with k=100). The performance on the test set reached a R2 score
of 38.5% (k=100); thus this method outperformed the hill-climbing
optimized pairwise similarity learning approach, that reached an
R2 score of 32.6%.
4</p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSION</title>
      <p>The evaluation shows, that both approaches yield robust results
for the image-text re-matching. The results slightly outperform
the best results from MediaEval NewsImages 2020. The CCA-based
approach reaches a recall@100 score of 23.6% on the evaluation set;
the hill climbing approach based on pairwise similarity learning
yields a recall@100 score of 20.6%. Due to limited resources, we have
tested only a restricted set of parameter configurations; we think
that a further parameter optimization will improve the performance.
For our CCA model, we observe a performance diference between
training and testing set, indicating the presence of overfitting. The
overfitting could be tackled by applying regularization or adding a
dropout layer (eliminating features with a low impact). Furthermore,
a penalty-based component could be used for boosting articles
considering the margin-based error. The ranking of the top-k images
for an article could then be optimized significantly.</p>
      <p>Furthermore, an alternative image labeling component should be
considered to getting a more detailed image description that could
be matched with the article text. This is based on the observation
that we observed a better performance when considering more
images labels.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Honnibal</surname>
          </string-name>
          , Ines Montani, Sofie Van Landeghem,
          <string-name>
            <given-names>and Adriane</given-names>
            <surname>Boyd</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>spaCy: Industrial-strength Natural Language Processing in Python</article-title>
          . (
          <year>2020</year>
          ). https://doi.org/10.5281/zenodo.1212303
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Kille</surname>
          </string-name>
          , Andreas Lommatzsch, and
          <string-name>
            <given-names>Özlem</given-names>
            <surname>Özgöbek</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>NewsImages: The role of images in online news</article-title>
          .
          <source>In Proceedings of the MediaEval Benchmarking Initiative for Multimedia Evaluation</source>
          <year>2020</year>
          . CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2882</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Kille</surname>
          </string-name>
          , Andreas Lommatzsch, Özlem Özgöbek, Mehdi Elahi, and
          <string-name>
            <surname>Duc-Tien</surname>
          </string-name>
          Dang-Nguyen.
          <year>2021</year>
          .
          <article-title>News Images in MediaEval 2021</article-title>
          .
          <source>In Proceedings of the MediaEval Benchmarking Initiative for Multimedia Evaluation</source>
          <year>2021</year>
          . CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2882</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Fouad</given-names>
            <surname>Omran</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Treude</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Choosing an NLP Library for Analyzing Software Documentation: A Systematic Literature Review and a Series of Experiments</article-title>
          . (05
          <year>2017</year>
          ). https://doi.org/10.1109/ MSR.
          <year>2017</year>
          .42
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Dimitris</given-names>
            <surname>Papadias</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Hill Climbing Algorithms for Content-Based Retrieval of Similar Configurations</article-title>
          .
          <source>In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery</source>
          ,
          <fpage>240</fpage>
          -
          <lpage>247</lpage>
          . https://doi.org/10.1145/345508.345587
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Melanie</given-names>
            <surname>Siegel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Francis</given-names>
            <surname>Bond</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>OdeNet: Compiling a German Wordnet from other Resources</article-title>
          .
          <source>In Proceedings of the 11th Global Wordnet Conference (GWC</source>
          <year>2021</year>
          ).
          <fpage>192</fpage>
          -
          <lpage>198</lpage>
          . https://www.aclweb.org/ anthology/2021.gwc-
          <volume>1</volume>
          .
          <fpage>22</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Google</given-names>
            <surname>Translator</surname>
          </string-name>
          .
          <year>2020</year>
          . (
          <year>2020</year>
          ). https://pypi.org/project/googletrans/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Xing</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>Yang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            <given-names>Zuo</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fumin Shen</surname>
          </string-name>
          , and Heng Tao Shen.
          <year>2020</year>
          .
          <article-title>Cross-Modal Attention With Semantic Consistence for Image-Text Matching</article-title>
          .
          <source>IEEE Transactions on Neural Networks and Learning Systems</source>
          <volume>31</volume>
          ,
          <issue>12</issue>
          (
          <year>2020</year>
          ),
          <fpage>5412</fpage>
          -
          <lpage>5425</lpage>
          . https://doi.org/10.1109/ TNNLS.
          <year>2020</year>
          .2967597
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Ying</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Huchuan</given-names>
            <surname>Lu</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep Cross-Modal Projection Learning for Image-Text Matching</article-title>
          . Springer International Publishing, Cham,
          <fpage>707</fpage>
          -
          <lpage>723</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>