<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DL-TXST NewsImages: Contextual Feature Enrichment for Image-Text Rematching</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuxiao Zhou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andres Gonzalez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Parisa Tabassum</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jelena Tešić</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department, Texas State University</institution>
          ,
          <addr-line>San Marcos, TX</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>In this paper, we describe our multiview approach to news image rematching to text for the news article run submission. The feature pool consists of provided features, baseline text and image features using pre-trained and domain-adapted modeling and contextual features for the news and image article. We have evaluated multiple modeling approaches for the features and employed a deep multilevel encoding network to predict a probability-like matching score of images for a news article. Our best results are the ensemble of proposed models, and we found that the URL for the image and related images provides the most discriminative context in this pairing task. Online news articles are multimodal; the textual content of an article is often accompanied by an image. The image is important to illustrate the content of the text and also to attract readers' attention. Existing research generally assumes a simple relationship between images and text, e.g. image captioning is often assumed to be a brief textual description of an image. On the contrary, when images accompany news articles, the relationship becomes less clear. In this research, we employ a state-of-the-art method that builds models to describe the connection between the textual content of articles and the images that accompany them. We evaluated our proposed model on the benchmark data set derived from four months from web server log files of a German news publisher. The performance of the proposed model is measured by image matching precision such as MRR and Mean Recall at different depths.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Recent work utilizes deep neural networks to capture the
visualsemantic similarity between image and text. Wang et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and
Faghri et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] map the image and the entire sentence to a
common vector space and compute the similarity between the
global representations. The fine-tuned version of the approach
uses a range of embedded information in news and images, for
example, extracted named entities and image features [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] or the
caption of the news image with named entities [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Semantic
concept learning [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and regional relationship reasoning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
approaches were shown to improve the discriminative ability of
unified embeddings.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>MATCHERS</title>
      <p>
        In this section we introduce the matchers we have tested for
image-text rematching task, as illustrated in
3.1 The Semantic Space Matcher matches text and image
embeddings in the semantic space using the cosine distance.
Previous work emphasized matching an image to the category
using the URL by which the image was downloaded [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. We
streamline the approach and fix the semantic space ahead. We
refine the classification layers of ResNet50 to produce probability
outputs for a fixed semantic space for 70 classes, creating a
70dimensional array. We combine the title and the text of the
article, normalize it, and feed it to a text classifier that produces a
probability that the text will describe one of the 70 semantic
classes. The probability of the category is a feature vector value
for that category. The result is two sets: one containing the text
feature vectors, one per text instance, and another one containing
the image feature vector, one per instance, in the same feature
space. Next, we match an image to the input text based on the
minimal cosine distance between the said text feature vector and
all image features.
3.2 The Face-Name Matcher correlates the names within the the
articles with the faces within images using 128-dimensional image
space embedding [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The Stanford Named Entity Recognizer
(NER) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] provided a named entity recognizer particularly for the
extraction of person names: of the 7530 given news articles in the
corpora, 24% of them included the person's name. We use open
source face detection FaceNet to connect the person's name from
the article to publicly available images and to create a
128dimensional face vector [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We use Google DeepFace to detect
the faces in the images, and encodes the detected faces to the
same space [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The image is matched to the article based on a
minimum cosine distance between the vectors for 24 % of the
articles that contain the actual names. For articles that do not
contain person names, image captioning is utilized.
3.3 The Image Captioning Matcher Based on the hypothesis that
the description of a new image is semantically like the matched
news title, we first adopted an image captioning model [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
pretrained with COCO dataset for image caption generation, and then
calculate the similarity score between the generated image
captions and the given news headlines. The pre-trained image
captioning model has three main components: 1. Image Feature
Extractor: the image caption model uses ResNet101, a
convolutional neural network (CNN) that is 101 layers deep for
feature extraction; 2. Transformer encoder: the extracted image
features are then passed to a Transformer-based encoder that
generates a new representation of the inputs;
3. Transformer Decoder: this component takes the encoder output
and the text data sequence as inputs and tries to learn to generate
the caption; 4. Text Similarity : we employ Word Mover's
Distance (WMD) to compare the similarity between image
captions and article titles. The WMD algorithm uses normalized
Bag-of-Words and word embeddings to calculate the distance
between documents and sentences. The wmdSimilarity is simply
the negative wmd between the image caption and the title.
3.4 The Filename Token Matcher The filenames of the images
extracted from the URLs and the filenames of the articles
extracted from the URLs encode semantic connection (as the
names are likely crafted by humans). We propose tokenizing the
filenames, and discover image-text match based on the number of
overlapping tokens, as illustrated in Figure 1.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4 RESULTS AND ANALYSIS</title>
      <p>Data The MediaEval 2021 Image-Text Re-Matching benchmark
provides four batches of data, which consist of the headline and a
text snippet of German news articles and their accompanying
images. The first three batches are used for training and the last
one is used for testing. We split the training data set into the actual
training set and the validation set. The training set included 5135
records, while the validation set included 2384 records. The
findings of the training data are shown in Figure 2. Filename
Token Matcher produces the best overall results on the training
and validation dataset. Based on our findings, we have submitted
3 runs:
Run1 combines three different methods. Equal weights are
assigned to the categorization-based method and a combination of
face-name matching and image captioning-based methods. The
ranking of a candidate image in Run1 is as follows:
 Run1 =0.5 Categorization +0.5 ( Face+ caption)</p>
      <p>Figure 3 Matcher’s MRR@100 during Training
Run2 combines all proposed methods. The first three models are
assembled using the same approach as in Run1. This ensemble
model is used to create the initial top 100 image list. Then we
append the result, which is generated from the filename token
matcher, to the end of the top 100 image list.</p>
      <p>Run3 is like Run2. The only difference is that we append the
result of the last method to the head of the top 100 image list.
Result Our proposed approach uses an ensemble design, so our
submissions are combined results from three or four models.
5</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS</title>
      <p>The filename token matching method recovered almost 50% of
the ground truth (R@100) in the test set, as outlined in Table 1.
This is consistent with the findings of the training phase described
in Figure 3. This experiment demonstrates the depiction gap in
automated image-text correspondence. Human reasoning depicts
the same piece of information in different modalities to
compliment, not duplicate, the presented information. Semantic
image-text connections are unconscious imprinted by humans in
the filenames of images and articles.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Diao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L.Ma, and
          <string-name>
            <given-names>H.</given-names>
            <surname>Lu</surname>
          </string-name>
          .
          <article-title>Similarity Reasoning and Filtration for Image-Text Matching</article-title>
          ,
          <fpage>36</fpage>
          -
          <lpage>44</lpage>
          . In the Thirty-
          <source>Fifth AAAI Conference on Artificial Intelligence (AAAI-21)</source>
          ,
          <year>2021</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Wang</surname>
            , Liwei,
            <given-names>Yin</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>and Svetlana</given-names>
          </string-name>
          <string-name>
            <surname>Lazebnik</surname>
          </string-name>
          .
          <article-title>"Learning deep structure-preserving image-text embeddings."</article-title>
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , pp.
          <fpage>5005</fpage>
          -
          <lpage>5013</lpage>
          .
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Faghri</surname>
            , Fartash,
            <given-names>David J.</given-names>
          </string-name>
          <string-name>
            <surname>Fleet</surname>
            , Jamie Ryan Kiros, and
            <given-names>Sanja</given-names>
          </string-name>
          <string-name>
            <surname>Fidler</surname>
          </string-name>
          .
          <article-title>"Vse++: Improving visual-semantic embeddings with hard negatives</article-title>
          .
          <source>" arXiv preprint arXiv:1707.05612</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Huang</surname>
            , Yan, Qi Wu, Chunfeng Song, and
            <given-names>Liang</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>"Learning semantic concepts and order for image and sentence matching."</article-title>
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pp.
          <fpage>6163</fpage>
          -
          <lpage>6171</lpage>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Kunpeng</given-names>
          </string-name>
          , Yulun Zhang, Kai Li,
          <string-name>
            <given-names>Yuanyuan</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Yun</given-names>
            <surname>Fu</surname>
          </string-name>
          .
          <article-title>"Visual semantic reasoning for image-text matching."</article-title>
          <source>In Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          , pp.
          <fpage>4654</fpage>
          -
          <lpage>4662</lpage>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Nguyen-Quang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>TDH</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen-Ho</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duong</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoang-Xuan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen-Truong</surname>
            ,
            <given-names>V. T.</given-names>
          </string-name>
          , ... &amp;
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          (
          <year>2020</year>
          ). HCMUS at MediaEval 2020:
          <article-title>Image-Text Fusion for Automatic News-Images Re-Matching.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yumeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jing</surname>
          </string-name>
          , G. Shuo, and
          <string-name>
            <given-names>L.</given-names>
            <surname>Limin</surname>
          </string-name>
          ,
          <article-title>"News Image-Text Matching With News Knowledge Graph,"</article-title>
          <source>in IEEE Access</source>
          , vol.
          <volume>9</volume>
          , pp.
          <fpage>108017</fpage>
          -
          <lpage>108027</lpage>
          ,
          <year>2021</year>
          , doi: 10.1109/ACCESS.
          <year>2021</year>
          .
          <volume>3093650</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <fpage>2021</fpage>
          ,
          <article-title>Stanford Named Entity Recognizer (NER)</article-title>
          . https://nlp.stanford.edu/software/CRF-NER.html
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9] 2021,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Taigman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ranzato</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <article-title>"DeepFace: Closing the Gap to Human-Level Performance in Face Verification,"</article-title>
          <source>2014 IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1701</fpage>
          -
          <lpage>1708</lpage>
          , doi: 10.1109/CVPR.
          <year>2014</year>
          .
          <volume>220</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <year>2021</year>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Luo</surname>
          </string-name>
          , G. Shakhnarovich,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cohen</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <article-title>"Discriminability Objective for Training Descriptive Captions,"</article-title>
          <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>6964</fpage>
          -
          <lpage>6974</lpage>
          , doi: 10.1109/CVPR.
          <year>2018</year>
          .
          <volume>00728</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Schroff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kalenichenko</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Philbin</surname>
          </string-name>
          ,
          <article-title>"FaceNet: A unified embedding for face recognition and clustering,"</article-title>
          <source>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>815</fpage>
          -
          <lpage>823</lpage>
          , doi: 10.1109/CVPR.
          <year>2015</year>
          .
          <volume>7298682</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>