<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Machine Learning Research</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>MeaningCloud at TASS 2018: News Headlines Categorization for Brand Safety Assessment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Javier Herrera-Planells</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Villena-Román MeaningCloud LLC</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>jherrera</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>jvillena}@meaningcloud.com</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>2172</volume>
      <fpage>807</fpage>
      <lpage>814</lpage>
      <abstract>
        <p>This paper describes the participation of MeaningCloud in Task 4 at TASS 2018 (Martínez-Cámara et al., 2018), which is focused on Brand Safety assessment. The objective of systems is to predict whether ads should be hidden for specific news articles, depending on the topics covered and potential negative emotions that could be triggered. Based on the output of our APIs for lemmatization, topics extraction and sentiment analytics, different Natural Language Understanding techniques combined with Machine Learning were tested in our experiments. The experiment that achieved the best result consisted of a Deep Learning algorithm based on Word Embeddings and CNN, trained with features based on the headlines, plus entity extraction, and topic and sentiment analysis.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In the online advertising context, Brand Safety
refers to practices and tools allowing to ensure
that an ad will not appear in a context that could
affect negatively or directly damage the
advertiser’s brand.</p>
      <p>An article may be considered unsafe for
advertising if it triggers negative feelings in the
reader. The creation of a system that detects
these cases faces some challenges. First,
different feelings might be triggered in each
reader, depending on their view on topics like
religion, economy, or sports. In addition,
combinations of pseudo-thematic classifications
and sentiment analysis are involved. For
example, a reduction of traffic accidents implies
a negative feeling because of the mention to car
accidents, but the reduction in number actually
represents good news.</p>
      <p>This paper describes the participation of
MeaningCloud in the Task 4 Good Or Bad News
of TASS 2018 workshop (Martínez-Cámara et
al., 2018), where prediction models have been
built for the categorization of news articles
headlines into two categories (Safe and Unsafe).</p>
      <p>In this task, lexical diversity among Spanish and
Latin American newspapers is also considered.</p>
      <p>Copyright © 2018 by the paper's authors. Copying permitted for private and academic purposes.</p>
      <p>Two subtasks had to be fulfilled for this Task
4. The first subtask required and evaluated a
training algorithm which was fed with a corpus
of 1500 headlines in Spanish from various
countries. It was afterwards tested against two
sets of 500 and 15000 headlines in Spanish from
various countries too. The second subtask
evaluated the generalization capacity of the
algorithm between Spanish from Spain and
Spanish from diverse American countries: the
training was fed with 250 headlines from
newspapers of Spain and tested against 400
headlines from newspapers of Latin America.</p>
      <p>The tagged corpus provided was quite well
balanced between training, development and test
sets with respect to country representation
(number of instances), although slightly
unbalanced with a higher number of samples in
the Unsafe category (64%).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Our Approach</title>
      <p>Our approach is composed by two steps: first,
multiple features are extracted from each
headline. Then, each feature vector is fed to a
machine learning model which finally produces
the Safe/Unsafe prediction.
2.1</p>
      <sec id="sec-2-1">
        <title>Feature Generation</title>
        <p>Features are extracted using the following public
APIs in our text analytics platform: entity
extraction and anonymization, lemmatization,
and sentiment and topic detection,
2.1.1</p>
        <p>Entity Extraction and Anonymization
Entities are detected using the MeaningCloud
topics extraction API. This service has been used
with raw headlines as the following one:
Vídeo muestra cómo Daesh mata a 4
soldados en Níger.</p>
        <p>For this headline, the entities Daesh and
Níger are detected along with their classes:
Organization&gt;TerroristOrganization and
Location&gt;Country.</p>
        <p>With this information, anonymized versions
of the headlines are generated to abstract from
references to actual entities that could bias the
analysis. Numbers are masked too. For instance:
Vídeo muestra
#TerroristOrganization#
soldados en #Location#.
mata
cómo
a 0</p>
        <p>In addition to this anonymized version of the
headline, the training algorithms were also fed,
separately, with the detected entities (Daesh and
Níger).
2.1.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Text Lemmatization</title>
        <p>Headlines are lemmatized after the previous
anonymization. The MeaningCloud
lemmatization, PoS and parsing API has been
used for this task. For the previous example, the
lemmatized form is:
vídeo mostrar cómo terroristorganization
matar a 0 soldado en location.
2.1.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Sentiment Detection</title>
        <p>Sentiment is detected using the MeaningCloud
sentiment analysis API. For instance, given the
following raw headline:</p>
        <p>Animales mueren en zoológico
Venezuela por falta de comida.
de</p>
        <p>For this headline, a global sentiment N+ is
detected. This score belongs to a scale [P+, P,
Neutral, N, N+] grading from positive to
negative, which we map to a range 0 (P+) to 4
(N+).</p>
        <p>The API also provides token-level sentiment
(morir with sentiment N+, por falta de comida
with a sentiment N) and subjectivity data (in this
case, OBJECTIVE), features which have been
omitted in this task.
2.1.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Topic Detection</title>
        <p>News topics (thematic categorization) can be
detected either with the MeaningCloud text
classification API, using one of the predefined
models such as IPTC for news categorization or
IAB for advertising market, or aggregating the
thematic information returned by the topics
extraction API for each detected entity. For this
task, we used this second approach. For instance,
for the following raw headline:</p>
        <p>En plena distensión por los Juegos
Olímpicos, Kim Jong-Un invitó al
presidente de Corea del Sur a Pyongyang.</p>
        <p>The topics extraction API detects four
entities: Juegos Olímpicos (Event), Kim Jong-un
(Person), Corea del Sur and Pyongyang (both
Location). Two of them provide thematic
information: Juegos Olímpicos belongs to sports
and Kim Jong-un belongs to politics. So finally
sports and politics are selected as topics.
2.2 Classifiers for Monolingual
Classification (subtask 1)
2.2.1</p>
      </sec>
      <sec id="sec-2-5">
        <title>Run 1: Machine Learning</title>
        <p>The starting point of our first experiment was a
training set of headlines that were anonymized
and lemmatized.</p>
        <p>We performed an iteration over all of them,
creating a list of n-grams that are frequently
found in the Unsafe category. Each n-gram was
assigned with a higher score if it appeared more
frequently in Unsafe than in Safe
headlines. Some of the top-ranked Unsafe
n-grams in the list after the training process were
morir, denunciar, asesinar, caso de corrupción
and the placeholder for the anonymized entity
terroristorganization. These n-grams are then
used to generate the features for each headline.</p>
        <p>One headline is represented by the following
feature vector:
•
•
•
•</p>
        <p>The sum of the scores of Unsafe
unigrams found in the headline, weighted
to the length of the headline in words.</p>
        <p>Scores for Unsafe bigrams, trigrams and
4-grams separately, in the same way as
unigrams.</p>
        <p>The sentiment score (0 to 4), extracted as
described in section 2.1.3.</p>
        <p>Features for each of the most frequent
topics, such as sports, politics or religion,
extracted as described in section 2.1.4.</p>
        <p>The most informative features, as shown by
the Extra-Trees algorithm (Geurts et al., 2006),
were the following: unigram scores (43%),
bigram scores (21%), sentiment scores (19%),
trigram scores (6%), 4-gram scores (2%),
politics topics (2%), football topics (1%) and
economy topics (1%).</p>
        <p>
          Then, several machine learning algorithms
have been tested for making predictions on these
feature vectors, including: KNN
          <xref ref-type="bibr" rid="ref1">(Altman, 1992)</xref>
          ,
random forests
          <xref ref-type="bibr" rid="ref2">(Breiman, 2001)</xref>
          , multilayer
perceptron, logistic regression, SVM (Vapnik et
al., 1995), XGBoost
          <xref ref-type="bibr" rid="ref3">(Chen et al., 2016)</xref>
          , and
AdaBoost
          <xref ref-type="bibr" rid="ref5">(Freund et al., 2003)</xref>
          .
        </p>
        <p>The accuracy of the different algorithms was
evaluated using development set. XGBoost was
finally chosen as the top-performant algorithm
for this experiment. The final settings were a tree
booster, learning rate of 0.1, minimum child
weight of 1 and maximum depth of 3.</p>
        <p>The resulting experiment for this task, trained
with 1500 samples and tested with 500 samples
(L1 corpus), had the following performance:
71.4% accuracy, 71.7% Macro-F1, 71.3%
Macro-Precision, and 72.2% Macro-Recall.</p>
        <p>The confusion matrix is shown in Table 1. As
it can be observed, the main reason for the errors
is the incorrect prediction of Unsafe news as
Safe, accounting for 19% of total errors and 32%
of the errors in the Unsafe category.</p>
        <p>Actual</p>
        <p>Safe</p>
        <p>Safe
Unsafe
Unsafe
This experiment follows the same principle as
the first one: we generate a n-gram score list in
our training step, and feature vectors use these
scores along the sentiment and topic statistics.</p>
        <p>This second experiment extends the n-gram
score features by using extra n-gram lists. These
additional lists are generated using the
nonanonymized version of the headlines. Some of
the top scoring resulting n-grams are FARC,
Jones Huala or caso de Edu Saettone.</p>
        <p>Using this information, which is derived from
non-anonymized entities, becomes helpful when
categorizing headlines within a similar period.</p>
        <p>This approach resulted in an increase of
performance over the previous experiment:
73.2% accuracy, 72.5% Macro-F1, 72.3%
Macro-Precision, and 72.7% Macro-Recall.</p>
        <p>The confusion matrix is shown in Table 2.
The detection of Unsafe category has noticeably
improved, though the accuracy of Safe category
has decreased.
This experiment, opposite to the previous ones,
feeds the machine learning algorithm with a set
of words/tokens. A deep learning model based
on word embeddings and a convolutional neural
network is then used for making predictions.</p>
        <p>The following headline will be used for
describing the process used in this experiment:
Al menos 25 civiles muertos deja ataque
contra el Daesh en Siria.</p>
        <p>The same features as described in the
previous experiment are generated: preprocessed
text, sentiment, topics and entities. All of them
are encoded as a set of tokens:
al menos 0 civil muerto dejar ataque
contra el terroristorganization en
location entdaesh entsiria sentiment3
topicpolitics</p>
        <p>The information contained in these tokens is
the following:</p>
        <p>Anonymized and lemmatized text: al
menos 0 civil muerto dejar ataque contra
el terroristorganization en location.</p>
        <p>Non anonymized entities: entdaesh,
entsiria.</p>
        <p>Sentiment: sentiment3, meaning a
sentiment with score 3 (negative, N).</p>
        <p>Topics found: topicpolitics.</p>
        <p>
          Afterwards, a deep learning model developed
using the Keras framework
          <xref ref-type="bibr" rid="ref4">(Chollet et al., 2015)</xref>
          is trained on this set of tokens. The
implementation of the model for this experiment
has the following settings:
        </p>
        <p>Input sequences with length of 23 words
(two times 11.5, the average word
length), padded for shorter texts with
PAD placeholders at the end.</p>
        <p>Embedding generation: 300-dimensional
embeddings for the 2241 most frequent
words (two thirds of the total 3362
different words). UNK placeholder for
words out of selected vocabulary.</p>
        <p>
          Convolutional neural network,
calculating a convolution with 3 different
region sizes
          <xref ref-type="bibr" rid="ref6">(Zhang and Wallace, 2015)</xref>
          and 2 filters for each region size. Kernel
size of {3,4,5}x300, ReLU activation
function (Nair and Hinton, 2010) and
max-pooling strategy.
        </p>
        <p>Two final densely-connected layers with
a dropout of 0.25 (Srivastava et al.,
2014). The second layer acts as the
output layer using a Softmax function.
•
•
•
•
•
•
•
•</p>
        <p>The resulting model, trained with 1500
samples and tested with 500 samples, showed a
high increase in performance: 77.6% accuracy,
76.7% Macro-F1, 76.7% Macro-Precision, and
76.7% Macro-Recall.</p>
        <p>Table 3 again shows the confusion matrix.
Actual Predicted</p>
        <p>Safe Safe</p>
        <p>Safe Unsafe
Unsafe Unsafe
Unsafe Safe</p>
        <p>Count
145 (29% all, 72% Safe)
56 (11% all, 28% Safe)
243 (49% all, 81% Unsafe)
56 (11% all, 19% Unsafe)
There is an improvement in both classes from
the previous experiments, most noticeable in the
Unsafe category. This confusion matrix is the
best among the previous ones if we take into
account the risk of considering an Unsafe article
as Safe. In this case, ads would be shown by
mistake.
2.2.4</p>
      </sec>
      <sec id="sec-2-6">
        <title>Overall Results</title>
        <p>Table 4 shows the overall results for subtask 1 of
our three experiments, sorted by Macro-F1
which is the comparison metric among
participants.</p>
        <p>Run Id
Run 1
Run 2
Run 3</p>
        <p>Macro-F1
0.717
0.725
0.767</p>
        <p>Next table 5 shows the final ranking in terms
of Macro-F1 for the best run by all participants,
sorted by Macro-F1. Our best experiment ranked
4th among 7 participants.</p>
        <p>Group
INGEOTEC
ELiRF-UPV
rbnUGR
MEANINGCLOUD
SINAI
lone_wolf
TNT-UA-WFU
Finally, the results over the L2 corpus
(including 13 152 headlines), tagged by pooling
submissions and based on the vote of majority,
are shown in Table 6. Our best experiment
ranked 4th again among all participants. The
improvement of results with respect to the other
corpus (L1) may be not real because of the
pooling (the decision of the majority may be
wrong anyway).</p>
        <p>Group
ELiRF-UPV
rbnUGR
INGEOTEC
MEANINGCLOUD
SINAI
TNT-UA-WFU
2.3 Classifiers for Multilingual
Classification (subtask 2)
Run 3 was, apparently, the top performant
among the other experiments in subtask 1, so it
was selected for subtask 2. The model was
trained with 250 headlines from newspapers of
Spain, and tested against 408 headlines from
newspapers of America.</p>
        <p>The results were the following: 65.8%
accuracy, 65.1% Macro-F1, 64.7%
MacroPrecision, and 65.4% Macro-Recall.</p>
        <p>The confusion matrix is shown in Table 7.
Obviously, results are considerably worse than
in the first task, as the information available for
training is extremely reduced.</p>
        <p>Actual</p>
        <p>Safe</p>
        <p>Safe
Unsafe
Unsafe</p>
        <p>Predicted</p>
        <p>Safe
Unsafe
Unsafe</p>
        <p>Safe</p>
        <p>Count
99 (24% all, 63% Safe)
57 (14% all, 37% Safe)
169 (42% all, 67% Unsafe)
84 (20% all, 33% Unsafe)
Finally, Table 8 shows the ranking in terms
of Macro-F1 in subtask 2 for the best run by all
participants, sorted by Macro-F1. Our best
experiment ranked 4th among 5 participants.</p>
        <p>Results in this subtask are, as expected, lower
than for subtask 1, for all teams. The ranking
among teams stays the same, so, apparently, the
lack of information (or the lack of generalization
of the models) affects the same to all groups.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>In this paper we described a system for detecting
headlines of news articles that might be unsafe
for advertising. We have incorporated different
preprocessing techniques, such as text
lemmatization, entity extraction and
anonymization, topic detection and sentiment
analysis.</p>
      <p>Then, we have evaluated several
classification algorithms, from n-gram scoring to
embeddings and deep learning models.</p>
      <p>Three techniques were found to significantly
improve the accuracy of the model: providing
both the anonymized text and the
nonanonymized entities separately, include the
sentiment pre-detection and the use of a deep
learning approach for the model training.</p>
    </sec>
    <sec id="sec-4">
      <title>Disclaimer</title>
      <p>MeaningCloud is one of the co-organizers of
TASS since the first edition in 2012, and,
specifically this year, of Task 4 Good Or Bad
News. Our participation in this task has been
completely blind, without making use of any
information or dataset not provided to the rest of
the participants.</p>
      <p>We are also sponsoring TASS 2018 with
prizes for the best teams. Obviously, as insiders,
we were never eligible for the prize, should our
experiments had been the top-performant.
GitHub.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Altman</surname>
            ,
            <given-names>N. S.</given-names>
          </string-name>
          <year>1992</year>
          .
          <article-title>An Introduction to Kernel and Nearest-Neighbor Non-Parametric Regression</article-title>
          .
          <source>The American Statistician</source>
          ,
          <volume>46</volume>
          (
          <issue>3</issue>
          ),
          <fpage>175</fpage>
          -
          <lpage>185</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2001</year>
          .
          <string-name>
            <given-names>Random</given-names>
            <surname>Forests</surname>
          </string-name>
          .
          <source>Machine learning</source>
          ,
          <volume>45</volume>
          (
          <issue>1</issue>
          ),
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Xgboost: A Scalable Tree Boosting System</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          (pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Chollet</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2015</year>
          . Keras. https://github.com/keras-team/keras Cortes,
          <string-name>
            <surname>C.</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Support-vector networks</article-title>
          .
          <source>Machine learning</source>
          ,
          <volume>20</volume>
          (
          <issue>3</issue>
          ),
          <fpage>273</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Freund</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Iyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>An Efficient Boosting Algorithm for Combining Preferences</article-title>
          . The Srivastava, N.,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Dropout: a simple way to prevent neural networks from overfitting</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1929</fpage>
          -
          <lpage>1958</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Wallace</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A sensitivity analysis of (and practitioners' guide to) convolutional neural networks for sentence classification</article-title>
          .
          <source>CoRR</source>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>