<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Concept Representations for Concept Drift Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oliver Becher</string-name>
          <email>becher@cwi.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Hollink</string-name>
          <email>hollink@cwi.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Desmond Elliott</string-name>
          <email>d.elliott@ed.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centrum Wiskunde &amp; Informatica</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Edinburgh</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>We present an approach to estimating concept drift in online news. Our method is to construct temporal concept vectors from topicannotated news articles, and to correlate the distance between the temporal concept vectors with edits to the Wikipedia entries of the concepts. We find improvements in the correlation when we split the news articles based on the amount of articles mentioning a concept, instead of calendar-based units of time.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 INTRODUCTION</title>
      <p>
        Concepts in Knowledge Organisation Systems (KOSs) are used
to provide structured annotations and background knowledge in
a wide variety of applications. They enhance interoperability
between datasets and enable structured access to annotated document
collections. These benefits, however, are compromised when
concept change (or drift) occurs. Wang et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] define three types of
concept drift: (1) change in the intension of the concept, defined
as the definition or the properties of the concept; (2) change in the
extension, or the instances, of a concept; and (3) change in the label
of the concept. Each type of concept drift may lead to problems
for applications working with KOSs. For example, an annotation
of a document may become invalid if the intension of the concept
changes. Correspondences between two concepts in diferent KOSs
may become incorrect if the extension of one of them changes.
A user’s keyword query on a historic corpus may be interpreted
incorrectly if the (prevalent) label to refer to a concept has changed.
      </p>
      <p>
        Significant progress has been made in the detection of meaning
change of words (e.g., [
        <xref ref-type="bibr" rid="ref3 ref9">3, 9</xref>
        ]). They are based on distributional
methods, where the meaning of a word is defined as the context in
which it appears. A change in context over time may then signify a
change in meaning. In this paper, we study change in the meaning of
concepts in a KOS. Drawing inspiration from work on word-change
detection, we aim to explore whether the change of a concept can
© 2017 Copyright held by the author/owner(s).
      </p>
      <p>SEMANTiCS 2017 workshop proceedings: Drift-a-LOD
September 11-14, 2017, Amsterdam, Netherlands
be measured from changes in how it appears in the context of a
document collection. This is diferent from other work on concept
change in KOSs in the sense that we ignore changes in the structure
of the KOS.</p>
      <p>
        This paper is an initial step towards understanding how the
context of a concept can be represented to efectively capture concept
change. Our representation is based on the co-occurrence between
concepts that appear as annotations of documents in a diachronic
collection: if two concepts co-occur if they are annotations of the
same document. Hence, a concept can be seen as a vector of
cooccurrence counts with other concepts in the KOS. Concept change
can then be measured by comparing vectors created for diferent
time spans in the collection. We experiment with various versions
of this basic idea, and apply it to detect change in an annotated
document collection: the ION dataset of 300k online news articles,
annotated with Wikipedia pages [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>To evaluate our method, we use Wikipedia edit counts. This is
based on the idea that a Wikipedia article is edited when a change
to the page was needed; hence, a higher number of edits may signify
a change in the underlying concept. Generally speaking, evaluation
of concept drift detection methods is hampered by a lack of large
scale evaluation datasets. Wikipedia edits are not to be seen as a
gold standard of concept drift. While some edits might be due to
a change in the concept, others might be, for example, additions
of missing information or corrections of previous mistakes. Our
assumption is that even though Wikipedia edit counts are a noisy
signal with respect to concept change, a correlation between our
change scores and the edits counts does say something about the
efectiveness of our method.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>REPRESENTING CONCEPTS</title>
    </sec>
    <sec id="sec-3">
      <title>Creating Concept Vectors</title>
      <p>Given a concept vocabulary C with N concepts, we create vector
representations of the concepts through their usage in a document
collection.</p>
      <p>We assume there is a collection of time-ordered documents D.
A document di is annotated with M topic annotations t1, . . . , tM ,
drawn from a total of T topics. Each document in the collection
can be represented as a binary document topic vector, di ∈ R1xT .
An element in the document topic vector takes a value of 1 if the
document has been annotated with that topic. We also assume a
function f: T → C that maps between the topic annotation and
concept vocabulary.</p>
      <p>We construct a concept vector cj for each concept in our
vocabulary c1, . . . , cN from co-occurrence counts of the topic annotations
in documents in the document collection. The set of concept
vectors forms a sparse matrix C ∈ RN ×N , where each row defines a
concept through co-occurrence with other concepts.</p>
      <p>Our concept vectors are co-occurrence counts. We reduce the
efect of frequently occurring concepts by re-weighting the
vectors using a TF-IDF-like weighting scheme, so that tf-idf(ci , cj ) =
t f (ci , cj ) ∗ id f (ci ), where t f (ci , cj ) is the number of times that
concept ci co-occurs with concept cj and id f (ci ) = loд d f (CN,ci ) + 1,
with d f (C, ci ) as the count of ci concept annotations in the entire
concept vocabulary C.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Temporal Concept Vectors</title>
      <p>Recall that we are interested in measuring the change in the
meaning of a concept over time. We redefine C to include a temporal
dimension, V ∈ RN x N x K , where the third dimension represents
K units of time, and Ík Vk = V ∈ RN x N . There are many ways
to define K: the document collection can be split into days, weeks,
months, or any other valid approach to splitting the collection
according to the sequential ordering of the documents. Note that the
co-occurrence statistics over topic annotations needs to be
calculated such that only documents timestamped between consecutive
units of time are used in the calculation, i.e. t=s1 and t=s2 are used
to define a temporal concept vector vj,s2 at t=s2.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Temporal Vector Distance</title>
      <p>We measure the change in the meaning of concepts by comparing
the vectors in the temporal concept matrix between subsequent
units of time. Specifically, we measure the change in a concept cj
between time k and k − 1 using a similarity metric sim(·, ·):
distance(vj, s, s-1) = sim(vj,s, vj,s−1) (1)</p>
      <p>
        We experiment with two similarity metrics: cosine similarity,
previously used to detect concept drift [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and KL-divergence (when
the vectors represent distributions).
3
3.1
      </p>
    </sec>
    <sec id="sec-6">
      <title>APPLICATION TO AN ANNOTATED NEWS</title>
    </sec>
    <sec id="sec-7">
      <title>COLLECTION</title>
    </sec>
    <sec id="sec-8">
      <title>Dataset and Model Application</title>
      <p>
        We explore our method for constructing concept representations
and measuring concept change with a dataset of online news
articles [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This data set contains news articles together with topic
annotations and images in their natural textual context. The
richness of information and meta data in this dataset can give many
ways to define and explore concepts, while a defined structure of
the data helps to use it reliably and consistently.
      </p>
      <p>The dataset contains articles published online between August
2014 – August 2015. In total, it includes more than 300K articles from
ifve publishers across British and US English sources: Daily Mail,
The Independent, New York Times, Hufington Post, and the
Washington Post. The articles are annotated with topics using TextRazor1.
TextRazor uses Wikipedia as a topic vocabulary. This vocabulary
ranges from narrowly defined concepts, e.g., The United States
Women’s Soccer Team or Electromagnetism, to broader concepts,
e.g., Sport or Science. The average number of topic annotations per
article is 25 broad ’Category pages’ and 5 specific (non-category)
pages, giving in total 122,000 distinct topic annotations on all
articles.
1http://www.textrazor.com</p>
      <p>We define our concept vocabulary C as a subset of TextRazor’s
topic vocabulary T: we retain only topics that are associated with at
least 2 articles. In preliminary experiments, we found that concepts
that are associated with too few articles have sparse representations
resulting in unrealistic change scores between the representations.
This leaves us with N=70,000 concepts. The mapping function f: T →
C is trivial in this case. However, the structured nature of Wikipedia,
and the links that it provides to other concept vocabularies, provide
starting points for other mapping functions, allowing us to explore
other concept vocabularies in the future.</p>
      <p>We construct concept vectors using the method outlined in
Section 2. The vocabulary of the concept vectors is defined over the
Wikipedia entries, therefore it is trivial to map the topic annotations
to the concept vectors.
3.2</p>
    </sec>
    <sec id="sec-9">
      <title>Visualization</title>
      <p>
        To visualize the change that a concept c has undergone, we create
a stream graph [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] of the temporal concept vectors of c. Figure 1,
for example, plots the temporal vectors of the Wikipedia concept
Police. Each ‘stream’ represents a concept that co-occurs with Police
in the document collection. The thickness of the line represents
the co-occurrence count at a certain time period. Since stream
graphs are suited to convey changes over time of only a limited
number of concepts, we select only those that occur most frequently.
Specificaly, we create ’streams’ for only those concepts that are
among the top 5 most frequently co-occurring concepts in any of
the temporal concept vectors of concept c.
      </p>
      <p>In figures 1, 2, and 3 we plot two concepts for which the average
change is low (measured as a high average cosine similarity between
12 temporal vectors) and one where the average change score is
high. Figure 1 shows that Police is a stable concept: the top five
most frequently occurring concepts remain frequent troughout the
year, and the volume of documents in which they co-occur hardly
lfuctuates. However, a concept might change on a larger time scale
than given in the data. Nonetheless, Police seems to be more stable
than other concepts in the time span.</p>
      <p>The concept Labour_Party (Figure 2) is stable as well: although
there is a burst in the volume of documents about this concept, there
is hardly a change in which concepts co-occur in these documents.
In other words, there is change in how much reporting there is
about the Labour_Party, but not in how they are reported.</p>
      <p>Figure 3 shows the streamgraph of the New York University. We
can see that the most co-occurring topics are constantly
changing in the streamgraph, both in periods with a high volume of
documents and in periods with a low volume of documents. This
suggests changes in how much and how New York University has
been reported in the news.
4
4.1</p>
    </sec>
    <sec id="sec-10">
      <title>TOWARDS A QUANTITATIVE EVALUATION</title>
    </sec>
    <sec id="sec-11">
      <title>Measuring Concept Change</title>
      <p>
        Concept change detection is hard to evaluate for a lack of gold
standard datasets [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Kenter et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] use a small sets of 21
humanjudged change scores. Frermann and Lapata [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] indirectly evaluate
change detection by using it in an application for which a gold
standard exists, namely the SemEval task for dating text. To the best
of our knowledge, large scale datasets to directly valuate change
detection, do not exist.
      </p>
      <p>For our application, we explore the use of Wikipedia edit rates to
evaluate our method of concept change representation. We believe
that the act of editing a Wikipedia page can signal a change in the
information that is relevant to that entry.</p>
      <p>Specifically, given a concept c and a pre-defined K units of time,
we measure changes scores as the consecutive temporal vector
distances for concept c (Section 2.3); then, we count the number of
Wikipedia edits to the aligned article during each of the K units of
time. We evaluate our method by measuring the Spearman
correlation between the change scores (i.e. the temporal vector distances)
and the Wikipedia edit counts. The higher the correlation, the more
accurately the temporal concept vectors can estimate the rate of
change of the Wikipedia entries.</p>
      <p>We perform an experiment on 964 concepts. Since it seems likely
that the number of articles that a concept is related to plays a role,
we draw a stratified random sample from our concept vocabulary to
include both frequently and infrequently used concepts. We select
three diferent strati of even size. Group 1 contains concepts which
are related to more than 500 articles. Group 2 contains concepts
which are related to at least 200 articles but not more than 500.
Group 3 contains concepts with at least 24 articles but less than 200.
The sample includes only concepts that map to ‘regular’ Wikipedia
pages and not Category pages.</p>
      <p>
        Figure 4 plots the number of articles that a concept is related to
against the average cosine similarity between the temporal vectors
of that concept. This shows that the more frequent a concept is used
as an annotation, the higher the average cosine similarity, i.e., the
lower the change. This is analogous to the change of meaning of
words [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], where the semantic changes of words scale with inverse
frequency, known as the law of conformity.
      </p>
      <p>We compare four models, each with diferent settings regarding
the way that time units are set, the use of TF-IDF, and the choice
of similarity measure (either cosine similarity or KL-divergence).
4.2</p>
      <p>Models
4.2.1 Fixed Time Bins (Cosine). Starting with the most basic
setup of our method, we calculate temporal concept vectors for time
frames (or bins) of a fixed duration. With n time frames, each frame
covers an n/year th of the dataset. For example, with 52 frames,
each frame covers exactly one week. We use the cosine similarity
to calculate change scores between each temporal concept vector.</p>
      <p>4.2.2 Flexible Time Bins (Cosine). In this model, we calculate
temporal concept vectors for time periods that each cover a fixed
amount of articles. Thus, time frames difer in length of days rather
than amount of data. The amount of articles per bin depends on the
total amount of articles available per concept. Analogous to Fixed
Time Bins, we create n bins, therefore we assign a nth of the total
amount of articles to each bin. However, a concept may have such
an amount of articles that does not split evenly into n bins. Thus, it
may be split into more than n bins. With these vectors, we use the
cosine similarity to calculate change scores. We use the same time
frames to bin the Wikipedia edits and estimate a correlation.</p>
      <p>Run / n bins &gt; 100 &gt; 52 &gt; 24 &gt; 12 &gt; 6</p>
      <p>Fixed Time Bins (Cos) 0.07 0.18 0.22 0.36 0.26
Flexible Time Bins (Cos) -0.2 -0.19 -0.14 0.03 0.33
- TF-IDF (Cos) -0.2 -0.2 -0.13 0.0 0.26</p>
      <p>Flexible Time Bins (KL) 0.23 0.25 0.29 0.19 -0.3
Table 1: Average Spearman correlation between concept
similarity scores and Wikipedia edits. Negative correlations are
good for cosine similarity; positive correlations are good for
KL-divergence.</p>
      <p>4.2.4 Flexible Time Bins (KL-divergence). This model is identical
to Flexible Time Bins except we measure the distance between
temporal concept vectors using Kullback-Leibner divergence (KL)
instead of cosine similarity.</p>
    </sec>
    <sec id="sec-12">
      <title>4.3 Results</title>
      <p>We collect Spearman correlation coeficients for 964 concepts using
diferent numbers of time frames (6, 12, 24, 52, and 100). Table 1
shows the average correlation over concepts that are significantly
correlated with Wikipedia edits. Note that the experiments with
Cosine similarity measure between temporal concepts should return a
negative correlation, while the experiments with the KL-divergence
distance should return a positive correlation. The results in Table 1
show that the performance of the models decreases as we decrease
the number of time bins.</p>
      <p>The Fixed Time Bins (Cosine) model only returns positive
correlations, indicating that fixed units of time (in this case, splitting
the articles into months) does not act as a reliable proxy for
concept change in our dataset. The Flexible Time Bin experiments
(Cosine) and (-TF-IDF) are better correlated with Wikipedia edits
than the Fixed Time Bin model. We do not find a diference in not
re-weighting the concept vectors using TF-IDF. Finally, we find
a small improvement from using KL-divergence as the temporal
vector distance metric instead of Cosine similarity. Throughout, we
can see that the number of temporal bins n is a crucial parameter
in our experiment.</p>
      <p>We performed a follow-up analysis of the efect of the number
of temporal bins. The histograms in Figures 5a to 5b show the
distributions of the Spearman correlations for the Flexible Time Bins
(KL) model with n=12 or n=100. We find that the ratio of positively
correlations to negative correlations is substantially reduced by
having more time bins. More time bins clearly improves the quality
of the concept vectors.</p>
    </sec>
    <sec id="sec-13">
      <title>5 CONCLUSION AND FUTURE WORK</title>
      <p>We explored concept change using vector space concept
representations. The concept vectors were constructed from topic
cooccurrence in a large collection of online news articles. We
introduced a temporal aspect to the vectors by requiring the
cooccurrences to happen within pre-defined windows of time. We
explored to what extend concept change can be evaluated by
correlating the distance between its temporal concept vectors and edits
to the Wikipedia article corresponding to the concept.</p>
      <p>We found that a flexible approach to defining a window of time
was more successful than using calendar-based windows of time.
We also found that having more windows of time resulted in better
correlations between the temporal vector distances and Wikipedia
article edits.</p>
      <p>Future work includes an analysis of which types of concepts
correlate to Wikipedia edits counts, to get more insights into the
use of Wikipedia as an evaluation tool. Similarly, we could look into
the types of edits made on Wikipedia to distinguish actual change
from simple growth of an article.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Lee</given-names>
            <surname>Byron</surname>
          </string-name>
          and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Wattenberg</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Stacked Graphs - Geometry and Aesthetics</article-title>
          .
          <source>IEEE transactions on visualization and computer graphics 14 6</source>
          (
          <issue>2008</issue>
          ),
          <fpage>1245</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Lea</given-names>
            <surname>Frermann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mirella</given-names>
            <surname>Lapata</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A Bayesian Model of Diachronic Meaning Change</article-title>
          .
          <source>Transactions of the ACL 4</source>
          (
          <year>2016</year>
          ),
          <fpage>31</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>William</surname>
            <given-names>L Hamilton</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jure Leskovec</surname>
            , and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Diachronic word embeddings reveal statistical laws of semantic change</article-title>
          .
          <source>arXiv preprint arXiv:1605.09096</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Laura</given-names>
            <surname>Hollink</surname>
          </string-name>
          , Adriatik Bedjeti, Martin van Harmelen,
          <string-name>
            <given-names>and Desmond</given-names>
            <surname>Elliott</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A Corpus of Images and Text in Online News</article-title>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Laura</given-names>
            <surname>Hollink</surname>
          </string-name>
          , Sándor Darányi,
          <source>Albert Meroño Peñuela, and Efstratios Kontopoulos</source>
          .
          <year>2017</year>
          . First Workshop on Detection,
          <article-title>Representation and Management of Concept Drift in Linked Open Data: Report of the Drift-a-</article-title>
          <string-name>
            <surname>LOD2016 Workshop</surname>
          </string-name>
          : Front Matter..
          <source>In Knowledge Engineering and Knowledge Management. EKAW 2016 (Lecture Notes in Computer Science)</source>
          , Vol.
          <volume>10180</volume>
          .
          <fpage>15</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Tom</given-names>
            <surname>Kenter</surname>
          </string-name>
          , Melvin Wevers, Pim Huijnen, and Maarten de Rijke.
          <year>2015</year>
          .
          <article-title>Ad hoc monitoring of vocabulary shifts over time</article-title>
          .
          <source>In Proceedings of the 24th International Conference on Information and Knowledge Management</source>
          .
          <fpage>1191</fpage>
          -
          <lpage>1200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Astrid</surname>
            <given-names>van Aggelen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laura Hollink</surname>
          </string-name>
          , and Jacco van Ossenbruggen.
          <year>2016</year>
          .
          <article-title>Combining distributional semantics and structured data to study lexical change</article-title>
          .
          <source>In European Knowledge Acquisition Workshop</source>
          . Springer,
          <fpage>40</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Shenghui</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Stefan Schlobach</surname>
            , and
            <given-names>Michel</given-names>
          </string-name>
          <string-name>
            <surname>Klein</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Concept drift and how to identify it</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web 9</source>
          ,
          <issue>3</issue>
          (
          <year>2011</year>
          ),
          <fpage>247</fpage>
          -
          <lpage>265</lpage>
          . https://doi.org/10.1016/j.websem.
          <year>2011</year>
          .
          <volume>05</volume>
          .003
          <string-name>
            <given-names>Semantic</given-names>
            <surname>Web Dynamics Semantic Web Challenge</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Yating</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Adam Jatowt, and
          <string-name>
            <given-names>Katsumi</given-names>
            <surname>Tanaka</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Towards understanding word embeddings: Automatically explaining similarity of terms</article-title>
          .
          <source>2016 IEEE International Conference on Big Data (Big Data)</source>
          (
          <year>2016</year>
          ),
          <fpage>823</fpage>
          -
          <lpage>832</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>