<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Web Enabled Record Linkage Attacks on Anonymized Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jacob Miracle</string-name>
          <email>miracle.13@wright.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michelle Cheatham</string-name>
          <email>michelle.cheatham@wright.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DaSe Lab, Wright State University</institution>
          ,
          <addr-line>Dayton, OH 45435</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Big Data analytics holds the promise of enabling new discoveries in medicine, more e cient business practices, and other important advances. However, much of the data involved in such analyses contains personally identi able information (PII) that needs to be removed or obscured prior to release in order to protect individuals' privacy. Anonymizing a dataset is not as easy as it seems, however, and many supposedly anonymous datasets are vulnerable to a record linkage attack. Most of these attacks are currently conducted manually and can be labor-intensive, but as semantic web technologies continue to gain popularity, the potential for automating various aspects of these attacks increases. This paper explores the components of a record linkage attack and how semantic web technologies could play a role in facilitating them.</p>
      </abstract>
      <kwd-group>
        <kwd>Anonymization</kwd>
        <kwd>Semantic Web</kwd>
        <kwd>Privacy</kwd>
        <kwd>De-anonymization Attacks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Humanity is producing more data than ever before, and in the hands of
researchers this data has led to important insights in a variety of domains, from
medicine to construction to marketing. Privacy concerns are obviously an issue
in this environment, and most data sets that are made public or used in research
have been through some sort of anonymization process. Unfortunately, this
process frequently consists entirely of removing explicit identi ers, such as name,
email address, and social security number, from the data. In many cases, such an
anonymization process still leaves the data vulnerable to a record linkage attack.</p>
      <p>Consider a dataset that contains the following elds: name, social security
number, job title, gender, age, zip code, and salary. A company wishes to make
this dataset available to the media to show that there is no pay gap between
men and women in the organization. To avoid revealing the salary of individual
employees, the names and social security numbers are removed. However, the
remaining elds can still be used to link individuals to their salaries given the
presence of an appropriate secondary dataset. For example, if the company also
keeps its employees' CVs on its website that contain their names, job titles, and
addresses (and a person's gender can often be inferred from their name), then
it might be possible to link the two data sources and thereby associate names
with salaries. The privacy of employees who have an uncommon job title or live
in a sparsely populated zip code is particularly at risk in this scenario.</p>
      <p>
        Tables 2a and 2b show an example of a record linkage attack involving these
datasets. In this case, gender, job title, and zip code are quasi-identi ers that
can be used to link records across the two datasets. For example, it can be
inferred that Jonathan Wilson makes $54,750 because he is the only male Software
Developer living in the 24932 zip code, but it is not known who makes $37,500
because there are several people with the same combination of gender, job title,
and zip code: Abby Johnson, Victoria Stevens, and Stephanie Lewis. The risk
to a particular person's privacy depends on the number of people who share
that individual's quasi-identi er values. Latanya Sweeney called this number of
people k and established the concept of k-anonymity, in which the values in a
dataset are generalized such that there are at least k records for each
combination of quasi-identi er values [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In this example, if only the rst four digits
of the zip codes are included in the data, then Jonathan Wilson's salary can no
longer be determined, because he is now indistinguishable from another person.
This increased privacy comes at a cost { since the full zip codes are no longer
included in the data, an analysis of a gender equality for salary can no longer
control for di erences in local standard of living costs as accurately.
      </p>
      <p>
        Other researchers have since expanded upon the initial model of k-anonymity
[
        <xref ref-type="bibr" rid="ref1 ref10 ref9">1, 10, 9</xref>
        ]. The underlying idea remains the same throughout these works: even
when explicit identi ers have been removed from a dataset, some identities may
be discovered if another dataset that contains explicit identi ers shares some
elds and individuals with the anonymous dataset can be found. Finding such
a dataset is possible in a surprising number of circumstances. In our own work,
we have deanonymized annual workforce surveys from the Ohio Board of
Nursing by linking them with voter registration rolls and the state licensing website.
Finding an appropriate dataset with which to link the target data can be di
cult though, involving a lot of manual search and evaluation of possibilities. In
this position paper we discuss how the widespread adoption of Semantic Web
technologies may make conducting record linkage attacks quicker and easier in
the future. Awareness of the potential negative as well as the positive uses of
such technologies is an important rst step towards designing and developing
privacy safeguards on the Semantic Web.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data Availability</title>
      <p>Record linkage attacks require a dataset against which to link the anonymized
data. Historically, most data was inconvenient to use since it was available only
as databases or on le servers as spreadsheets, CSV les, or tables in PDF
documents. Retrieving the data could also be di cult. For instance, some repositories
might be accessible via websites or structured query mechanisms while others
required a login and use of secure le transfer protocols. Financial drawbacks
also inhibited data integration. Some data might be stored using proprietary
formats that required expensive software licenses to read. These obstacles made
nding and retrieving data related to an attacker's target dataset di cult.</p>
      <p>
        The rise of linked data has changed this situation drastically. Linked data is
expressed as RDF and can be accessed using standard protocols. It is also proli c.
The Linked Open Data (LOD) Cloud now contains over 31 trillion triples across
295 datasets, with more than 503 million links across datasets [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The most
represented domains in the cloud are social network information and government
data, composing over 51% and 18%, respectively of the total [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This is a huge
amount of information about many di erent aspects of people's lives, and the
potential for its misuse should not go unconsidered. For example, Table 2 shows
data from the Open University in the United Kingdom.1 This information can be
downloaded in various formats or accessed via a SPARQL endpoint. It includes
information about a person's job title, groups they belong to, and publications
they have co-authored. While this data is not generally considered sensitive,
it could be used as a quasi-identi er for a target dataset. Additionally, some
information included in this data, including usernames on social media platforms
such as Twitter and LinkedIn and a list of other linked datasets in which this
person appears, can be used to nd more information about this person.
      </p>
      <p>Another quickly growing type of data on the Semantic Web is that annotated
with schema.org markup.2 Schema.org is an initiative by major search engine
companies to facilitate the description of entities and the relationships between
them using a basic syntax expressed as RDFa, Microdata, or JSON-LD. As of
2014, more than 36% of websites in Google's crawl contained schema.org markup</p>
      <sec id="sec-2-1">
        <title>1 http://data.open.ac.uk/page/context/people/pro les</title>
      </sec>
      <sec id="sec-2-2">
        <title>2 https://schema.org</title>
        <p>
          [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. It is particularly commonly used to describe people, businesses, events, and
reviews. Fields relevant to people include many likely quasi-identifying elds,
such as birth date, birth place, gender, nationality, and a liation.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data Relevance</title>
      <p>While the rise of linked data and schema.org markup has made much more data
available in an easily accessible manner, a record linkage attack relies on nding
datasets that include relevant information about the individuals in the target
dataset. Finding an appropriate dataset is often the most time-consuming aspect
of a record linkage attack. However, some research currently underway can speed
up this process, thereby lowering the barrier to deanonymization.</p>
      <p>
        Many linked datasets have very complex or extremely simple schemas. It is
often di cult for a potential user of a dataset to quickly identify whether or not
the data is useful for their purpose, but numerous methods for summarizing a
linked dataset speed up this process. For example, the linked data summarization
approach described in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] ranks the axioms within an ontology based on
graphbased measures such as centrality, while Loupe is an online tool that provides
statistics regarding the usage of properties to describe instances of particular
types within a linked dataset.3
      </p>
      <p>Visualization tools are another avenue for quickly determining the general
content and structure of a dataset. Many visual interfaces for data exploration
on the Semantic Web involve displaying the RDF data as a graph.
Unfortunately, graph-based representations frequently place entities based on graph
metrics such as centrality or density, rather than according to their semantic
meaning. They also have di culty scaling to large datasets without becoming
unwieldy. Kow and his colleagues have attempted to move beyond this towards
more semantic-based layout algorithms with their idea of an \information
landscape," which places similar concepts near one another and labels clumps of</p>
      <sec id="sec-3-1">
        <title>3 http://loupe.linkeddata.es/loupe/</title>
        <p>
          entities with terms describing the group. Users can select areas of interest that
seem likely to contain relevant entities, which automatically lters the mappings
shown in the list. This method of ltering allows users to systematically explore
an ontology at a high level of detail without losing track of the big picture [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
Other approaches handle the problem of scalability by providing an RDF triple
browser interface rather than attempting to show the entire dataset at once [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          Ontology Design Patterns (ODP) provide another means for quickly
determining linked dataset relevance to attackers. ODPs are self-contained, reusable
patterns that model concepts that commonly occur across di erent ontologies.
A well designed ODP describes the key aspects, and only the key aspects, of the
concept being modeled [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. If an attacker can quickly isolate the part of a complex
schema most related to an ODP of interest, or quickly determine whether or not
data with little schema information ts into the ODP model, they would have a
better idea of whether or not the dataset in question was useful. For example,
a person's communications often reveal much about them. Blomqvist posted an
ODP on the website ontologydesignpatterns.org to model a \Communications
Event." A simpli ed illustration of this ODP is shown in Figure 1.
        </p>
        <p>
          The dataset containing information about the 2012 European Semantic Web
Conference available at http://data.semanticweb.org/dumps/conferences/
contains information about, among other things, the keynote talks given at the
conference, including their start and end times, the speaker, the topic, the
setting in which the talk occurred, and the title and subject matter. A method
for detecting the presence of an ODP (such as Communications Events) in a
linked dataset was proposed by Khan and Blomqvist in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. For datasets with a
signi cant schema, an ontology alignment system (described in Section 4) could
also be used to recognize that this dataset contains Communications Events.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data Linking</title>
      <p>Once a relevant dataset has been identi ed, the target dataset must be joined
with it based on the quasi-identi er values. This seems straightforward, but can
actually be quite di cult in practice because the schemas for the two datasets
were likely developed by di erent people, for di erent purposes. Because of this,
even two ontologies that represent the same domain will generally not be the
same. They may use synonyms for the same concept or the same word for
different concepts, they may be at di erent levels of abstraction, they may not
include all of the same concepts, and they may not even be in the same language.
Furthermore, the classes and properties in the ontologies may not be used
consistently when describing the entities within the dataset. The goal of ontology
alignment is to determine when an entity in one ontology is semantically related
to an entity in another ontology, despite these challenges.</p>
      <p>
        The Ontology Alignment Evaluation Initiative (OAEI) is a set of
benchmarks for evaluating the performance of alignment systems. The initiative has
held evaluations annually since 2005. Over that time, the accuracy and the
variety of problems handled by alignment systems have increased, while runtimes
have decreased.4 The top performing alignment systems include two that are
available online: AgreementMakerLight5 and LogMap.6 These systems achieve
an F-measure of .76 and .73, respectively, on an OAEI track based on aligning
ontologies related to conference organization. These results are approaching the
level of consensus that humans have when performing alignment tasks [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
implying that the dataset linking phase may be an aspect of record linkage attacks
that could be automated in the near future. Additionally, alignment systems
could be used to attempt to align datasets to ODPs representing key concepts,
such as a Person, in order to re ne a collection of possibly-relevant datasets for
further analysis. Aligning a dataset against an ODP rather than another dataset
can be easier, due to the limited scope and application-neutral nature of an ODP.
      </p>
      <p>
        Coreference resolution algorithms attempt to determine when the same
instance (i.e. individual) is referred to in two in di erent ways. For instance, is John
Q Public in one dataset the same person as J.C. Publick in another? This is the
Semantic Web technology most closely related to deanonymization: determining
whether a person whose name and social security number have been replaced
with random strings, for example, is present within an external dataset such
as voter registration records is precisely what coreference resolution algorithms
attempt. Most current approaches use string similarity metrics to compare two
instances based on their property values (e.g. zip code, age, height) or their
property values together with the names of those properties. Top performing systems
on the mainbox instance matching task include the aforementioned LogMap and
RiMOM [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], with F-measures of .83 and .91, respectively.
      </p>
      <sec id="sec-4-1">
        <title>4 http://oaei.ontologymatching.org</title>
      </sec>
      <sec id="sec-4-2">
        <title>5 https://github.com/AgreementMakerLight/AML-Jar</title>
      </sec>
      <sec id="sec-4-3">
        <title>6 http://csu6325.cs.ox.ac.uk</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Data Inferencing</title>
      <p>Traditional record linkage attacks sometimes involve speci c or general
knowledge or assumptions an attacker has about a target. This can be made
significantly easier using automated reasoners. For example, assume the attacker is
working with a medical records dataset organized according to a schema which
includes the statements below.
&lt;exSchema:hasDisease&gt; &lt;rdfs:domain&gt; &lt;exSchema:Person&gt;
&lt;exSchema:hasDisease&gt; &lt;rdfs:range&gt; &lt;exSchema:Disease&gt;
&lt;exSchema:LungDisease&gt; &lt;rdfs:subClassOf&gt; &lt;exSchema:Disease&gt;
&lt;exSchema:HeartDisease&gt; &lt;rdfs:subClassOf&gt; &lt;exSchema:Disease&gt;
&lt;exSchema:hasEthnicity&gt; &lt;rdfs:domain&gt; &lt;exSchema:Person&gt;
&lt;exSchema:hasEthnicity&gt; &lt;rdfs:range&gt; &lt;xsd:string&gt;
&lt;http://data.ex.org/person/12345&gt; a &lt;exSchema:Person&gt;
&lt;http://data.ex.org/person/12345&gt; &lt;rdfs:nameFull&gt; "Zhang Lu"</p>
      <p>If the attacker knows that Zhang Lu has some disease and wants to determine
what it is, he can add some additional statements to the knowledge base that
re ect his assumptions and then use a reasoner to check whether or not it is
possible to infer the disease Mr. Zhang has, based on those assumptions. For
instance, the attacker may know that Mr. Zhang is of Asian ancestry. He could
then add the following fact to the knowledge base:
&lt;http://data.ex.org/person/12345&gt; &lt;exSchema:hasEthnicity&gt; "Asian"</p>
      <p>The attacker could further assume that people of Asian ancestry are unlikely
to get heart disease (based on statistical knowledge). The attacker would add
the following fact to the knowledge base:
(exSchema:Person and (exSchema:hasEthnicity Asian)) SubClassOf:
not (exSchema:hasDisease some exSchema:HeartDisease)</p>
      <p>The attacker could then use an automated reasoner to determine whether
or not Mr. Zhang's disease could be inferred. While space constraints force this
example to be relatively simplistic, it shows that the attacker can use existing
Semantic Web languages and tools to quickly explore the rami cations of any
assumptions he would like to make.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>This position paper explored the potential for Semantic Web technologies to
facilitate record linkage attacks against anonymized datasets. The emergence of
linked data and schema.org markup has increased the availability of data that
can be used to conduct attacks. Tools for data summarization and visualization
can assist an attacker in sifting through this data to nd a relevant dataset
with which to link the target dataset. Meanwhile, the emergence of automated
techniques for ODP identi cation, ontology alignment, coreference resolution,
and reasoners hold the potential to one day fully automate record linkage attacks.</p>
      <p>
        The rami cations of Semantic Web technologies on privacy are likely to be
profound. Aggarwal showed that typical approaches to anonymize data break
down in the face of high dimensionality [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which is precisely what linked data
provides. Dealing with this may require di cult decisions about how to publish
sensitive data, potentially involving perturbing the sensitive values [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which
has corresponding impacts on its utility. These concerns are present whether the
sensitive data is published as linked data or in a database, CSV le, or other
traditional format, because as we have seen, even innocuous data about a person
available on the Semantic Web can be used to deanonymize a standalone dataset.
      </p>
      <p>In future work on this topic, we plan to assess the volume of data containing
potential quasi-identi ers that currently exists on the Semantic Web and develop
a data vulnerability assessment tool based on Semantic Web technologies.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aggarwal</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          :
          <article-title>On k-anonymity and the curse of dimensionality</article-title>
          .
          <source>In: Proceedings of the 31st International Conference on Very Large Databases</source>
          . pp.
          <volume>901</volume>
          {
          <issue>909</issue>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Aggarwal</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Philip</surname>
          </string-name>
          , S.Y.:
          <article-title>A general survey of privacy-preserving data mining models and algorithms</article-title>
          . In:
          <article-title>Privacy-preserving data mining</article-title>
          , pp.
          <volume>11</volume>
          {
          <fpage>52</fpage>
          . Springer (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Anja</surname>
            <given-names>Jentzsch</given-names>
          </string-name>
          , Richard Cyganiak,
          <string-name>
            <surname>C.B.</surname>
          </string-name>
          :
          <article-title>State of the lod cloud</article-title>
          (
          <year>September 2011</year>
          ), http://lod-cloud.net/state/ [Online; accessed 29-February-2016]
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cheatham</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The Properties of Property Alignment on the Semantic Web</article-title>
          .
          <source>Ph.D. thesis</source>
          , Wright State University (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cheatham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hitzler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Conference v2.
          <article-title>0: An uncertain version of the oaei conference benchmark</article-title>
          .
          <source>International Semantic Web</source>
          Conference pp.
          <volume>33</volume>
          {
          <issue>48</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Erling</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikhailov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Faceted views over large-scale linked data</article-title>
          .
          <source>LDOW</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomqvist</surname>
          </string-name>
          , E.:
          <article-title>Ontology design pattern detection-initial method and usage scenarios</article-title>
          .
          <source>In: SEMAPRO</source>
          <year>2010</year>
          ,
          <source>The Fourth International Conference on Advances in Semantic Processing</source>
          . pp.
          <volume>19</volume>
          {
          <issue>24</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kow</surname>
            ,
            <given-names>W.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabol</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Granitzer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kienrich</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lukose</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A visual soa-based ontology alignment tool</article-title>
          .
          <source>In: Proceedings of the 6th International Conference on Ontology Matching</source>
          . pp.
          <volume>242</volume>
          {
          <fpage>243</fpage>
          .
          <string-name>
            <surname>CEUR-WS. org</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkatasubramanian</surname>
          </string-name>
          , S.:
          <article-title>t-closeness: Privacy beyond k-anonymity and l-diversity</article-title>
          . In: ICDE (Purdue University) (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Machanavajjhala</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kifer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gehrke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkitasubramaniam</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>l-diversity: Privacy beyond k-anonymity</article-title>
          (
          <year>March 2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mika</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potter</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Metadata statistics for a large web corpus</article-title>
          .
          <source>LDOW</source>
          <volume>937</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chung</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          :
          <article-title>Rimom-im: A novel iterative framework for instance matching</article-title>
          .
          <source>Journal of Computer Science and Technology</source>
          <volume>31</volume>
          (
          <issue>1</issue>
          ),
          <volume>185</volume>
          {
          <fpage>197</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sweeney</surname>
          </string-name>
          , L.:
          <article-title>k-anonymity: A model for protecting privacy</article-title>
          .
          <source>International Journal on Uncertainty, Fuzziness and Knowledge-based Systems 10 (May</source>
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , X., Cheng, G.,
          <string-name>
            <surname>Qu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Ontology summarization based on rdf sentence graph</article-title>
          .
          <source>In: Proceedings of the 16th International Conference on World Wide Web</source>
          . pp.
          <volume>707</volume>
          {
          <fpage>716</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>