<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Challenges in Detecting Privacy Revealing Information in Unstructured Text</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Deutsche Telekom Chair of Mobile Business and Multilateral Security, Goethe University Frankfurt</institution>
          ,
          <addr-line>Theodor-W.-Adorno-Platz 4, 60323 Frankfurt am Main</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Welderufael.Tesfay</institution>
          ,
          <addr-line>Jetzabel.Serna, Sebastian.Pape</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper discusses the challenges in detecting privacy revealing information using ontologies, natural language processing and machine learning techniques. It reviews current de nitions, and sketches problem levels towards identifying the main open challenges. Furthermore, it elicits that the current notion of personally identi able information lacks robustness to be used in varying contexts and user perceptions, and shows the need to additionally consider privacy sensitive information.</p>
      </abstract>
      <kwd-group>
        <kwd>Privacy</kwd>
        <kwd>personally identi able information</kwd>
        <kwd>privacy sensitive information</kwd>
        <kwd>privacy revealing information</kwd>
        <kwd>ontology</kwd>
        <kwd>machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        These days many of our daily activities leave tremendous digital traces in the
internet. While these data can be useful to solve societal challenges in areas such
as health, transportation; it also presents threats to our personal and societal
sovereignty with regard to privacy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In this regard, the research community has
been proposing di erent legal and technical mechanisms to safeguard privacy [
        <xref ref-type="bibr" rid="ref1 ref5">5,
1</xref>
        ]. At the centre of most of these proposals is the idea of anonymising the identity
of users and reducing the disclosure of personal data. In order to reduce the
disclosure of privacy revealing information (PRI), personalised privacy protection
tools for internet users need to be developed. In this regard, Shah and Manisha
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] recommend machine learning techniques to detect PRI which often has a
certain pattern. Similarly, Caliskan-Islam et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] showed that ontologies can
also be useful to detect privacy revealing textual data. However, regardless of the
fact that personally identi able information (PII) is often used as benchmark
for privacy prevention measures, there is still no clearly established notion of
what they consist of. This is due to complexities and context dependencies
since the question of which information makes a user personally identi able
also depends on the group considered. This makes it particularly challenging
to research in PRI detection, especially those using ontology, natural language
processing and classi cation algorithms. Furthermore, there is the notion of
privacy sensitive information (PSI), which highly depends on the user's own
perception. For instance, in many social media sites the user is not anonymous,
and thus should be careful not to reveal PSI. This is even more important if
the user is participating in the social media site over a long time and PSI could
be obtained by a simple correlation of her postings. But like PII, the notion
of PSI is not only fuzzy by itself, it also depends on whom the user is sharing
the information with. For example, in anonymous support-groups each member
reveals privacy sensitive information (the purpose of the support-group) but does
not want the knowledge to di use outside of this group. In real-world groups the
user is not anonymous to other members. However, in online anonymous support
groups, e.g. an online forum with pseudonyms, the user may keep her identity
private. Then it is particularly important, that the user does not reveal PII when
posting in the forum, because that would threaten her anonymity. This should
hold again, if the user is member of the group over a long time.
      </p>
      <p>This shows, that { depending on the scenario/use-case { PII or PSI may be
revealed (un)intentionally and the user's privacy is at risk as soon as someone is
able to gather the matching part. We elaborate on this topic in Sect. 2.</p>
      <p>In summary this paper identi es the challenges in detecting PII and PSI
taking into consideration current approaches so far. The scope is limited to only
textual content created by users, e.g. postings on social media sites such as
Reddit. Thus the following data sources are out of scope of this work:
{ Any (meta-)data collected on the user during the communication (e.g. with a
social media site). This includes IP-address or ngerprinting the user's client
{ Any other media than text. E.g. pictures or video les, even though they
may contain meta-data such as exchangeable image le format (Exif) data.
{ Information from other persons or users</p>
      <p>The remainder of the paper is organised as follows. Sect. 2 discusses the
notion and relation of personally identi able and privacy sensitive information.
Sect. 3 structures the underlying problems. In Sect. 4 related work is given and
Sect. 5 focuses on open challenges. Sect. 6 concludes our work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Terminology</title>
      <p>
        A commonly used classi cation of privacy sensitiveness of data is based on whether
the given data ful ls the de nition of PII [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Identity theft and other privacy
violations exploit these identi ers along with other background information of
the target to compromise their privacy. However, privacy research hasn't yet
developed an all-embracing de nition for PII that will consider speci c contexts.
This induces a "grey area" between what kind of information should be considered
PII and non-PII. Article 4 of the recently approved European General Data
Protection Regulation (GDPR) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] de nes 'personal data' as any information
relating to an 'identi ed' or identi able natural person ('data subject'). An
identi able natural person can be identi ed, directly or indirectly, in particular
by reference to an identi er such as a name, an identi cation number or to one
or more factors speci c to the physical, physiological, genetic, mental, economic,
cultural or social identity of that natural person [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Furthermore, Shilton [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
stated that privacy decisions have multiple parts such as identity (who is interested
in the data), granularity (how much the data reveals about the target) and time
(for how long the data will be in use). Hasan et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] also identi ed six di erent
information types that could potentially identify a user: individual characteristics
(age, name, gender etc), knowledge (information associated to speci c knowledge
in a given domain), interests (hobbies, sports, professional activities etc), goals
(short term or long term users' wishes or intentions to achieve something in a
given context), behavior (online activity, mobility patterns etc), and context
information (spatio-temporal information).
      </p>
      <p>
        In their recent work, Schwartz and Solove [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] developed a PII 2.0 model
in which they propose a scale going from "no possibility of identi cation" to
"individual can be clearly identi ed". This scale is further divided into three
categories: (1) identi ed, (2) identi able or (3) non-identi able person or entity.
In this context identi ed means that a speci c person can be distinguished from
others. In the second category the linkage to the speci c person has not yet been
made, but is likely to happen with additional information. The last category
refers to data, which cannot be used to identify a person.
      </p>
      <p>
        Thus far, we have seen how PII have been de ned from technical and
nontechnical perspectives. However, as Narayanan and Shmatikov [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] have rightly
mentioned, relying on PII for privacy risks analysis, such as re-identi cation of
anonymised or pseudonymized information, at the age of big data is fallacious.
For this reason we go back to the de nition of personal data in the GDPR,
which is de ned as \any information relating to a [...] person". We argue that
for analysing the user's privacy \any information" is too far-reaching and one
should concentrate on sensitive data. However, analyzing privacy sensitiveness
involves the users' own perception about the given attribute and goes beyond
merely categorising data as PII or non-PII. Hence, we introduce the notion of
privacy sensitive information (PSI), which is any information that, depending on
the user's perception, has the consequence of revealing privacy of the individual.
While PII and PSI have some commonalities, not all PII is necessarily PSI, and
likewise not all PSI is PII. However, on the long run collections of PSI may result
in PII. For further clarity, we introduce PRI, which is the superset of PII and
PSI, and we recommend privacy research to consider PRI when dealing with
privacy revelation detection and development of mitigating mechanisms.
      </p>
      <p>Concluding, we note that { as already sketched in Sect. 1 { the critical path
is linking a user's PSI with her PII. If only PSI or PII about a user is known,
either the PSI cannot be linked to a user or there is no PSI about the identi able
user. This is in accordance with the de nition in the GDPR.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Problem Statement</title>
      <p>The task considered in this paper is to nd PII and PSI in unstructured texts.
We assume those texts are created by the user, e.g. as a posting in social media
or a chat message. We do not consider structured texts, e.g. the elds in a form
of a social media or chat user pro le when the user is asked for certain attributes
such as age, gender, location. The task to check whether the input of a certain
eld belongs to the questioned category and is a truthful answer is fundamentally
di erent than identifying relevant information in unstructured texts.</p>
      <p>
        We have identi ed four di erent levels of di culty depending on the amount
of data considered.
1. Identify PRI in an unstructured text : Given an unstructured text, the
challenge is to identify PII and PSI within the text. As already stated, the notion
of PII depends on the context/domain and the notion of PSI additionally
depends on the user perceptions, which information she classi es as sensitive.
2. Identify PRI with historical information : Given a number of unstructured
texts, the challenge is to identify PII and PSI by analysing the set of
(historical) texts. Since social media or chats usually involve a series of postings, in
this level, historical information (previous unstructured texts) also needs to
be considered. This results in the analysis of multiple texts, their timeline
and the analysis of possible PII or PSI spread among them.
3. Identify PRI with side information : Given an unstructured text and publicly
available information, the challenge is to identify PRI within the text by
considering additional sources of information. PRI may be derived if the
information contained in that text is combined with public knowledge, e.g.
any kind of demographic data or public information such as Open Data or
Wikipedia (cf. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]). This is especially di cult due to the huge amount of
available data and the large number of possible combinations of sources to
infer PII or PSI from the text with (a subset of) the available side information.
4. Identify PRI with side information and historical information : Given a
number of unstructured texts, the challenge is to identify PRI within the
texts by considering additional publicly available information. This is the
combination of available information in the problem levels 2 and 3.
      </p>
      <p>The di culty of the described levels forms a partial order. Obviously, levels 2
and 3 are more di cult than level 1 and level 4 is more di cult than levels 2
and 3. It is unclear if level 2 or 3 is more di cult than the other (cf. Fig. 1).</p>
      <p>4
with history &amp; side information
w2ithhistory</p>
      <p>with side 3information</p>
    </sec>
    <sec id="sec-4">
      <title>Related Work</title>
      <p>
        Detecting the presence and degree of sensitivity of private information is the
rst step towards empowering users with support in privacy decision making [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
In this regard, Wang et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] have introduced a tool to support social-science
researchers, which performs a real-time analysis of unstructured short texts in
order to extract conceptual associations. Authors focused on content generated
in Twitter and implemented an association extraction module to determine the
relevance of each word to the target keyword (e.g. privacy) using the real-time
pointwise mutual information statistical association measure. However, the main
outcome is not to identify privacy sensitive information, but instead to discover
which concepts or topic users associate with privacy.
      </p>
      <p>
        Mao et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] presented an analysis of privacy leaks on Twitter. Their main
contribution was to provide an initial understanding on what type of privacy
information users reveal on their tweets. The authors limited the scope of their
study to three privacy-related topics, namely, vacation plans, in uence of alcohol
and medical conditions. Furthermore, for each of those categories they built up a
classi er in order to determine whether the tweet could be classi ed as privacy
sensitive according to the detected content.
      </p>
      <p>
        Jindal et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] applied semi-supervised machine learning techniques to
identify privacy sensitive data in medical texts. Their approach relied on information
contained in the hierarchical structure of a large medical encyclopaedia. Following
a similar direction, Caliskan-Islam et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] proposed a privacy detective tool able
to detect a broader range of privacy sensitive information. Authors combined a
number of techniques, namely, topic modelling, named entity recognition, privacy
ontology, and sentiment analysis in order to represent privacy features and trained
a classi er based on Naive Bayes. They further analyzed Twitter users' time-line
and computed a privacy scoring for each classifying them according to their
sharing information behavior of privacy related information.
      </p>
      <p>
        Gill et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] developed a privacy dictionary able to distinguish between privacy
and non-privacy content. The authors state that their privacy dictionary has six
state categories. However, their categorisation lacks contextual and probabilistic
rules to encompass more privacy related words. Zhang et al. [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] applied ontologies
for rule-based privacy protection in pervasive computing environments. They
used two properties (Data is and and Disclose when) to check for type of data
and ful lment of condition before data disclosure. However, the Data is class
consists of PII as de ned in the Platform for Privacy Preferences Project (P3P)
and was only extended by a location attribute. Therefore, this approach still
lacks user's personal perception of privacy.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Challenges</title>
      <p>In what follows, we illustrate the main challenges we have identi ed in the
problem domain. Based on their mappings into the problem levels from Sect. 3,
the challenges are further categorised into two, namely: general and speci c ones.
General Challenges: these challenges regard all problem levels from Sect. 3.
Users' privacy perception - depending on di erent parameters such as
educational background, previous privacy incidence/experience, perceived privacy
risk, etc users have varying levels of concern and perceptions of privacy.
Therefore, privacy detection systems built on the one-size- ts-all notion do
not fairly address privacy choices of individuals. This is particularly
challenging to rule-based ontologies and supervised machine learning approaches.</p>
      <p>Additionally, users' privacy perception may change over time.</p>
      <p>
        Privacy paradox - users exhibit a complex and paradoxical dichotomy between
their privacy (declarations) concerns and actual behaviour [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. This is
challenging to privacy disclosure analysis because on the one hand the decision
involves whether to consider actual or declared behaviour, while on the other
hand, prediction models can only grasp from the user activities, hence the
actual behaviour, unable to accommodate the user's behavioural wishes.
Information privacy sensitiveness classi cation - when detecting PSI, for
example illness, the degree of sensitiveness of a cold/ u might be substantially
di erent from that of cancer [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Even though both attributes appear to ful l
privacy sensitiveness class, it is technically quite challenging to di erentiate
between the two. This is essentially related to the challenge of regarding the
user's privacy perception except that this one focuses on the technicalities.
Context dependence - users privacy preferences are quite dependent on the
context, e.g., a person can in a given situation be indi erent to privacy while
the same person becomes very concerned about privacy in another situation
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. If we consider the location data, a user may worry less about sharing it
when in work place than when in a bar. In this case, ontologies and machine
learning models can easily detect it as PII, but whether it is PSI or not
strongly depends on the context.
      </p>
      <p>
        Domain speci c - there is a lack of comprehensive privacy dictionaries,
ontologies, etc. which are not only context-aware [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], but have domain speci c
knowledge about PRI. Current ontologies are restricted for a speci c target
(e.g. legal compliance [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) or a certain usecase (e.g. service orientation [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
Language issues - at the basic level text analysis requires proper sentence
structures. Thus, text normalisation tools are needed before a proper analysis
could be done. At the most advanced level, language variability also needs
to be considered as languages are rich and there exist di erent ways to
express certain meanings. Furthermore, most research works have focus on
the English language and therefore, there is a lack of available resources
User preferences - users have di erent requirements on settings, e.g. there
need to be di erent levels for experts or laymen. Getting this wrong, as
well as mistakenly classifying PRI has a major e ect on the usability, user
acceptance and adoption of the provided tools.
      </p>
      <p>Speci c Challenges: these challenges map into the speci c problem levels (cf.
Sect. 3) as stated in each challenge below.</p>
      <p>
        Data linkages resulting in PRI - by analysing multiple pieces of information
new PRI could be found which are commonly not considered as such. With
the rise of publicly available information, especially open data (cf. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]),
linking seemingly harmless information with side information (cf. Sect. 3,
problem level 3) could result in privacy breaches. This is one of the reasons
why de-anonymization attacks (cf. Sweeney [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]) work.
      </p>
      <p>
        Buildup of information - some information may not be privacy revealing if
only a small quantity of data is considered, but may be sensitive in larger
amounts. This holds for PII, e.g. revealing a location just one time might
be safe, while you can build movement pro les from a continuous report of
location data, e.g. by a mobile phone [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. But this also holds for PSI, e.g.
revealing only one base pair from your DNA might be safe while revealing a
longer sequence might lead to implications [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The di cult question behind
this challenge is, when is the aggregated amount of data privacy revealing?
This is related to considering historical information (problem level 2).
      </p>
      <p>
        PII and PSI may also be derived from other media, such as images or videos.
Further work should consider the media itself, e.g. by using machine learning
techniques [
        <xref ref-type="bibr" rid="ref19 ref22">22, 19</xref>
        ] to annotate/describe the content of the le. Additionally,
further information within the les like exif-meta-data should be considered.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Summary and Conclusion</title>
      <p>In this work, we investigated the challenges and opportunities in inferring privacy
sensitive information from textual data using ontology, classi cation and natural
language processing based techniques. While personal identi able information and
privacy sensitive information should be both considered, we showed fundamental
di erences between them. To structure the challenges in this area, we have
de ned di erent problem levels based on the information which needs to be
regarded. A closer look at the challenges revealed that more research is needed,
especially to consider speci c privacy perceptions by di erent users when using
the aforementioned techniques.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <article-title>Eu general data protection regulation. Legal document</article-title>
          ,
          <source>EU</source>
          (
          <year>2016</year>
          ), http://eur-lex.europa.eu/legal-content/EN/TXT/?uri=uriserv:OJ.L_ .
          <year>2016</year>
          .
          <volume>119</volume>
          .01.0001.01.ENG&amp;toc=OJ:L:
          <year>2016</year>
          :119:TOC, accessed
          <year>July</year>
          2016
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Acquisti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brandimarte</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loewenstein</surname>
          </string-name>
          , G.:
          <article-title>Privacy and human behavior in the age of information</article-title>
          .
          <source>Science</source>
          <volume>347</volume>
          (
          <issue>6221</issue>
          ),
          <volume>509</volume>
          {
          <fpage>514</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Callahan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Us dhs handbook for safeguarding sensitive personally identi able information</article-title>
          . Washington, DC (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Casellas</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nieto</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Meron~o,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Roig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Torralba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Reyes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Casanovas</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Ontological semantics for data privacy compliance: The NEURONA project</article-title>
          .
          <source>In: Intelligent Information Privacy Management</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dwork</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Di erential privacy</article-title>
          .
          <source>In: Automata, languages and programming</source>
          , pp.
          <volume>1</volume>
          {
          <fpage>12</fpage>
          . Springer (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toledo</surname>
            ,
            <given-names>M.B.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Capretz</surname>
            ,
            <given-names>M.A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allison</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blair</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grace</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flores</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Towards a base ontology for privacy protection in service-oriented architecture</article-title>
          .
          <source>In: 2009 IEEE International Conference on Service-Oriented Computing and Applications (SOCA)</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gill</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasalou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papoutsi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joinson</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          :
          <article-title>Privacy dictionary: a linguistic taxonomy of privacy for content analysis</article-title>
          .
          <source>In: Proceedings of the SIGCHI conference on human factors in computing systems</source>
          . pp.
          <volume>3227</volume>
          {
          <fpage>3236</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Habegger</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brunie</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bennani</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Damiani</surname>
          </string-name>
          , E.:
          <article-title>A discussion of privacy challenges in user pro ling with big data techniques: The eexcess use case</article-title>
          .
          <source>In: Proceedings of the IEEE International Congress on Big Data</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Islam</surname>
            ,
            <given-names>A.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walsh</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenstadt</surname>
          </string-name>
          , R.:
          <article-title>Privacy detective: Detecting private information and collective privacy behavior in a large social network</article-title>
          .
          <source>In: Proceedings of the 13th Workshop on Privacy in the Electronic Society</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jindal</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gunter</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Detecting privacy-sensitive events in medical text</article-title>
          .
          <source>In: Proceedings of the 5th ACM Conference on Bioinformatics</source>
          , Computational Biology, and Health Informatics, BCB '
          <volume>14</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kum</surname>
            ,
            <given-names>H.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahalt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Privacy by design: understanding data access models for secondary data</article-title>
          .
          <source>American Medical Informatics Association (AMIA) Joint Summits on Translation Science and Clinical Research Informatics</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shuai</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kapadia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Loose tweets: An analysis of privacy leaks on twitter</article-title>
          .
          <source>In: Proceedings of the 10th Annual ACM Workshop on Privacy in the Electronic Society</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Narayanan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shmatikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Myths and fallacies of personally identi able information</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>53</volume>
          (
          <issue>6</issue>
          ),
          <volume>24</volume>
          {
          <fpage>26</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Pape</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serna-Olvera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tesfay</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Why open data may threaten your privacy</article-title>
          . In: Workshop on Privacy and
          <article-title>Inference, co-located with KI</article-title>
          (
          <year>September 2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Annas</surname>
            ,
            <given-names>G.J.:</given-names>
          </string-name>
          <article-title>Dna testing, banking, and genetic privacy</article-title>
          .
          <source>New England Journal of Medicine</source>
          <volume>355</volume>
          (
          <issue>6</issue>
          ),
          <volume>545</volume>
          {
          <fpage>546</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solove</surname>
            ,
            <given-names>D.J.:</given-names>
          </string-name>
          <article-title>Pii 2.0: Privacy and a new approach to personal information</article-title>
          .
          <source>Privacy and Security Law Report</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valera</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Survey of sensitive information detection techniques: The need and usefulness of machine learning techniques</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Shilton</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Four billion little brothers?: Privacy, mobile phones, and ubiquitous data collection</article-title>
          .
          <source>Commun. ACM</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Siddiqui</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mishra</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verma</surname>
            ,
            <given-names>J.S.:</given-names>
          </string-name>
          <article-title>Article: A survey on automatic image annotation and retrieval</article-title>
          .
          <source>International Journal of Computer Applications</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Spitz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Tell-all telephone</article-title>
          . http://www.zeit.de/datenschutz/malte-spitzdata-retention (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Sweeney</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Simple demographics often identify people uniquely</article-title>
          .
          <source>Tech. rep.</source>
          , Carnegie Mellon University (
          <year>2000</year>
          ),
          <source>data Privacy Working Paper 3</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A survey on automatic image annotation and trends of the new age</article-title>
          .
          <source>Procedia Engineering</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quercia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seaghdha</surname>
            ,
            <given-names>D.O.</given-names>
          </string-name>
          :
          <article-title>Reading tweeting minds: Real-time analysis of short text for computational social science</article-title>
          .
          <source>In: Proceedings of the 24th ACM Conference on Hypertext and Social Media</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Young</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quan-Haase</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Privacy protection strategies on facebook: The internet privacy paradox revisited</article-title>
          .
          <source>Information, Communication &amp; Society</source>
          <volume>16</volume>
          (
          <issue>4</issue>
          ),
          <volume>479</volume>
          {
          <fpage>500</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>N.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Todd</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A privacy agent in context-aware ubiquitous computing environments</article-title>
          .
          <source>In: IFIP International Conference on Communications and Multimedia Security</source>
          . Springer (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>