<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Multi-Stage Approach to Detect Privacy Breaches in Physician Reviews</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Frederik S. Baumer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joschka Kersting</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Orlikowski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michaela Geierhos</string-name>
          <email>geierhosg@mail.upb.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Semantic Information Processing Group, Paderborn University, Germany https://go.upb.de/seminfo</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Physician Review Websites allow users to evaluate their experiences with health services. As these evaluations are regularly contextualized with facts from users' private lives, they often accidentally disclose personal information on the Web. This poses a serious threat to users' privacy. In this paper, we report on early work in progress on \Text Broom", a tool to detect privacy breaches in user-generated texts. For this purpose, we conceptualize a pipeline which combines methods of Natural Language Processing such as Named Entity Recognition, linguistic patterns and domain-speci c Machine Learning approaches which have the potential to recognize privacy violations with wide coverage. A prototypical web application is openly accesible.</p>
      </abstract>
      <kwd-group>
        <kwd>Detection of Privacy Violations</kwd>
        <kwd>Physician Reviews</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Due to progressive semantic enrichment, the Web is becoming a vast resource
for exciting data-driven applications. However, this also poses a threat to
individual users. As newly exposed data can be linked with existing resources
more and more e ectively, even implicitly disclosed individual pieces of personal
information may have harmful consequences for users. An important example
in this context are Physician Review Websites (PRWs), which enable users to
rate medical services on the Web. To provide an authentic rating, patients
often add private information, e.g. about locations, diseases or medication. This
makes them potentially identi able by third parties [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this paper, we present
work in progress on Text Broom1, a tool which detects and highlights privacy
breaches in user-generated texts to prevent accidental disclosure of information.
While possible privacy threats are obvious with respect to speci c entities (e.g.
locations), they are much less obvious in full texts [
        <xref ref-type="bibr" rid="ref1 ref2 ref5">1, 2, 5</xref>
        ]. Natural language
(NL) allows us to share information in numerous subtle ways, so that
information is more than a sum of words. Therefore, in order to prevent violations
of privacy and personality rights [
        <xref ref-type="bibr" rid="ref4 ref6 ref7">4, 7, 6</xref>
        ], a lot of open challenges need to be
1 A prototype of TextBroom is available under https://bit.ly/2vrDd5Q
adressed [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Related work [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] uses existing Named Entity Recognition (NER)
methods to detect explicitly mentioned entities in texts and remove or randomly
replace them. We substantially extend this approach by also considering
inherent private information. For example, the sentence\As mother of three girls "
discloses information about family relations and gender without stating them
explicitly. Similar sentences are very frequent in physicians' reviews [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Therefore,
we combine Natural Language Processing (NLP) methods (NER, linguistic rules
and patterns), knowledge resources (linked data, gazetteers) and domain-speci c
machine learning models to detect potential privacy violations. Note, that we
report on early work in progress and do not provide quantitative evaluations.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Text Broom: A</title>
    </sec>
    <sec id="sec-3">
      <title>Multi-Stage Approach</title>
      <p>
        Text Broom basically adapts the idea of Baumer et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], who recognize privacy
breaches in physician reviews using a pipeline of NLP approaches (multi-stage
approach) which provide di erent perspectives and levels of granularity. Baumer
et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] use a series of patterns which reach high precision, but low recall. They
also do not take into account the problem of ambiguities and the complexity of
grammatical constructions. For this reason, the Text Broom pipeline processes a
much wider range of linguistic information. We also follow Kleinberg and Mozes
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], who focus on the visualization of potential violations, and extend this
approach with ideas from explainability research.
      </p>
      <p>Preprocessing
POS-Tagging
Domain-specific</p>
      <p>Gazetteers
Lexical &amp; Syntactical</p>
      <p>Disambiguation
Language Detection</p>
      <p>I</p>
      <p>II
1. Level Detection
2. Level Detection</p>
      <p>III</p>
      <p>Summarization &amp; IV</p>
      <p>Cleaning
Semantic Role Labeling</p>
      <p>Information Extraction
Linguistic Patterns</p>
      <p>Phrase Classification
Named Entity
Recognition</p>
      <p>Information Gathering
&amp; Reclassification</p>
      <p>Highlighting
Explainability
Cleaning</p>
      <p>
        Gazetters are used in the preprocessing stage and contain extensive lists of
terms for drugs and diseases. These are technical terms which are not covered by
domain-unspeci c methods (e.g. general NER). In addition to doctor portals and
pharmacy websites, we also use Wikidata for the maintenance and extension of
the drug gazetteers. Since the tokens de ned in the gazetteers are applied
without taking into account further context, the found matches are exclusively used
as input for other components (e.g. linguistic patterns). Lexical and syntactic
disambiguation allows us to improve detection quality. Lexical disambiguation
provides information about which reading of a word is meant in the given
context and can minimize recognition errors. An example is \The doctor also treats
my mom ", where \mom" stands for mother (breach of privacy towards a third
party) and not e.g. Mars Orbiter Mission. More importantly, disambiguation
allows us to consider additional surface forms which way refer to the mentioned
entity (e.g. mom, mother, female parent). These may be more likely to occur
in gazetteers and thus potentially increase recall. Linguistic patterns such as
\(My)?+GAZETTEER+(ADV)?+VERB" use gazetteers in a prede ned context [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In
this example, the context is de ned by the optional adjective \My" and a
following verb. Theoretically, suitable verbs could also be maintained in a gazetteer,
but we aim to detect a broader range of candidate privacy breaches.
Consequently, this type of pattern tends to detect a lot of false positives. Therefore,
we combine several di erent detection methods. For example, in the case of \My
mother also goes to this doctor ", the word \mother" would also be recognized
by the NER component as a person, so that there is additional evidence of a
potential breach. Information gathering and reclassi cation is an
important processing step, since, as mentioned, di erent components discover many
and sometimes contradictory evidence for privacy breaches. This component can
combine and re-evaluate indications of potential privacy violations that are
related, merge and re-evaluate the di erent types of evidence and, in principle,
classify the identi ed entities. For example, disclosing a real name is clearly
more serious than merely mentioning the name of a drug in isolation and should
be marked accordingly. Explanation generation or explainability is located
in the fourth stages. This component generates explanations for why certain
words, phrases or sentences are potentially harmful from a privacy perspective.
It is motivated by research in the eld of fair, accountable and transparent (FAT)
machine learning and explainable arti cial intelligence. Since enabling users to
understand the privacy implications of indvidual statements is an explicit design
goal of Text Broom, we regard this component to be very important. It is still
in very early stages. Currently, we highlight segments of text which have been
detected as potential privacy breaches directly based on the output of the other
components.
      </p>
      <p>Fig. 2 shows Text Broom's web interface with exemplary input and output.
In the given example, our system detects a notable amount of evidence for
potential breaches, but there are also obvious problems. For example, the system
has not detected the drug \propranool" (Propranolol) due to a spelling mistake
and ignored the co-reference between \Dr. Nase" and \he". To tackle these and
similar problems in a rigorous fashion, we are currently working on annotating
a comprehensive evaluation dataset. But even at this early stage, Text Broom
showcases a number of promising approaches to improve online privacy. As soon
as the tool has matured, we will provide a server-sided interface (API).
Professional providers can then integrate Text Broom into their services to help users
to protect their privacy and also prevent their business from encountering legal
issues due to data protections laws, especially within the European Union.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S.</given-names>
            <surname>Afroz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Islam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stolerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Greenstadt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>McCoy</surname>
          </string-name>
          .
          <article-title>Doppelganger nder: Taking stylometry to the underground</article-title>
          .
          <source>In Proceedings of the 2014 IEEE Symposium on Security and Privacy. IEEE</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Balduzzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Platzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Holz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kirda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Balzarotti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Kruegel</surname>
          </string-name>
          .
          <article-title>Abusing Social Networks for Automated User Pro ling</article-title>
          .
          <source>In Proceedings of the 13th International Conference on RAID</source>
          , pages
          <volume>422</volume>
          {
          <fpage>441</fpage>
          , Berlin / Heidelberg,
          <year>2010</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Ba</surname>
          </string-name>
          umer,
          <string-name>
            <given-names>N.</given-names>
            <surname>Grote</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kersting</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Geierhos</surname>
          </string-name>
          .
          <article-title>Privacy matters: Detecting nocuous patient data exposure in online physician reviews</article-title>
          .
          <source>In Proceedings of the 23rd ICIST</source>
          <year>2017</year>
          , Communications in Computer and Information Science, volume
          <volume>756</volume>
          , pages
          <fpage>77</fpage>
          {
          <fpage>89</fpage>
          ,
          <string-name>
            <surname>Druskininkai</surname>
          </string-name>
          , Lithuania,
          <year>2017</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>R.</given-names>
            <surname>Bild</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Prasser</surname>
          </string-name>
          .
          <article-title>SafePub: A truthful data anonymization algorithm with strong privacy guarantees</article-title>
          .
          <source>Proceedings on PET</source>
          ,
          <year>2018</year>
          (1),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>G.</given-names>
            <surname>Danezis</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Troncoso</surname>
          </string-name>
          .
          <article-title>You cannot hide for long</article-title>
          .
          <source>In Proceedings of the 12th ACM workshop on Workshop on privacy in the electronic society. ACM</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>O.</given-names>
            <surname>Ferrandez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>South</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Friedlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Samore</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Meystre</surname>
          </string-name>
          .
          <article-title>BoB, a best-of-breed automated text de-identi cation system for VHA clinical documents</article-title>
          .
          <source>JAMIA</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <volume>77</volume>
          {
          <fpage>83</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>M.</given-names>
            <surname>Geierhos</surname>
          </string-name>
          and
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Ba</surname>
          </string-name>
          <article-title>umer. Erfahrungsberichte aus zweiter Hand: Erkenntnisse uber die Autorschaft von Arztbewertungen in Online-Portalen</article-title>
          .
          <source>In Book of Abstracts der DHd-Tagung</source>
          <year>2015</year>
          , pages
          <fpage>69</fpage>
          {
          <fpage>72</fpage>
          ,
          <string-name>
            <surname>Graz</surname>
          </string-name>
          ,
          <year>2015</year>
          . DHd.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>B.</given-names>
            <surname>Kleinberg</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mozes</surname>
          </string-name>
          .
          <article-title>Web-based text anonymization with node.js: Introducing NETANOS (named entity-based text anonymization for open science)</article-title>
          .
          <source>The Journal of Open Source Software</source>
          ,
          <volume>2</volume>
          (
          <issue>14</issue>
          ):
          <fpage>293</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>