<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hyperbolic Embeddings for Preserving Privacy and Utility in Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oluwaseyi Feyisetan</string-name>
          <email>sey@amazon.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tom Diethe</string-name>
          <email>tdiethe@amazon.co.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Drake</string-name>
          <email>draket@amazon.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amazon</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>Guaranteeing a certain level of user privacy in an arbitrary piece of text is a challenging issue. However, with this challenge comes the potential of unlocking access to vast data stores for training machine learning models and supporting data driven decisions. We address this problem through the lens of d -privacy, a generalization of Di erential Privacy to non Hamming distance metrics. In this work, we explore word representations in Hyperbolic space as a means of preserving privacy in text. We provide a proof satisfying d -privacy, then we de ne a probability distribution in Hyperbolic space and describe a way to sample from it in high dimensions. Privacy is provided by perturbing vector representations of words in high dimensional Hyperbolic space to obtain a semantic generalization. We conduct a series of experiments to demonstrate the tradeo between privacy and utility. Our privacy experiments illustrate protections against an authorship attribution algorithm while our utility experiments highlight the minimal impact of our perturbations on several downstream machine learning models. Compared to the Euclidean baseline, we observe &gt; 20x greater guarantees on expected privacy against comparable worst case statistics.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Copyright ©2020 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0). Presented at the PrivateNLP 2020
Workshop on Privacy in Natural Language Processing Colocated with 13th ACM
International WSDM Conference, 2020, in Houston, Texas, USA.</p>
      <p>PrivateNLP ’20, February 7, 2020, Houston, TX, USA
© 2020
SUMMARY
• User’s goal: meet some specific need with
respect to a query x
• Agent’s goal: satisfy the user’s request
• Question: what occurs when x is used to
make other inferences
• Mechanism: Modify the query to protect
privacy whilst preserving semantics
• Our approach:</p>
      <p>Hyperbolic Metric Differential Privacy
METRIC DIFFERENTIAL PRIVACY</p>
      <p>P[M(x) 2 E]  e" dh(x,x0)P [M (x0) 2 E] + .
HYPERBOLIC WORD EMBEDDINGS</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>