<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recognizing and Reducing Bias in NLP Applications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dirk Hovy Universita Bocconi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italia dirk.hovy@unibocconi.it</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>As NLP technology becomes used in ever more settings, it has ever more
impact on the lives of people all around the world. As NLP practitioners, we
have become increasingly aware that we have the responsibility to evaluate
the e ects of our research and prevent or at least mitigate harmful outcomes.
This is true for academic researchers, government labs, and industry
developers. However, without experience of how to recognize and engage with
the many ethical conundrums in NLP, it is easy to become overwhelmed and
remain inactive. One of the most central ethical issues in NLP is the impact
of hidden biases that a ect performance unevenly, and thereby disadvantage
certain user groups.</p>
      <p>This tutorial aims to empower NLP practitioners with the tools spot these
biases, and a number of other common ethical pitfalls of our practice. We will
cover both high-level strategies, as well as go through speci c case sample
exercises. This is a highly interactive workshop with room for debate and
questions from the attendees. The workshop will cover the following broad
topics:</p>
      <p>Biases: Understanding the di erent ways in which biases a ect NLP
data, models, and input representations, including including strategies
to test for and reduce bias in all of them.</p>
      <p>Dual Use: Learning to anticipate how a system could be repurposed
for harmful or negative purposes, rather than its intended goal.
Privacy: Protecting the privacy of users both in corpus construction
and model building.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>