<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NERPII: A Python Library to Perform Named Entity Recognition and Generate Personal Identifiable Information</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simona Mazzarino</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Minieri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Gilli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Clearbox AI</institution>
          ,
          <addr-line>Turin</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Nowadays, the convergence of Artificial Intelligence and data privacy is of crucial importance. This paper introduces NERPII, a Python library utilizing Named Entity Recognition (NER) and synthetic data generation to identify and protect Personal Identifiable Information (PII). We discuss the architecture of NERPII and provide a concise tutorial on its application, demonstrating how to extract entity information from datasets containing personal data and synthesize new PII while preserving data characteristics. Additionally, the study discusses the library's potential contributions and implications for future research. In conclusion, NERPII emerges as a practical tool for addressing ethical concerns related to data privacy in the AI domain.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Named Entity Recognition</kwd>
        <kwd>Personal Identifiable Information</kwd>
        <kwd>Synthetic Data</kwd>
        <kwd>Data Privacy</kwd>
        <kwd>Python library</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In the era of data-driven insights and technological progress, the intersection of Artificial
Intelligence (AI) and data privacy has assumed an increasingly important role. AI has
profoundly reshaped societal dynamics, influencing everything from finance and healthcare to
transportation and entertainment. However, this progress has not been devoid of ethical
concerns, particularly those related to the exposure of Personal Identifiable Information (PII) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
PII refers to any data or information that can be used to identify, contact, or locate a specific
individual [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. PII includes a wide range of data elements, and it can be either sensitive or
non-sensitive. Sensitive PII includes information such as Social Security numbers, driver’s
license numbers, financial account numbers, and medical records, while non-sensitive PII may
include names, addresses, phone numbers, email addresses, and other information that, when
combined, can be used to identify a person.
      </p>
      <p>
        The protection of PII is a critical aspect of privacy and data security, as it involves safeguarding
individuals’ personal information from unauthorized access, use, or disclosure. Various laws
and regulations, such as the European Union’s General Data Protection Regulation (GDPR) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
and the United States’ Health Insurance Portability and Accountability Act (HIPAA) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], govern
the collection, storage, and handling of PII to ensure individuals’ privacy rights are respected
and their information is kept secure.
      </p>
      <p>
        In this perspective, some Natural Language Processing (NLP) techniques, such as Named
Entity Recognition (NER), a challenging learning problem that involves processing a text and
identifying certain occurrences of words or expressions as belonging to particular categories
of Named Entities (NE) [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ], if combined with the potential of synthetic data generation,
emerge as a tool that can be used with an ethical purpose, that is, to identify and categorize
entities, including names, phone numbers, credit card number, and other potentially sensitive
information that then can be replaced with synthetic data. The synthesis of data is a technique
that involves generating artificial data while preserving the statistical characteristics of real
data [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. So, combining synthetic data with NER can be a strong strategy to mitigate the risk
of data breaches and preserve data privacy.
      </p>
      <p>So, NER, as mentioned before, is a technique often used on unstructured data, namely texts.
However, companies or institutions often possess entire structured datasets, such as Excel
or CSV files, containing numerous PII records. The challenge, therefore, was to create a tool
capable of performing NER on tabular data in order to recognize PII and replacing it with
synthetic data.</p>
      <p>
        In literature, numerous tools have been created to associate semantic types with table columns.
For instance, Hulsebos et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] introduced Sherlock, a deep learning model that employs neural
networks to analyze various feature sets, including word embeddings, character embeddings,
and global statistics derived from individual column values. Extending this work, Zhang et al.
[11] developed Sato, which enhances Sherlock by incorporating table context and structured
output prediction to more efectively capture the correlations between columns within the same
table. In addition, there are tools that make use of pre-trained language models to annotate
columns: for instance, Deng et al. [12] developed TURL, a Transformer-based pre-training
framework for table understanding tasks, while Suhara et al. [13] developed Doduo, a multi-task
learning framework designed to take the entire table as input and uses a single model to predict
column types and relationships.
      </p>
      <p>Although these tools have achieved high performance on state-of-the-art benchmarks, their
architecture makes them computationally expensive. Furthermore, these tools only enable the
recognition of entities and relationships within a table, without the capability to regenerate a
potential PII once identified in order to ensure privacy.</p>
      <p>Considering this, we introduce NERPII1, a Python library that leverages NER methods to
efectively identify PII within structured data formats, such as CSV files, and subsequently
regenerate them in a privacy-preserving manner. To perform NER, we used Presidio [14], a
Microsoft SDK that provides a fast identification for private entities such as credit card numbers,
names, locations, phone numbers, financial data and more, and a BERT model (
dslim/bert-baseNER) [15] to identify additional entities. The use of the BERT model thus makes the tool more
robust, allowing for the expansion of the entities recognized by the library. Moreover, in order
to generate synthetic PII, we use Faker [16], a Python library created to generated fake PII. The</p>
      <sec id="sec-1-1">
        <title>1The library is available at the following link: https://github.com/Clearbox-AI/nerpii</title>
        <p>idea is to assign an entity to each column which contains PII in a dataset, and then replace each
values in that column with a coherent fake PII.</p>
        <p>In this paper, we describe the architecture of NERPII and we show a short tutorial on how to
use the library to extract entity information from a dataset containing personal information and
to synthesize new PII that maintains data characteristics while safeguarding privacy. Finally,
we discuss conclusion and future directions.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. NERPII Architecture</title>
      <p>The NERPII library consists of two distinct classes: NamedEntityRecognizer and FakerGenerator.</p>
      <p>NamedEntityRecognizer is used to perform Named Entity Recognition on structured data,
typically in the form of a CSV file. An instance of NamedEntityRecognizer requires several
parameters: a Pandas DataFrame containing the data to be subjected to NER, an optional count
indicating the desired number of samples to be processed (with a default value of 500), and
a NaN filler—a string used to fill NaN values in the dataset (set to ’?’ by default). Within the
NamedEntityRecognizer, a Presidio analyzer is initialized, which attempts to assign a named
entity from those supported by the library to each column of the dataset, and the BERT model
that tries to assign the organization entity to the columns. To assign an entity to each column,
the analyzer and the model first assign an entity to each value within the column. The entity
ultimately assigned to the column is the one that has been most frequently assigned to the
values contained within that column. When an instance of NamedEntityRecognizer is created,
the NER is performed on the dataset, resulting in a dictionary accessible through the instance’s
attribute dict_global_entities. Within this dictionary, for each column, the assigned entity and
the confidence score with which that entity has been assigned are recorded.</p>
      <p>FakerGenerator is used to generate synthetic PII. The parameters of a FakerGenerator object
include the dataset previously analyzed by the NamedEntityRecognizer and the associated
dict_global_entities dictionary. The generator divides columns for which the recognizer was
able to assign an entity from those without an associated entity. For those columns with an
associated entity, the generator replaces each value within the column with synthetic data. The
FakerGenerator can regenerate the following Named Entities: address, phone number, email
address, first and last name, city, state, URL, zipcode, credit card number, Social Security Number
(SSN), and country.</p>
      <p>Consequently, the output of the generator is a new dataset in which the originally present
PII has been substituted with synthetic PII.</p>
    </sec>
    <sec id="sec-3">
      <title>3. How to use NERPII</title>
      <p>Suppose you have a dataset containing personal information of several people, such as first
name, last name, phone number, address, e-mail, etc., like the one shown in Table 1. You need to
anonymize the PII contained in the dataset, so firstly, you have to install the library by cloning
the github repository or by simply installing it via pip.
pip install nerpii
Once you have installed the library, you can import the class NamedEntityRecognition by using
the following line of code.
from nerpii.named_entity_recognizer import NamedEntityRecognizer
Then, you can create a recognizer passing as parameter the path to the CSV file that contained
your dataset or directly your dataset in Pandas DataFrame format.
recognizer = NamedEntityRecognizer(’./csv_path.csv’)
Once you have created your recognizer, you can performed NER using the following functions.
recognizer.assign_entities_with_presidio()
recognizer.assign_entities_manually()
recognizer.assign_organization_entity_with_model()
These functions assign an entity to most of the columns. The final output is a dictionary, like
the one shown below, accessible with
recognizer.dict_global_entities
in which column names are given as keys and assigned entities and a confidence score as values.
{’first name’: {’entity’: ’PERSON’, ’confidence_score’: 0.9127725856697819},
’last name’: {’entity’: ’PERSON’, ’confidence_score’: 0.8625},
’address’: {’entity’: ’ADDRESS’, ’confidence_score’: 0.8926174496644296},
’city’: {’entity’: ’LOCATION’, ’confidence_score’: 0.8731343283582089},
’state’: {’entity’: ’LOCATION’, ’confidence_score’: 0.976},
’zip’: {’entity’: ’ZIPCODE’, ’confidence_score’: 1.0},
’phone’: {’entity’: ’PHONE_NUMBER’, ’confidence_score’: 0.888},
’email’: {’entity’: ’EMAIL_ADDRESS’, ’confidence_score’: 1.0}}
After performing NER on your dataset, you can generate new PII using Faker. You can import
the class FakerGenerator by using the following command.
from nerpii.faker_generator import FakerGenerator</p>
      <sec id="sec-3-1">
        <title>Then, you can create a generator as follows.</title>
        <p>generator = FakerGenerator(dataset, recognizer.dict_global_entities)</p>
      </sec>
      <sec id="sec-3-2">
        <title>Finally, to generate new PII you can run this command line.</title>
        <p>generator.get_faker_generation()
At the end of the whole process you will have obtained a dataset, identical to the original (see
Table 2), where the values in the various columns will have been replaced with synthetic PII.</p>
        <p>Overall, the library has been tested on two openly available and two proprietary data sets.
The Classic Models data set represents a Customer Relationship Management database 2, while
the AWS Honeypot data set 3 comes from the cybersecurity domain. The proprietary data sets
were a synthetic table depicting a financial fraud detection use case and a table containing user
data collected by an IT department. The following table contains an overview of the metrics
achieved on the aforementioned tests.</p>
      </sec>
      <sec id="sec-3-3">
        <title>2Classic Models data set: https://relational.fit.cvut.cz/dataset/ClassicModels 3AWS Honeypot data set: https://datadrivensecurity.info/blog/pages/dds-dataset-collection.html</title>
        <p>first name last name</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion and Future Directions</title>
      <p>The evolution of AI has undeniably transformed numerous sectors, contributing to the creation
of several new technologies and showing rapid progress. However, this transformation has not
been without its ethical implications, particularly regarding the exposure of sensitive personal
data. The synthesis of data coupled with NLP techniques, ofers a promising solution that ensure
the protection of PII.</p>
      <p>NERPII, as detailed in this paper, tries to combine the remarkable potential of AI with the
importance of data privacy. By employing NER methods and harnessing the power of synthetic
data generation, the library adeptly identifies and categorizes PII within structured data formats.
The integration of Presidio and a BERT model in the identification process, along with the
usage of the Faker library for synthetic data generation, demonstrates a comprehensive and
innovative approach.</p>
      <p>The strength of NERPII lies in its ability to work with structured data, which is often
overlooked by other solutions that focus more on unstructured data such as texts. Additionally, with
NERPII, the privacy of personal data is guaranteed, without sacrificing the semantic component
of the data. Indeed, there are other solutions, like Presidio itself, that allow for the identification
of PII in texts and their anonymization using predefined or customizable tags. However, in our
view, this approach results in a significant loss of informative content in the data. On the other
hand, the use of synthetic data enables complete anonymization without any loss of meaning.</p>
      <p>A promising direction for the future development of this library is to adapt it to other
languages in addition to English, such as Italian, using Faker providers for the Italian language.
Currently, the library is designed to recognize Named Entities in structured data in English and
to regenerate PII in the American format (for example, if it needs to regenerate a Social Security
Number, it will do so according to the American nine-digit format). Moreover, it could be useful
to expand the number of Named Entities recognized by the NamedEntityRecognizer and those
regenerated by the FakerGenerator. Finally, a future development could examine the aspect of
coherence among diferent columns: in the current version of the library, data synthesis occurs
independently for each column, despite it being evident that some columns are semantically
correlated (for example, the column containing the city with the one containing the state).</p>
      <p>In conclusion, therefore, NERPII can represent a practical solution for addressing some of the
ethical issues related to data privacy in the field of artificial intelligence.
[11] D. Zhang, Y. Suhara, J. Li, M. Hulsebos, Ç. Demiralp, W.-C. Tan, Sato: Contextual semantic
type detection in tables, arXiv preprint arXiv:1911.06311 (2019).
[12] X. Deng, H. Sun, A. Lees, Y. Wu, C. Yu, Turl: Table understanding through representation
learning, ACM SIGMOD Record 51 (2022) 33–40.
[13] Y. Suhara, J. Li, Y. Li, D. Zhang, Ç. Demiralp, C. Chen, W.-C. Tan, Annotating columns
with pre-trained language models, in: Proceedings of the 2022 International Conference
on Management of Data, 2022, pp. 1493–1503.
[14] O. Mendels, C. Peled, N. Vaisman Levy, T. Rosenthal, L. Lahiani, et al., Microsoft Presidio:
Context aware, pluggable and customizable pii anonymization service for text and images,
2018. URL: https://microsoft.github.io/presidio.
[15] J. Devlin, M. Chang, K. Lee, K. Toutanova, BERT: pre-training of deep bidirectional
transformers for language understanding, CoRR abs/1810.04805 (2018). URL: http://arxiv.
org/abs/1810.04805. arXiv:1810.04805.
[16] D. Faraglia, Other Contributors, Faker, 2014. URL: https://github.com/joke2k/faker.
[17] P. M. Schwartz, D. J. Solove, The pii problem: Privacy and a new concept of personally
identifiable information, NYUL rev. 86 (2011) 1814.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bassignana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramponi</surname>
          </string-name>
          , Preface to the
          <source>Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI)</source>
          ,
          <source>in: Proceedings of the Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI</source>
          <year>2023</year>
          )
          <article-title>co-located with 22th International Conference of the Italian Association for Artificial Intelligence (AI* IA</article-title>
          <year>2023</year>
          ),
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Floridi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Taddeo</surname>
          </string-name>
          , What is data ethics?,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Shmatikov</surname>
          </string-name>
          ,
          <article-title>Myths and fallacies of" personally identifiable information"</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>53</volume>
          (
          <year>2010</year>
          )
          <fpage>24</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>European</given-names>
            <surname>Parliament</surname>
          </string-name>
          ,
          <article-title>Council of the European Union, Regulation (EU) 2016/679 of the European Parliament</article-title>
          and of the Council,
          <year>2016</year>
          . URL: https://data.europa.eu/eli/reg/2016/ 679/oj.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Centers</surname>
            <given-names>for Medicare</given-names>
          </string-name>
          &amp; Medicaid
          <string-name>
            <surname>Services</surname>
          </string-name>
          , Online at http://www.cms.hhs.gov/hipaa/,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Lample</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Subramanian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kawakami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyer</surname>
          </string-name>
          ,
          <article-title>Neural architectures for named entity recognition</article-title>
          ,
          <source>arXiv preprint arXiv:1603.01360</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nadeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sekine</surname>
          </string-name>
          ,
          <article-title>A survey of named entity recognition and classification</article-title>
          ,
          <source>Lingvisticae Investigationes</source>
          <volume>30</volume>
          (
          <year>2007</year>
          )
          <fpage>3</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>El Emam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mosquera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoptrof</surname>
          </string-name>
          ,
          <article-title>Practical synthetic data generation: balancing privacy and the broad availability of data,</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T. E.</given-names>
            <surname>Raghunathan</surname>
          </string-name>
          ,
          <article-title>Synthetic data</article-title>
          ,
          <source>Annual review of statistics and its application 8</source>
          (
          <year>2021</year>
          )
          <fpage>129</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hulsebos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bakker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Zgraggen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Satyanarayan</surname>
          </string-name>
          , T. Kraska, Ç. Demiralp,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hidalgo</surname>
          </string-name>
          ,
          <article-title>Sherlock: A deep learning approach to semantic data type detection</article-title>
          ,
          <source>in: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1500</fpage>
          -
          <lpage>1508</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>