<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GDPR Article Retrieval based on Domain-adaptive and Task-adaptive Legal Pre-trained Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Simeri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Tagarelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. Computer Engineering</institution>
          ,
          <addr-line>Modeling, Electronics, and Systems Engineering (DIMES)</addr-line>
          ,
          <institution>University of Calabria</institution>
          ,
          <addr-line>87036 Rende (CS)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>63</fpage>
      <lpage>76</lpage>
      <abstract>
        <p>The General Data Protection Regulation (GDPR) is an European regulation on data protection and privacy for all individuals within the European Union (EU) and the European Economic Area (EEA), and for all foreign subjects dealing with European citizens data. Therefore, the GDPR has important legislation implications that hold beyond EU member states. In this paper, we address the problem of GDPR article retrieval through the use of pre-trained language models (PLMs). Our approach features several key aspects, which include both domain-general and domain-specific pre-trained BERT models, further powered by self-supervised task-adaptive pre-training stages, with or without data enrichment based on recitals. Our study endeavors to demonstrate the potential of PLMs in addressing the challenges posed by the GDPR's intricate legal framework, thus ultimately facilitating eficient access to GDPR provisions for government agencies, law firms, legal professionals, and citizens alike.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;law article retrieval</kwd>
        <kwd>domain adaptation</kwd>
        <kwd>legal language models</kwd>
        <kwd>artificial intelligence and law</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>representing a valuable resource for determining the meaning and scope of the GDPR articles.
The recitals are 173 in total, each associated to one or more articles; in turn, each article is
associated with zero, one or more recitals.</p>
      <p>The GDPR hence encompasses a comprehensive set of norms and principles that regulate
the collection, processing, and transfer of personal data. Its provisions, such as the right to
be forgotten, consent requirements, and data subject rights, have brought about significant
changes in the digital landscape. Indeed, the complexity and scope of the GDPR pose challenges
for government agencies, law firms, legal professionals, and citizens seeking to navigate its
intricacies and access relevant regulations. Automating the search for GDPR information
is demanding to address the need for eficient access to the GDPR contents. By employing
advanced NLP technologies, the process of accessing and understanding GDPR provisions can
be greatly facilitated.</p>
      <p>In this regard, pre-trained language models (PLMs) such as BERT and GPT like models are
the most helpful and attractive tools, given their widely known remarkable capabilities in
various NLP tasks, including text classification, question answering, and document retrieval. In
particular, we notice that opting for BERT-like models over GPT-like models for GDPR article
retrieval ofers several key advantages. BERT-like models prioritize precision, accuracy, and
contextual understanding, mitigating the risks associated with hallucinations and ensuring
reliable interpretations of the GDPR contents. Their emphasis on pre-training and fine-tuning,
coupled with higher transparency and explainability than GPTs, make BERT-like models
wellsuited for addressing the complex task of article retrieval in the GDPR context, empowering
users to access and comprehend data protection regulations with confidence and accuracy.</p>
      <p>
        However, despite the potential benefits of leveraging PLMs for GDPR article retrieval, we
are not aware of PLMs that have been specifically trained on GDPR texts to date. In this work,
we aim to fill this gap by training BERT-based models for the task of GDPR article retrieval.
More specifically, we train and fine-tune a pool of BERT models for a sequence classification
task on the GDPR articles, with or without data enrichment based on the GDPR recitals. Our
selected BERT models include not only the general domain (i.e., base) BERT but also the legal
BERT models in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], particularly the from-scratch pre-trained and further pre-trained versions.
Furthermore, we originally propose two self-supervised task-adaptive pre-training strategies,
namely Related Sentence Prediction and Multiple Choice Answering, which show key advantages
on diferent query sets, which vary in terms of source, length, and lexical characteristics.
      </p>
      <p>By harnessing the power of PLMs’ contextual language understanding, we aim to provide
an eficient and efective means for government agencies, law firms, legal professionals, and
citizens to access and retrieve GDPR regulations. The outcomes of our research might contribute
to enhancing accessibility, comprehension, and application of the GDPR, benefiting a wide
range of stakeholders in their eforts to comply with data protection regulations and uphold
individuals’ rights.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>Most existing approaches have focused on checking the GDPR compliance of privacy policies.</p>
      <p>
        [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] proposes a conceptual model for characterizing the content of privacy policies in terms
of information elements that one can expect to find in them (e.g., controller’s identity and
contact). Based on named entity recognition and Glove word embeddings, such information
elements are extracted and used to train a SVM model for a task of multi-label classification.
The approach in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] distinguishes between coarse-grained and fine-grained practices based on
the OPP-15 taxonomy, and model them as a directed acyclic graph with a three level structure
(i.e.; categories, attributes and values) so that the extraction of data practices is treated as a
hierarchical multi-label classification task converted into two text-to-text tasks, one for each
level of the label hierarchy, based on a T5 model. The extracted information are fed into a
rule-based system that encodes the GDPR articles 13 and 14 under the supervision of legal
experts in accord with the OPP-115 taxonomy.
      </p>
      <p>
        [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] identifies four types of regulatory entities within policies of web services, which are
used to define 16 classes. A BiLSTM is trained for a multi-class classification task, and a BERT
summarizer is applied prior to the evaluation of context similarity for adhering vs. non-adhering
policies. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] exploits a supervised variant of Latent Dirichlet Allocation (i.e., Labeled LDA) to
model topics associated with various types of violations of GDPR articles within real privacy
incidents. The most relevant words associated with the topics induced by Labeled LDA are used
to augment the training instances from the CMS.Law GDPR Enforcement Tracker database1 to
learn an LSTM classifier at GDPR article level. [ 6] leverages FastText word embeddings and a
CNN model to predict privacy disclosure requirements according to GDPR articles 13 and 14.
[7] is concerned with compliance of data processing agreements (DPAs), i.e., legally binding
agreements that regulate the data processing activities according to GDPR. In relation to a
predefined set of 45 requirements extracted from the GDPR provisions relevant to DPA, the
proposed approach aims to assess whether a DPA is GDPR compliant, by comparing
semanticrole-based representations of DPAs against predefined representations of the requirements.
Also, the approach further provides recommendations about missing information in the DPA.
In the context of GDPR compliance concerning the Italian Public Administration, [8] proposes
a framework to detect security breaches related to unlawful disclosure of health information in
public documents. Personally identifiable information as named entities are extracted and used
to feed a machine learning classifier (e.g., SVM, XGBoost) for a binary classification task (i.e.,
compliant or non-compliant).
      </p>
      <p>Unlike our work, the above studies have made limited use of PLMs and only focused on
completeness or compliance/violation w.r.t. the GDPR; moreover, with regard to the latter
aspect, the requirements checking is often carried out only within few GDPR articles (e.g., 13
and 14), although other articles contain fundamental GDPR requirements as well. By contrast,
our work is the first to embrace the more general task of GDPR article retrieval by leveraging
PLMs, and also powering them through self-supervised task-adaptive pre-training schemes.
While our models can serve as a basis for further tasks like compliance checking — in fact,
we recognize it as ongoing work (cf. Conclusions) — they ofer a more general and versatile
solution to automate and ease access to GDPR.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Self-supervised task-adaptive pre-training strategies</title>
      <p>
        Task-adaptive fine-tuning is known to be the commonly used approach to enable a direct
application of a pre-trained model to a downstream task. On the other hand, domain-adaptive
pre-training allows for a more extensive customization of a pre-existing out-of-the-box model
to a specialized language domain. In the legal domain, two approaches to domain-adaptive
pretraining have been adopted: one is to continue pre-training the model using a legal corpus, while
the other is to start the pre-training process from scratch using a legal corpus. An exemplary
study adopting both approaches is provided in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for the development of the well-known
Legal-BERT models.
      </p>
      <p>However, the availability of legal data for a particular task may be limited, which can hinder
the efective training of the model. Consequently, the model may struggle to fully grasp the
meaning of legal texts and generalize the acquired knowledge in enough detail to handle
unknown inputs successfully. It has been demonstrated that pre-training a model on a legal
corpus does not always guarantee significant improvements over fine-tuning a corresponding
domain-general pre-trained model on the target task (e.g., [9]). According to [10, 11], the
advantages of domain-specific pre-training are particularly evident when dealing with
lowresource downstream tasks.</p>
      <p>Nonetheless, there exists another form of pre-training, which is to train a (pre-trained) model
on a smaller corpus of the specialized domain such that the corpus is directly related to the target
task but its documents are not annotated with the target class labels. This form of unsupervised
pre-training, called task-adaptive pre-training, has shown competitiveness in comparison to
domain-adaptive pre-training and can also enhance performance when combined with it for
the downstream task [11]. However, these findings have not been proven specifically for the
legal domain, presenting opportunities for further research in this area.</p>
      <p>In this work, we aim to fill the above gap by developing the first approaches to task-adaptive
pre-training of BERT models tailored to the GDPR article retrieval task. We shall describe our
two proposed approaches in the following sections.</p>
      <sec id="sec-3-1">
        <title>3.1. Related Sentence Prediction</title>
        <p>Besides Masked Language Modeling, BERT was unsupervisedly trained on another pre-training
objective, called Next Sentence Prediction (NSP), to cover a variety of downstream tasks
involving sentence pairs. Given two word sequences as input, the NSP task is to determine if the
second sequence is subsequent to the first in a document.</p>
        <p>Inspired by the idea underlying NSP, we propose the Related Sentence Prediction (RSP) approach
to task-adaptive pre-training. Basically, RSP is to predict if two given sentences in input are
related to each other or not; the relatedness concept can be defined in more ways, and in this
work we shall consider the location of two sentences within the same GDPR article.</p>
        <p>Let us denote with  and ℛ the sets of GDPR articles and recitals, respectively. Each article 
can be modeled as a sequence  = (,1, . . . , , ), where s denote textual units belonging to
, i.e., sentences – following an analogy with NSP – or, more generally, paragraphs constituting
. For any  ∈ , we define RSP(), with  ∈ {, ℛ}, as a meta-function expressing
relatedness between  and diferent portions of the GDPR, thus producing a set of training
instances as associations between textual units of  and textual units from the document
collection . This means that, depending on the choice of , sentences/paragraphs of  are
coupled with either sentences/paragraphs from other article(s) than , or with recitals. We
refer to the first case (i.e.,  = ) as article-level RSP, and to the second case (i.e.,  = ℛ) as
recital-level RSP. In both cases, two types of training instances are built, which are contrastive to
each other: the ones referring to “positive” relatedness (denoted with superscript +) and the
other ones referring to “negative” relatedness (denoted with superscript − ). We shall provide
our definitions next.</p>
        <p>Article-level RSP. For any article  ∈ , we define</p>
        <p>RSP() = RSP+(, ) ∪ RSP− (,  ∖ ),
(1)
where RSP+(, ) is a function producing a set of intra-article pairings over the textual units
of  as positive relatedness associations, and RSP− (,  ∖ ) is a function producing a set
of inter-article pairings between textual units of  and textual units from articles diferent
from , as negative relatedness associations. The above functions are specified so as to satisfy
the following minimum requirements. First, to ensure balance between positive and negative
associations, the number of training instances as pairs derived from , which is bounded
by ||(|| − 1)/2, is used to constrain the number of training instances derived by pairing
units from  and units from any other article  , with  ̸= . Second, the choice of  is, by
default, made uniformly at random, although constraints could be added to get  more or less
“topically distant” from  (e.g., selecting  from the same chapter of  or from a diferent
one). Third, multiple choices of  are made if | | &lt; || (i.e., multiple articles need to be
involved to form a number of negative associations to equal the number of positive ones); in
general, using multiple articles against  can be useful to diversify the negative associations
with , thus obtaining a mix of negative training instances at diferent hardness levels.
Recital-level RSP. Leveraging recitals for training a model on an RSP task is strongly justified
since they are originally conceived in the GDPR as essential complements for the articles. Based
upon this, we provide a function definition analogous to the article-level one, which is as follows:
RSPℛ() = RSP+(, ()) ∪ RSP− (, ℛ ∖ ()),
(2)
where () denotes the set of recitals associated with article , the positive relatedness function
RSP+(, ()) yields a set of training instances as pairs obtained by the textual units of 
and the recitals in (), and the negative relatedness function RSP− (, ℛ ∖ ()) yields a set
of training instances as pairs obtained by textual units of  and recitals not in (). It should
be noted that we specify associations between portions of articles and the entire recitals, as we
want to provide a maximal context based on recitals for each of the articles.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Multiple Choice Answering</title>
        <p>Our second proposed task-adaptive pre-training approach is Multiple Choice Answering (MCA),
which is defined as choosing the correct answer from a set of possible answers relating to an
input query. In a sense, this task can be seen as a blend between the RSP task and the Masked
Language Modeling task, which is at the core of BERT-like pre-training. In fact, the latter
requires to choose the correct word to fill in a mask from a set of possible options based on the
context of an input sentence. Analogously, MCA requires to choose from a set of sentences that
are presented as either positively or negatively related to an input one, in a similar fashion to
the RSP task. Also, we again consider the opportunity of enriching the training through the
recitals, therefore we shall distinguish between article-level MCA and recital-level MCA.
Article-level MCA. Given an integer  &gt; 1 as the number of choices, for any article  ∈ ,
we define</p>
        <p>||
MCA,() = ⋃︁ ⟨MCA+(, , ), MCA− (, ,  ∖ )⟩,</p>
        <p>=1
where MCA+(, , ) yields a pairing between , and another unit randomly chosen from 
(i.e., the positive or correct choice for , ), and MCA− (, ,  ∖ ) yields a ( − 1)-sized set
of pairings between , and units each selected from randomly chosen articles  , with  ̸= .
Overall, for each , a set of training instances is computed, where each training instance is a
tuple of size  + 1 (i.e., a sentence/paragraph and its relating  choices). Note that the positions
of the  choices are actually randomly shufled so that the position of the correct choice is
variable through all training instances.</p>
        <p>Recital-level MCA. By changing the answering choice context from articles to recitals, we
have the following definition:</p>
        <p>||
MCAℛ,(()) = ⋃︁ ⟨MCA+(, , ()), MCA− (, , ℛ ∖ ())⟩,</p>
        <p>=1
where () denotes the set of recitals associated with article , and the positive, resp. negative,
MCA functions have analogous definitions to the corresponding article-level ones. Note however
that, consistently with the recital-level RSP definition, recitals are considered as atomic units.
(3)
(4)</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Training and evaluation methodologies</title>
      <sec id="sec-4-1">
        <title>4.1. Model selection and settings</title>
        <p>
          Our study is versatile w.r.t. the choice of PLM to deal with the GDPR search and retrieval.
As happened in the past for other novel applications of PLMs, demonstration through BERT
(bert-base-uncased) [12] represents a primary choice. Moreover, this allows us to naturally
couple evaluation based on a domain-general BERT model with evaluation based on legal
specialized counterparts, which are the well-known family of models in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Specifically, we
use (i) the main model, named LegalBERT (legal-bert-base-uncased), which was
pretrained from scratch on large corpora including EU legislation, US contracts and cases, and
UK legislation; (ii) a EU specific model, named EULegalBERT (bert-base-uncased-eurlex),
which was further pre-trained starting from BERT base using EU legislation only.2
        </p>
        <p>Table 1 provides a summary of our developed models. Sufix -r is used to denote the
recitalenriched models, i.e., the recital-level RSP or MCA based models, as well as the base BERT,
2LEGAL-BERT models are available at https://huggingface.co/nlpaueb/legal-bert-base-uncased
LegalBERT, and EULegalBERT which were fine-tuned based on articles and recitals as the
training data. Note also that, in the fine-tuning stage, each of the models was trained for 10
epochs, using cross-entropy as loss function, AdamW optimizer and initial learning rate selected
within [1e-5, 5e-5] on batches of 256 examples.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Training data preparation</title>
        <p>Data enhancement. GDPR articles exhibit three distinctive traits in their logical structure.
First, the text of an article is commonly organized as a numbered sequence of paragraphs
(commas), each sometimes formatted as an enumerated list of points. Second, an article often
contains references to specific paragraphs or points of the same article as well as of other articles.
Third, one or more recitals can be associated with specific paragraphs, subparagraphs, or points
within the same article.</p>
        <p>Such features of the GDPR articles prompted us to carry out some preprocessing aimed to
enhance them for feeding a language model. In particular, we pursued a threefold goal: (i) to
enrich the article contents by expanding the references to (portions of) other articles or chapters;
(ii) to refactor structured parts of an article in order to resolve anaphoric passages; and (iii) to
produce a recital-based labeling of the articles at the finest level, by leveraging associations of
recitals to the individual paragraphs of an article, when available. While the latter required no
particular efort, the first two objectives were accomplished through a semi-automatic process
with manual supervision to produce a reliable outcome. Specifically, with regard to objective (i),
we resolved each reference by replacing it either with the original text of the referred part or
with an inferred short description, in the form of a citation (i.e., enclosed by quotation marks).
Concerning objective (ii), any enumerated list of points in an article was either replaced by as
many paragraphs as the number of points, each equally preceded by the common premise of
the point list, or just flattened by keeping the premise once followed by the enumeration, in
case of relatively short texts as points in the list.</p>
        <p>Fine-tuning article labeling schemes. Creating a training dataset for the downstream
task, i.e., GDPR article retrieval, requires that the entire corpus must be used to embed its
knowledge fully, and each class label must correspond to a specific GDPR article since we want
to learn how to classify at the article level. To this purpose, in order to build the training set
for the fine-tuning task, we resort to an unsupervised article-labeling scheme proposed in [ 13]
which is designed to select and combine portions of each article, while ensuring balance of
the contributions of each article, which clearly have diferent lengths. This scheme applies a
round-robin method to iterate over replicas of the same group of training instances per article
until a minimum number  of instances (to be produced for each article) is reached. In
this work, we used the unigram with parameterized emphasis on the title scheme [13] creates a
set of training instances for each article which is comprised of round-robin selected sentences
from the article, along with replicas of the article’s title; also,  was set to 64 instances.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Query sets</title>
        <p>We are not aware of any publicly available benchmark for evaluating retrieval models on GDPR
data. Therefore, we built our own query data as test sets, by varying them in terms of source,
length, and lexical characteristics. We define the following query sets: QA, which contains
sentences randomly extracted from the GDPR articles; QpA, where each sentence in QA is
paraphrased through an English-Spanish-Italian-English machine translation of the queries
(via Google Translate); QR and QpR, which are analogous of QA and QpA, respectively, but
replacing articles with recitals; QC, which contains expert commentary texts related to the
GDPR articles, i.e., a set of opinions and comments provided by experts in data protection and
privacy;3 QCs, where each comment in QC is broken down into its constituting sentences.</p>
        <p>Each of the QA and QpA, resp. QR and QpR, sets contains 661, resp. 138, queries, with an
average of 60 (± 35), resp. 112 (± 61), words per query. Also, QC, resp. QCs, contains 45, resp.
272, queries, with an average of 169 (± 52), resp. 28 (± 13), words per query.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Assessment criteria</title>
        <p>Each query is associated with one article (ground-truth). For each article , we first computed
the following statistics: the recall for  (), i.e., the number of queries s.t.  was correctly
predicted out of all queries actually pertinent to , the precision for  ( ), i.e., the number
of queries s.t.  was correctly predicted out of all predictions of , and the F-measure for
, i.e.,  = 2 /(  + ). Then, we averaged over all articles to obtain the
per-article average precision ( ), recall (), micro-averaged F-measure (  ) as the average
over all s, and macro-averaged F-measure (  ) as the harmonic mean of   and .</p>
        <p>In addition, we accounted for the top-3 predictions and the position (rank) of the correct
article in predictions: the former is the fraction of correct article labels that are found in the
top-3 predictions (i.e., top-3-probability results in response to each query), and averaging
over all queries, which is the recall@3 (@3); the latter is the mean reciprocal rank ( )
considering for each query the rank of the correct prediction over the classification probability
distribution, and averaging over all queries.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>We organize our presentation of the results into two parts: the first reporting the performance
of our models trained on articles only, and the second considering the recital-enriched models.
5.1. Training on articles only
 
0.985
0.990
0.946
0.929
0.951
0.812
0.596
0.658
0.451
0.550
0.662
0.291
0.390
0.419
0.262
0.590
0.806
0.399
 
 
 
Comparison of BERT models. Table 2 shows results obtained by the base BERT, LegalBERT,
and EULegalBERT on the various query sets.</p>
      <p>First, all models achieve high scores across all metrics over QA and QpA queries, indicating
their ability to accurately retrieve relevant information; clearly, this is not surprising since QA
and QpA queries contain information seen during the models’ training. More challenging are
the QR/QpR and QCs/QC queries which are based on contents from recitals and commentaries,
respectively. Generally, LegalBERT consistently outperforms BERT, which would indicate the
benefits of adapting the models to the legal domain prior to the GDPR article retrieval task.
However, EULegalBERT shows significantly lower performance compared to the other two
models. This performance gap should be ascribed to the fact that EULegalBERT results from a
further pre-training of BERT over EU specific resources, which limited knowledge expansion
w.r.t. that gained by LegalBERT over a much larger set of legal corpora in relation to BERT.
Impact of task-adaptive pre-training. Let us now focus on the impact of task-adaptive
pre-training on the GDPR article retrieval task, whose results are summarized in Tables 3–5; on
each of those tables, we also report the scores achieved by the corresponding model without
task-adaptive pre-training (cf. Table 2).</p>
      <p>Considering first BERT models (Table 3), we find that BERT-MCA achieves the highest scores
for most query types, and the advantage over the other two models are particularly evident for
the most dificult query sets. BERT-RSP generally performs better than the base BERT in terms
of precision, recall, and f-measures, with the exception of QA/QpA query sets, which would
indicate that task-adaptation based on the RSP pre-training task can even worsen performance
on article retrieval when the query content does not deviate much from the training data.
Moreover, BERT-RSP consistently falls behind BERT-MCA, which highlights the superiority of
the latter form of task-adaptive pre-training when applied to BERT.</p>
      <p>As reported in Table 4, task-adaptive pre-training of LegalBERT showcases quite diferent
trends from the BERT counterpart. While MCA maintains a significant advantage over RSP
especially for the commentary-based queries (i.e., QCs and QC), the same does not hold for the
other queries, especially the recital-based queries (i.e., QR and QpR). Also, on the latter query
types and on QC, both variants of task-adaptive pre-training are not able to improve performance
over the base LegalBERT. This would suggest MCA (and RSP) might not necessarily take
advantage when, as it is the case for LegalBERT, the base model has a pre-training knowledge
on a legal domain that would enclose the targeted one (i.e., GDPR in our setting).</p>
      <p>Table 5 shows how EULegalBERT appears to take advantage when task-adapted via RSP, on
query sets QA- QpR, or via MCA, on commentary-based queries. However, compared to the
previous results, EULegalBERT models are still outperformed by LegalBERT and BERT models
(apart from very few exceptions, such as precision on QCs and QC queries).</p>
      <p>Overall, our findings suggest that task-adaptive pre-training, especially based on MCA,
can yield better results when applied to base BERT and its legal pre-trained models, and this
particularly holds in terms of @3 and   criteria.</p>
      <sec id="sec-5-1">
        <title>5.2. Training with recital data enrichment</title>
        <p>Table 6 shows results corresponding to the recital-enriched models. For each query set and
criterion, the table reports the percentage variation (i.e., increase/decrease) achieved by a
recitalenriched model w.r.t. its corresponding non-recital-enriched model; moreover, the last row
of each query-set subtable shows the absolute best-performing model, considering both the
recital-enriched models and the previously analyzed models.</p>
        <p>First, we notice that while the improvements are negligible for the QA/QpA query types, some
significant increase in performance holds for the QCs/QC query types, particularly for the BERT
and EULegalBERT models. Also, it clearly does not come to our surprise that leveraging recitals
is beneficial for all domain-adaptive and task-adaptive pre-trained models when evaluated on
recital-based queries (i.e., QR and QpR). In such cases, the absolute best model is LegalBERT-r,
which also indicates that task-adaptive pre-training is not needed for the target task as its lack
can be well compensated by the recital data enrichment.</p>
        <p>More importantly, task-adaptive pre-training, especially based on MCA, reveals to be essential
to maximize performance in relation to the non-recital-based query sets. Moreover, a
combination of MCA and recital-enrichment leads to the absolute best model for the QCs type. This is
-0.28%
+0.69%
+0.28%
-0.58%
-0.28%
-0.20%
+2.69%
+5.04%
+4.72%
0.996
-5.91%
-0.72%
-3.08%
-4.29%
-0.96%
-1.59%
-7.21%
+14.61%
+14.79%
0.959
-0.15%
+0.46%
+0.31%</p>
        <p>0%
-0.15%
-0.30%
+2.67%
+3.14%
+3.46%
0.998*
-0.16%
+0.64%
+0.16%
+0.31%
+0.16%
-0.16%
-2.86%
+7.23%
+7.40%
0.982
+34.31%
+5.88%
+10.78%
+23.42%
-2.70%
+2.70%
+63.10%
+22.62%
+26.19%
0.993
+36.36%
+6.06%
+11.11%
+28.04%
+1.87%
+6.54%
+74.03%
+31.17%
+32.47%
0.993
-6.45%
-1.94%
+8.39%</p>
        <p>0%
+10.74%
+12.08%
-16.842
+64.21%
+65.26%
0.614
+8.57%
+2.86%
+8.57%
+2.50%
+2.50%
+2.50%
-22.58%
+22.58%
+12.90%
0.911*
another remarkable aspect since it supports our intuition of the beneficial efect of integrating
the complementary role of recitals for the GDPR articles into task-adaptive pre-training of
models to be fine-tuned for the article retrieval task.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>Summary. We addressed the problem of GDPR article retrieval through the use of PLMs, which
include both domain-general and domain-specific pre-trained BERT models, further powered
by self-supervised task-adaptive pre-training stages, with or without data enrichment based on
recitals. To the best of our knowledge, this is the first study that explores a number of aspects
concerning both domain-adaptive and task-adaptive legal pre-trained language models for the
task of GDPR article retrieval.</p>
      <p>
        Ongoing work. We are currently working on an evaluation of our proposed models on the
CMS.Law GDPR Enforcement Tracker database. Preliminary experimental results have shown
the efectiveness of our models, both in absolute terms and in relation to the article classification
approach proposed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] relying on labeled LDA and an LSTM model (cf. Related Work).
      </p>
      <p>As previously discussed, our models for GDPR article retrieval can serve as a basis for a variety
of similarity based tasks, including question-answering. Indeed, we are working on further
developments of our models to deal with GDPR compliance and violation checking tasks. In
particular, we are developing a hybrid framework based on a combination of BERT-like models
and ChatGPT-like models in order to take advantage of similarity search and classification
capabilities of the former and conversational functionality of the latter.
&amp; Cloud Computing, Sustainable Computing &amp; Communications, Social Computing &amp;
Networking (ISPA/BDCloud/SocialCom/SustainCom), IEEE, 2021, pp. 1617–1624. doi:10.
1109/ISPA-BDCloud-SocialCom-SustainCom52081.2021.00216.
[6] T. A. Rahat, M. Long, Y. Tian, Is your policy compliant?: A deep learning-based empirical
study of privacy policies’ compliance with GDPR, in: Proc. of the 21st Workshop on
Privacy in the Electronic Society (WPES 2022), ACM, 2022, pp. 89–102. doi:10.1145/
3559613.3563195.
[7] O. Amaral, M. I. Azeem, S. Abualhaija, L. C. Briand, NLP-based Automated Compliance
Checking of Data Processing Agreements against GDPR, CoRR abs/2209.09722 (2022).
doi:10.48550/arXiv.2209.09722.
[8] F. Lorè, P. Basile, A. Appice, M. de Gemmis, D. Malerba, G. Semeraro, An AI framework to
support decisions on GDPR compliance, Journal of Intelligent Information Systems (2023)
1–28.
[9] A. Simeri, A. Tagarelli, Exploring domain and task adaptation of LamBERTa models for
article retrieval on the Italian Civil Code, in: Proc. of the 19th Conference on Information
and Research science Connecting to Digital and Library science (IRCDL 2023), volume
3365 of CEUR Workshop Proceedings, 2023, pp. 130–143. URL: https://ceur-ws.org/Vol-3365/
paper4.pdf.
[10] S. Wang, M. Khabsa, H. Ma, To pretrain or not to pretrain: Examining the benefits of
pretraining on resource rich tasks, in: Proc. of the Annual Meeting of the Association for
Computational Linguistics (ACL), ACL, 2020, pp. 2209–2213.
[11] S. Gururangan, A. Marasovic, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, N. A. Smith,
Don’t stop pretraining: Adapt language models to domains and tasks, in: Proc. of the
Annual Meeting of the Association for Computational Linguistics (ACL), ACL, 2020, pp.
8342–8360.
[12] J. Devlin, M. Chang, K. Lee, K. Toutanova, BERT: pre-training of deep bidirectional
transformers for language understanding, in: Proc. of the 2019 Conference of the North
American Chapter of the Association for Computational Linguistics: Human Language
Technologies, (NAACL-HLT 2019), Association for Computational Linguistics, 2019, pp.
4171–4186.
[13] A. Tagarelli, A. Simeri, Unsupervised law article mining based on deep pre-trained language
representation models with application to the Italian civil code, Artif. Intell. Law 30(3)
(2022) 417–473. Published: 15 September 2021. doi:10.1007/s10506-021-09301-8.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Chalkidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fergadiotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Malakasiotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Aletras</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Androutsopoulos</surname>
          </string-name>
          , LEGAL-BERT:
          <article-title>The muppets straight out of law school, in: Findings of the Association for Computational Linguistics: EMNLP 2020, Association for Computational Linguistics</article-title>
          ,
          <year>2020</year>
          , pp.
          <fpage>2898</fpage>
          --
          <lpage>2904</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .findings-emnlp.
          <volume>261</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Torre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abualhaija</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sabetzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Briand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Baetens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Forastier</surname>
          </string-name>
          ,
          <article-title>An AI-assisted Approach for Checking the Completeness of Privacy Policies Against GDPR</article-title>
          ,
          <source>in: Proc. of the 28th IEEE International Requirements Engineering Conference (RE</source>
          <year>2020</year>
          ), IEEE,
          <year>2020</year>
          , pp.
          <fpage>136</fpage>
          -
          <lpage>146</lpage>
          . doi:
          <volume>10</volume>
          .1109/RE48521.
          <year>2020</year>
          .
          <volume>00025</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Hamdani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mustapha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Amariles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Troussel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Meeùs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Krasnashchok</surname>
          </string-name>
          ,
          <article-title>A combined rule-based and machine learning approach for automated GDPR compliance checking</article-title>
          ,
          <source>in: Proc. of the Eighteenth International Conference for Artificial Intelligence and Law (ICAIL</source>
          <year>2021</year>
          ), ACM,
          <year>2021</year>
          , pp.
          <fpage>40</fpage>
          -
          <lpage>49</lpage>
          . doi:
          <volume>10</volume>
          .1145/3462757.3466081.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Elluri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S. L.</given-names>
            <surname>Chukkapalli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Finin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <article-title>A BERT based approach to measure web services policies compliance with GDPR, IEEE Access 9 (</article-title>
          <year>2021</year>
          )
          <fpage>148004</fpage>
          -
          <lpage>148016</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2021</year>
          .
          <volume>3123950</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Aleroud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Masalha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Saifan</surname>
          </string-name>
          ,
          <article-title>Identifying GDPR privacy violations using an augmented LSTM: toward an ai-based violation alert systems</article-title>
          ,
          <source>in: Proc. of the IEEE International Conference on Parallel &amp; Distributed Processing with Applications</source>
          , Big Data
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>