<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Topic Similarities in Rights and Duties across European Constitutions using Transformer-based Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Candida M. Greco</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Tagarelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. Computer Engineering</institution>
          ,
          <addr-line>Modeling, Electronics, and Systems Engineering (DIMES)</addr-line>
          ,
          <institution>University of Calabria</institution>
          ,
          <addr-line>87036 Rende (CS)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>47</fpage>
      <lpage>62</lpage>
      <abstract>
        <p>The use of language models in the legal NLP field has brought significant advances in the use of AI systems to support legal professionals. However, most of the eforts so far have focused on processing documents such as legal cases, contracts and statutes. There are several types of legal resources that are still overlooked, and these include constitutions. A constitution establishes the basic principles, structures, functions and powers of a country's governance. Several portions of the constitutions are devoted to rights and duties of the citizens, which are essential to define and protect the status of citizens as individuals and as members of the society. To this regard, in this work we focus on the range of topics covered in the European constitutions that guarantee rights and duties to citizens. We present the first study providing lexical and semantic similarity analysis of the European constitutions, which especially takes advantage of using several Transformer-based models.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;European constitutions</kwd>
        <kwd>topic similarity</kwd>
        <kwd>legal language models</kwd>
        <kwd>artificial intelligence and law</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        A constitution is a fundamental document that serves as the supreme law of a country or an
organization. It establishes the basic principles, rights, and rules that govern the functioning
of the entity it applies to. The legal domain is currently one of the major fields of application
of AI techniques for supporting experts in the analysis of documents, comprising mostly legal
cases, contracts and statutes. Surprisingly, the current literature on the application of AI-based
NLP to the processing and understanding of constitutions is quite limited. In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the authors
carry out a comparative analysis between US constitution and a number of constitutions of four
geographic regions (Africa, Asia, Europe, and Middle East), with the aim of detecting diferences
and similarities w.r.t. rights of citizens and the relationship between the major institutions of
their governments; however, the study employs basic techniques used in text mining, discarding
any reference to context-free as well as contextualized deep language models. Moreover, no
insights into the European countries’ constitutions are provided.
      </p>
      <p>In this paper, we aim to fill this gap in the literature by conducting a similarity analysis of
the European countries’ constitutions with a focus on the rights and duties of citizens, which are
essential to define and protect the status of citizens as individuals and as members of the society.
Our study is motivated by the opportunity of unveiling commonalities and diferences in the
constitutions of several European countries as a similarity search problem. In particular, we
pursue two main research objectives: (i) understanding how much European countries agree (or
difer) w.r.t. specific topics within the realm of citizen rights and duties and (ii) how pre-trained
language models are able to discern the diferences between topics, especially when they are
conceptually related. To achieve this, we employ methods that involve lexical and semantic
analysis of the text, with a specific focus on Transformer-based language models.</p>
      <p>
        We believe our work can pave the way for further developments on AI-based solutions for
NLP tasks involving the constitutions. This is supported by the evidence that similarity search
is broadly employed in the legal AI as an essential means to address more complex tasks, such
as statutory article retrieval [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], legal case retrieval [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], document review [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], legal judgement
prediction [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], summarization [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and many others.
      </p>
      <p>
        In addition, the use of Transformer-based language models for solving legal tasks is a
prevailing trend in recent years. For instance, in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] such models are trained on a topic similarity task
to predict the coherence among topics and to detect topical changes on legal texts. In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], legal
and general-purpose Transformer models are compared for legal document recommendation
addressed as a similarity task. Nonetheless, our study is the first to analyze the semantics of
European constitutions by leveraging Transformer-based language models.
Plan of this paper. The subsequent sections of the paper are organized as follows. In Section
2 we give a brief overview of the Transformer-based models used in this study. Section 3
provides a description of the dataset we built for supporting our study. In Section 4 we outline
the specific objectives in our work. In Section 5 we discuss our experimental evaluation and
provide an analysis of the achieved outcomes, whereas in Section 6 we summarize the main
ifndings of our analysis, as well as discussing limitations and future perspectives.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background on Transformers</title>
      <p>
        Transformer models [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] have emerged as a dominant paradigm for designing outstanding deep
learning models that have revolutionized the state-of-the-art in a wide range of challenging
Natural Language Processing (NLP) tasks. The fundamental aspect of the Transformer
architecture is the incorporation of attention mechanisms [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which encompass all hidden states
of a neural network at a time and assign suitable weights to capture the inter-dependencies
among words. Currently, the Transformer paradigm is widely adopted in NLP, with a significant
portion of state-of-the-art NLP models built upon this architecture.
      </p>
      <p>
        In this work, we focus on a set of Transformer-based language models (TLMs) that is
representative according to two key dichotomic aspects in our study: domain-generality vs.
domainspecificity, suitability to similarity search tasks vs. task generality. This has led us to select BERT
as domain-general model, legal-BERT s as domain-specific models, and Sentence-Transformers as
models designed for similarity search tasks. In the following, we recall main characteristics of
such models.
BERT. BERT [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is widely recognized as the pioneering TLM that has revolutionized natural
language understanding. The model’s advantages encompass bidirectional unsupervised
pretraining and a unified architecture that adeptly addresses a range of tasks. The bidirectionality
is achieved through the Masked Language Modeling (MLM) task, which involves predicting
masked input words from unlabeled text while considering both left and right context words.
Moreover, BERT is pre-trained using the Next Sentence Prediction (NSP) task, which aims to
determine if one sequence follows another in a given text. Over the years, BERT has played
as a catalyst for extensive research, which provided several variants and enhancements of the
model. This has culminated in a broad range of BERT-based models.
      </p>
      <p>
        Legal-BERT. BERT and BERT-based models are primarily designed for general domains.
However, their performance tends to degrade when applied to specific domains such as the legal
one. To address this limitation, several strategies have been adopted to adapt the Transformer
models to the legal domain. The main approaches include further pre-training the model or
conducting pre-training from scratch using a legal corpus. Chalkidis et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] were the first
to propose both further pre-training and pre-training from scratch BERT on legal corpora,
including EU and UK legislations, cases from the European Court of Justice (ECJ), cases from
the European Court of Human Rights (ECHR), US court cases and US contracts. In particular,
they developed Legal-BERT-FP models, obtained through further pre-training of BERT on
diferent sizes of the training legal corpora, and Legal-BERT-SC model, the result of training
BERT from scratch specifically for the legal domain.
      </p>
      <p>Sentence-Transformers. S-BERT [13] is a variant of BERT that has been specifically tailored
for tasks involving semantic textual similarity, clustering, and information retrieval through
semantic search. The model employs a fine-tuned siamese network architecture, wherein two
separate pre-trained BERT models, each dedicated to one input sentence, share tied weights that
are updated during fine-tuning. The siamese architecture enhances the generation of sentence
embeddings that encode meaningful semantic information. S-BERT is the core model from
which a series of Sentence-Transformers have been developed over the years, representing the
state-of-the-art for sentence embeddings.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>Data collection and structure. We retrieved the texts of the European constitutions from
the portal www.constituteproject.org [14], which provides free and public access to
constitutions from various countries around the world. The resources for the website come from the
Comparative Constitutions Project.</p>
      <p>The documents in the Comparative Constitutions Project were originally labeled according to
a number of topics they identify on the constitutions. The topics are organized on three levels,
whereby the labels correspond to the most specific topics (i.e., the third-level topics). Hence,
given a label, it is possible to identify the topic hierarchy to which it belongs. We use the same
labeling system to structure our dataset for the similarity analysis at multiple levels of depth.
Specifically, we narrowed down on the first-level topic “rights and duties” and on the European
countries. We therefore conducted our similarity analysis of texts associated with third-level
topics (hereinafter micro-topics, for short) and second-level topics (hereinafter macro-topics, for
short) referring to rights and duties.</p>
      <p>Table 1 shows the macro-topics and an overview of the corresponding micro-topics. Overall,
European constitutions encompass 9 macro-topics and 111 micro-topics. Note that the same
portion of a constitution can be assigned to multiple micro-topics, but also the same
microtopic can be associated to several, not necessarily contiguous parts of a constitution, in which
case the texts are concatenated. In summary, the resulting dataset consists of text portions of
constitutions, with each portion being hierarchically assigned the country name, the micro-topic
and the macro-topic.</p>
      <p>Data cleaning and chunking. Each text corresponding to a particular combination of
country and micro-topic may span from one sentence to multiple parts of a constitution.
Transformers generally have a maximum limit on the number of tokens they can process. To
handle this, we divide the text into chunks so that the input does not exceed the 512 tokens
limit imposed by BERT and BERT-based models. The chunking process was carried out so
as to keep as many sentences together as possible and ensuring not to break sentences and
paragraphs. In any case, just a few instances ended up to exceed 512 tokens and the splitting
consisted of mostly 2 or 3 chunks. When possible, a long portion was chunked on the basis
of the constitution’ structure (e.g., if the portion is the concatenation of parts from diferent
articles, the subdivision was carried out keeping together the sentences from the same article).
The chunks were associated with the same country, micro-topic and macro-topic. In general,
an instance of the dataset corresponds to one country, but if the text was divided into chunks,
there is one instance per chunk. The resulting dataset size is about 2580 instances.</p>
      <p>Moreover, a step of anonymization was carried out to debias the lexical analysis from specific
terms, while preserving the essential meaning in the sentences. In particular, we introduced
generic identifiers to replace particular occurrences in the text, such as: names of persons
(e.g., royals, secretaries who drew up documents) with “person”, inhabitants (e.g., Italians) with
“European people”, countries (e.g., Italy) with “geo-political European entity”, locations (e.g., the
Athos peninsula mentioned in the Greek constitution) with “location”, organizations (e.g., the
United Nations mentioned in the Croatia constitution) with “organization”. Analogously, all
legal references were replaced with a special token (law_ref ).</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <sec id="sec-4-1">
        <title>Lexical vs Semantic similarity across European countries. Firstly, we assess lexical and</title>
        <p>semantic diferences among European countries on the parts of constitutions that share the
same micro-topic. This is useful for understanding how much European countries agree (or
difer) with respect to a specific theme of interest.</p>
        <p>Multi-level semantic similarity analysis. A significant portion of our research eforts has
been devoted to an extensive and multi-faceted examination of the similarity between texts
extracted from European constitutions. The aim is to assess how well Transformer-based models
can capture the closeness of texts that share the same topic, but more importantly how well they
can discern diferences between texts that deal with diferent topics (albeit discussing rights
and duties) and whether and to what extent they are able to detect subtle nuances of texts from
diferent but similar topics. In particular, we compare our selected Transformer-based models
to conduct the following analysis tasks:
• Topic similarity of texts having the same micro-topic: the input corresponds to the
instances concerning the same topic. The purpose is to compare the language models to
assess the ability to detect similarities between countries.
• Topic similarity of texts having the same macro-topic: the input consists of all the
instances that are associated with micro-topics belonging to the same macro-topic. The
purpose is to compare the various language models to assess the ability to discern similar
but not identical topics.
• Topic similarity of texts across all the macro-topics: the input consists of all portions
of the countries dealing with micro-topics from all macro-topics. The purpose is to
conduct an overall assessment on the generic topic of rights and duties. Specifically, given
a micro-topic as a query, we evaluate the ability of the models to assign higher similarity
scores to instances related to the query and, conversely, to assign lower scores with respect
to instances related to other micro-topics. Ideally, the lower scores given to instances
from the other micro-topics should still reflect whether or not they belong to the same
macro-topic of the query, i.e., the scores given to instances from the same macro-topic
should be higher than the scores given to instances related to other macro-topics.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental evaluation</title>
      <sec id="sec-5-1">
        <title>5.1. Settings</title>
        <p>Topic similarity is measured as cosine similarity throughout all the conducted tests. For the
lexical analysis, the term-frequency inverse-document-frequency term relevance function (TF-IDF)
is adopted to get the document embeddings. In this case, since the sparse vectorial space
poses no limit to the input length, the text was not divided into chunks. On the other hand,
lemmatization1 and stemming2 operations were performed as pre-processing steps.
1https://spacy.io/api/lemmatizer
2https://www.nltk.org/howto/stem.html</p>
        <p>
          For the semantic analysis, the text embeddings are obtained by each of our models, using
the following implementations. bert-base-uncased is selected as the domain-general BERT.
The selected legal-specific models are developed by [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and available on Huggingface,3 namely
• legal-bert-base-uncased, a BERT model pre-trained from scratch on legal corpora,
which is referred to as legal-bert-sc in the original paper;
• legal-bert-500k, a BERT model further pre-trained on legal corpora, which is referred
to as legal-bert-fp in the original paper;
• bert-base-uncased-echr, which is legal-bert-fp fine-tuned on ECHR cases;
• bert-base-uncased-eurlex, which is legal-bert-fp fine-tuned on EurLex. 4
We did not consider other available legal models in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] since they are specific to U.S. law,
while our study is focused on European constitutions. Using all the aforementioned models, we
obtain the sentence embeddings applying mean pooling strategy on top of the contextualized
token embeddings.
        </p>
        <p>Finally, we consider sentence-Transformer models since they are highly applicable to
similarityrelated tasks. Currently, several models are available through the sentence-transformers
library5 and they are ranked according to the quality of sentence embeddings, based on the
performances achieved on diferent tasks and domains. 6. Based on the ranking, we chose models
that achieve good performance, but at the same time have a manageable size and a maximum
length of 512 tokens. We therefore opted for the following models:
• gtr-t5-large,7 based on T5 [15] and fine-tuned for semantic search,
• all-mpnet-base-v1,8 based on MPNet model [16] and fine-tuned on diferent
usecases,
• all-distilroberta-v1,9 based on a distilled RoBERTa model [17] and fine-tuned on
diferent use-cases.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Lexical vs Semantic similarity</title>
        <p>We first compute and analyze heatmaps of the similarity scores between TF-IDF embeddings
produced for texts related to the same micro-topic. Each heatmap entry refers to an instance
of the dataset dealing with the selected micro-topic and corresponds to a particular country.
Note that, in all heatmaps shown throughout this paper, lighter colors correspond to higher
similarity scores.</p>
        <p>Figure 1 shows representative examples corresponding to selected micro-topics; for the sake
of readability, we removed the country labels from the heatmaps. We notice diferent situations
depending on the micro-topic, but in general, there is a low similarity among countries, as it is
evident in, e.g., Figure 1a. In some cases, a number of countries show strong lexical similarity,
3https://huggingface.co/nlpaueb/legal-bert-base-uncased
4https://eur-lex.europa.eu/
5https://www.sbert.net/index.html
6https://www.sbert.net/docs/pretrained_models.html
7https://huggingface.co/sentence-transformers/gtr-t5-large
8https://huggingface.co/sentence-transformers/all-mpnet-base-v1
9https://huggingface.co/distilroberta-base
(a) micro-topic “Limits on
(b) micro-topic “Right to
(c) micro-topic “Prohibition
employment of children"
renounce citizenship"
of torture"
such as in Figure 1b and 1c in which there are peaks of high scores and also peaks of maximum
similarity. In Table 2, we report some examples of the most similar texts according to the
similarity scores w.r.t. the micro-topics discussed in Figure 1. We can notice that for the
microtopics “Right to renounce citizenship” and “Prohibition of torture” the texts associated with the
highest scores are structurally similar or almost identical.
and pairwise similarity of semantic-based embeddings (generated by all-distilroberta-v1),
corresponding to the micro-topics “Limits on employment of children” and “Right to renounce
citizenship”. It can be noticed that the lexical and the semantic heatmaps have markedly diferent
scores, with the former having significantly lower scores than the latter. In particular, in Figure
2(a-b), we can notice that all-distilroberta-v1 reveals some similarity matches between
countries which are absent in TF-IDF. We show some examples in Table 3. In Figure 2(c-d), the
heatmaps have a similar shape but, again, the semantic model associates significantly higher</p>
        <p>(b) all-distilroberta-v1
(c) TF-IDF</p>
        <p>(d) all-distilroberta-v1
all-distilroberta-v1 but dissimilar based on TF-IDF embeddings</p>
        <p>Moldova (Republic of) 1994 (rev. 2016) “All employees shall have the right to social protection of labour. The protecting
measures shall bear upon the labour safety and hygiene, working conditions for women and young people, the introduction
of a minimum wage per economy, week-ends and annual paid leave, as well as dificult working conditions and other specific
situations. The exploitation of minors and their involvement in activities, which might be injurious to their health,
moral conduct, or endanger their life or proper development shall be forbidden.”
Montenegro 2007 (rev. 2013): “Youth, women and the disabled shall enjoy special protection at work.</p>
        <p>A child shall be guaranteed special protection from psychological, physical, economic and any other exploitation or abuse.”
Albania 1998 (rev. 2016): “Every child has the right to be protected from violence, ill treatment, exploitation and use for work,
especially under the minimum age for work, which could damage their health and morals or endanger their life or normal
development.”
Croatia 1991 (rev. 2013): “Children may not be employed before reaching the legally determined age, nor may they be forced
or allowed to do work which is harmful to their health or morality.”
scores than the lexical model. By comparing the two models, it can be inferred that strong
similarities are captured by both, but the semantic model is able to detect more adequately the
common focus of the texts. High scores are, indeed, expected since the texts discuss the same
micro-topic. By contrast, the lexical model often has very low scores even on texts of the same
micro-topic, consequently it also fails to diferentiate texts of the same micro-topic from texts
of diferent micro-topics.
(b) legal-bert-500k
(c) legal-base-uncased
(e) legal-bert-eurlex (f) all-distilroberta-v1 (g) all-mpnet-base-v1</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Results at micro-topic level</title>
        <p>As the lexical model revealed to be unable to adequately detect similarities between texts
about the same micro-topic, hereinafter we focus on semantic similarity only, in the attempt of
identifying the best model in capturing commonalities and diferences between micro-topics.</p>
        <p>A first step is at the micro-topic level, that is, evaluating which model is able to associate
the highest scores for texts belonging to the same micro-topic. Figure 3 shows heatmaps
of all the considered Transformer-based models (BERT, Legal-BERT models, and
SentenceTransformers) for the micro-topic “Duty to pay taxes”. In general, they all provide high
scores, but legal-bert-base-uncased exhibits extremely high similarities. Among all,
legal-bert-echr and the Sentence-Transformers provide some slightly lower scores. The
models exhibit a similar pattern across all micro-topics; we omit the heatmaps due to space
limitation of this paper, nonetheless, in Table 4 we provide an overview of the models’ behavior.
More precisely, for each macro-topic, we provide the average values of the mean, minimum,
maximum, and median similarity scores calculated on the same micro-topics. For instance, for the
macro-topic “Physical Integrity Rights”, bert-base-uncased provides, on average, a mean
similarity score of 0.834 on texts related to the same micro-topic, while all-distilroberta-v1
provides an average maximum similarity score of 0.878 for the macro-topic “Civil and
Political Rights”. We observe that legal models have the highest values and the smallest range
between the minimum and maximum scores for most macro-topics. This may be due to an
over-specialization knowledge of legal domain, leading to high scores for texts having strong
Statistics on similarity scores within the same micro-topic, for each macro-topic.
legal concepts in common, or conversely to a poor ability to recognize diferences.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Results at macro-topic level</title>
        <p>Considering the aforementioned behavior of the models on each macro-topic, we explore
whether the high scores of legal models can be ascribed either to a robust understanding of the
micro-topic or to a potential inability to discern subtle semantic distinctions.</p>
        <p>For this purpose, we assess whether diferent micro-topics belonging to the same macro-topic
are indistinguishable, that is, whether the models can distinguish related but not identical topics.
In Figure 4, we show the heatmaps corresponding to the various micro-topics belonging to the
macro-topic “Economic Rights”. Note that each heatmap reports both the similarities between
texts belonging to the same micro-topic (which are concentrated on the main diagonal) and the
similarities between texts belonging to diferent micro-topics. For the sake of readability, we
replace the names of the countries with letters corresponding to their respective micro-topics;
when an explicit label is missing, it is inferred that the entry of the heatmap is associated to the
preceding label.</p>
        <p>It can be noticed that legal models (and even bert-base-uncased) are unable to perceive
diferent degrees of similarity, which should be higher for texts of the same micro-topic and
lower between texts of diferent micro-topics. On the contrary, the Sentence-Transformers
(particularly all-distilroberta-v1 and all-mpnet-base-v1) are able to distinguish the
diferent micro-topics much more clearly. However, there are micro-topics that are nearly
indistinguishable even for the Sentence-Transformers. Examining these challenging
microtopics, we observe that they often involve remarkably similar concepts. For example, in Figure
4g, there is a clear dificulty in distinguishing the micro-topics
, , and  , which, however,
have in common the aspect of addressing matters pertaining to property rights. Similarly, the
micro-topic  and  are practically indiscernible even for the Sentence-Transformers.</p>
        <p>The above is also evident in the boxplots in Figure 5. In this case as well, the micro-topics 
and  encompass a closely related concept, namely aspects related to business. Once again, the
behavior of the models is consistent across all macro-topics, with all-distilroberta-v1
and all-mpnet-base-v1 showing the best results. To provide an example, Figure 6 shows the
(b) legal-bert-base-500k (c) legal-bert-base-uncased
(e) legal-bert-eurlex
(f) all-distilroberta-v1
(g) all-mpnet-base-v1
boxplots of all-mpnet-base-v1 and legal-bert-uncased for the micro-topic “Prohibition
of slavery” against all the micro-topics of its macro-topic (“Physical Integrity Rights”). It can be
observed that, in the case of all-mpnet-base-v1, the boxplot corresponding to the texts of
the micro-topic under examination (the first one from the left) has a higher mean compared to
the other boxplots, which represent the similarity scores of texts from the “Prohibition of slavery”
topic compared to texts from other micro-topics of its macro-topic. On the contrary, the boxplots
of legal-bert-uncased demonstrate that the model does not perceive substantial diferences
between texts with diferent micro-topics compared to texts with the same micro-topic.</p>
      </sec>
      <sec id="sec-5-5">
        <title>5.5. Overall results on Rights and Duties</title>
        <p>Another crucial aspect concerns the models’ capability to distinguish between micro-topics that
belong to the same macro-topic, as opposed to micro-topics from diferent macro-topics. Given
a micro-topic as a query, we shall compute similarity scores for all micro-topics. This can be
seen as an overall evaluation of the Rights and Duties topic.</p>
        <p>Our analysis can be grouped into three categories: (1) similarity between texts of the same
micro-topic as the query, (2) similarity between texts of diferent micro-topics but within the
same macro-topic as the one associated to the query micro-topic, and (3) similarity between
texts with both a diferent micro-topic and macro-topic w.r.t. the query. The expected behavior
for the models is to assign high similarity scores for the first category and low scores for the
other two categories, but the scores associated with the second category should be higher
compared to the scores associated with the third category.</p>
        <p>Table 5 summarizes the analysis conducted on all micro-topics, providing an overall mean
value across the three categories. For instance, bert-base-uncased obtains, on average, a
mean similarity score of 0.853 on texts related to the same micro-topic, a mean similarity score
of 0.799 on texts related to diferent micro-topics but having the same macro-topic, and a mean
similarity score of 0.762 on texts related to diferent micro-topics and diferent macro-topics.
The diference between the values of the first and second categories (column | −  |), is of
0.053, for the second and the third categories (column | −  ′|) is of 0.037, and for the first
and the third categories (column | −  ′|) is of 0.091. Once again, it can be observed that the
generic bert-base-uncased and the legal models are the least efective in distinguishing
between diferent micro-topics. Even gtr-t5-large demonstrates limitations in this regard,
particularly in distinguishing between the second and third category. Among all the models,
all-mpnet-v1 achieves the most favorable results, demonstrating the largest disparity in all
scenarios (columns | −  |, | −  ′| and | −  ′|), followed by all-distil-roberta.</p>
        <p>As an illustrative example, we show the boxplots of all models for the micro-topic
“Prohibition of slavery” in Figure 7. The diferences across the boxplots are more evident with
all-mpnet-base-v1. Indeed, the first boxplot (related to the first category) exhibits the
highest and most uniform values, the second boxplot (related to the second category) is suficiently
lower than the first one but higher than the others (related to the third category).</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>The experimental results unveil a number of major findings. In general, we have found that
European constitutions share many topics in the context of Rights and Duties. Despite for most
micro-topics the lexical similarities are generally very low, it also happens that few pairs of
European countries apparently address the same micro-topic in a similar manner; however,
lexical analysis is not suficient to capture the similarities between countries at a fine grain. On
the other hand, the semantic analysis unveils quite diferent behavior of the language models. In
particular, it is evident that the legal models we consider are not directly applicable to similarity
tasks, whereas the Sentence-Transformers demonstrate significantly better results, despite not
being specifically trained on legal corpora. This is not actually surprising, since the constitutions
are not written in a highly technical legal language, as they should be easily comprehensible
even for non-experts. Secondly, the Sentence-Transformers are trained to generate sentence
(b) legal-bert-fp
(c) legal-bert-sc
(e) legal-bert-eurlex (f) all-distilroberta-v1 (g) all-mpnet-base-v1
embeddings, which capture the semantic meaning and context of the texts rather than relying
on specific legal terminology. As a result, the Sentence-Transformers can efectively capture
the similarities and nuances of the constitutions. Among them, all-mpnet-base-v1 has
proven to be the best performer, although all-distilroberta-v1 follows closely behind;
by contrast, gtr-t5-large performs significantly worse, likely due to its specialization in
semantic search and lack of fine-tuning for other use cases.</p>
      <p>This study has some limitations that leave room for future improvement. We notice that
when the topics are highly similar, all the models faced dificulties in perceiving their diferences.
A fine-tuning phase on the constitution data would help recognize subtle nuances in meaning.
Also, we are aware that there are many other legal models that could have been considered in the
experimentation, and their inclusion could have provided further insights into real performances
of legal architectures on similarity tasks. However, we opted to focus on a selected set of models
that are widely recognized and representative of the current state-of-the-art in the field. Future
research may include exploring a broader range of legal models to gain a comprehensive
overview of their capabilities and limitations in detecting topic similarities among constitutions.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>In this work, we presented a lexical and a semantic similarity analysis among the segments of
European constitutions. We investigated the ability of legal and Sentence-Transformer models
to recognize texts that share the same topic and to diferentiate texts that cover diferent topics.
We conducted a multi-faceted experimental evaluation and provided an analysis of the achieved
outcomes, highlighting main findings, limitations and further perspectives.
doi:10.18653/v1/2020.findings-emnlp.261.
[13] N. Reimers, I. Gurevych, Sentence-BERT: Sentence Embeddings using Siamese
BERTNetworks, in: Proc. of the 2019 Conference on Empirical Methods in Natural Language
Processing (EMNLP 2019), Association for Computational Linguistics, 2019.
[14] Z. Elkins, T. Ginsburg, J. Melton, R. Shafer, J. F. Sequeda, D. P. Miranker, Constitute: The
world’s constitutions to read, search, and compare, J. Web Semant. 27-28 (2014) 10–18.
[15] C. Rafel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu,
Exploring the limits of transfer learning with a unified text-to-text transformer, J. Mach.</p>
      <p>Learn. Res. 21 (2020) 140:1–140:67.
[16] K. Song, X. Tan, T. Qin, J. Lu, T. Liu, MPNet: Masked and Permuted Pre-training for
Language Understanding, in: Proc. of the Annual Conference on Neural Information
Processing Systems (NeurIPS 2020), 2020.
[17] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V.
Stoyanov, Roberta: A robustly optimized BERT pretraining approach, CoRR abs/1907.11692
(2019).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bayrak</surname>
          </string-name>
          ,
          <article-title>A comparative analysis of the world's constitutions: a text mining approach</article-title>
          ,
          <source>Soc. Netw. Anal. Min</source>
          .
          <volume>12</volume>
          (
          <year>2022</year>
          )
          <fpage>26</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Louis</surname>
          </string-name>
          , G. van Dijck,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Spanakis, Finding the Law: Enhancing Statutory Article Retrieval via Graph Neural Networks, in: Proc. of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL</article-title>
          <year>2023</year>
          ), Association for Computational Linguistics,
          <year>2023</year>
          , pp.
          <fpage>2753</fpage>
          -
          <lpage>2768</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          , W. Ma,
          <string-name>
            <given-names>K.</given-names>
            <surname>Satoh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ma, BERT-PLI:
          <article-title>modeling paragraphlevel interactions for legal case retrieval</article-title>
          ,
          <source>in: Proc. of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI</source>
          <year>2020</year>
          ),
          <year>2020</year>
          , pp.
          <fpage>3501</fpage>
          -
          <lpage>3507</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Shaghaghian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Y.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Jafarpour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pogrebnyakov</surname>
          </string-name>
          ,
          <article-title>Customizing contextualized language models for legal document reviews</article-title>
          ,
          <source>in: Proc. of the IEEE International Conference on Big Data (IEEE BigData</source>
          <year>2020</year>
          ), IEEE,
          <year>2020</year>
          , pp.
          <fpage>2139</fpage>
          -
          <lpage>2148</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Aletras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tsarapatsanis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Preotiuc-Pietro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lampos</surname>
          </string-name>
          ,
          <article-title>Predicting judicial decisions of the European Court of Human Rights: a Natural Language Processing perspective</article-title>
          ,
          <source>PeerJ Comput. Sci. 2</source>
          (
          <year>2016</year>
          )
          <article-title>e93</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Shukla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation, in: Proc. of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th</article-title>
          <source>International Joint Conference on Natural Language Processing (AACL/IJCNLP</source>
          <year>2022</year>
          ), Association for Computational Linguistics,
          <year>2022</year>
          , pp.
          <fpage>1048</fpage>
          -
          <lpage>1064</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Aumiller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Almasian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lackner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gertz</surname>
          </string-name>
          ,
          <article-title>Structural text segmentation of legal documents</article-title>
          ,
          <source>in: Proc. of the Eighteenth International Conference for Artificial Intelligence and Law (ICAIL</source>
          <year>2021</year>
          ), ACM,
          <year>2021</year>
          , pp.
          <fpage>2</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ostendorf</surname>
          </string-name>
          , E. Ash,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ruas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Rehm, Evaluating document representations for content-based legal literature recommendations</article-title>
          ,
          <source>in: Proc. of the Eighteenth International Conference for Artificial Intelligence and Law (ICAIL</source>
          <year>2021</year>
          ), ACM,
          <year>2021</year>
          , pp.
          <fpage>109</fpage>
          -
          <lpage>118</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS</source>
          <year>2017</year>
          ),
          <year>2017</year>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Bengio,</surname>
          </string-name>
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          ,
          <source>in: Proc. of the 3rd International Conference on Learning Representations (ICLR</source>
          <year>2015</year>
          ),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proc. of the</source>
          <year>2019</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, (NAACL-HLT</article-title>
          <year>2019</year>
          ), Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>I.</given-names>
            <surname>Chalkidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fergadiotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Malakasiotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Aletras</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Androutsopoulos</surname>
          </string-name>
          , LEGAL-BERT:
          <article-title>The muppets straight out of law school, in: Findings of the Association for Computational Linguistics: EMNLP 2020, Association for Computational Linguistics</article-title>
          ,
          <year>2020</year>
          , pp.
          <fpage>2898</fpage>
          --
          <lpage>2904</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>