<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Legal Summarization: to each Court its own model</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Flavia Achena</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Preti</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Venditti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leonardo Ranaldi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristina Giannone</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Massimo Zanzotto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Favalli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raniero Romagnoli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cognitive</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Almawave S.p.A.</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Università di Roma Tor Vergata</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the Italian Civil Law System, easily accessing legal judgments through massime is crucial. In this work, we compare extractive summarization models to produce massime in two Italian courts: the Constitutional Court and the Supreme Court. The aim of our study is to assess the efectiveness and eficiency of these models in summarizing the decisions of the two courts. Through a comprehensive analysis of two large datasets, we evaluate the quality of the summaries generated by each model and their ability to capture the key legal principles and linguistic features present in the courts' decisions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Legal Text Analysis</kwd>
        <kwd>BERT-based Summarization</kwd>
        <kwd>Legal NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>fying the portions of the text that contain the relevant
information to be reported in the massime becomes
chalIn civil and common law systems, accessing legal judg- lenging due to their length[5].
ments to retrieve legal decisions is crucial when lawyers Legal document summarization has seen rapid
have to defend clients, prosecutors have to build cases, progress in recent years, and several approaches[6, 7]
and judges have to draw decisions. To ensure widespread have been proposed to manage this kind of data, ranging
information on the decisions of the courts, in Italy, for from fine-tuned Transformer models on legal domain,
this purpose, a specific body drawn up massime. Reinforcement Learning to Generative Models[8].</p>
      <p>These massime present, in a short but detailed way, a In this paper, we compare extractive summarization
legal principle present in judgments. Hence, justice pro- models to produce massime in two diferent contexts:
fessionals can read these massime instead of the complete the Constitutional Court (Corte Costituzionale) and the
legal decisions. Supreme Court (Corte di Cassazione). We discuss
simi</p>
      <p>The process of analyzing judgments and extracting rel- larities and diferences about massime and how the kind
evant sentences can be significantly simplified through of court impacts the data in terms of their availability
the use of pre-trained models [1, 2], which serve as and privacy management. Then, we propose two
modversatile universal sentence/text encoders, capable of els tailored to the specific type of courts, discussing the
addressing various downstream tasks, including sum- approaches we implemented to circumvent the issue
remarization [3]. These models consistently outperform lated to lengthy documents. Results of the experiments
other approaches, especially after fine-tuning or domain- confirm that producing massime is a real challenge even
adaptation [4]. However, despite the success of pre- for dedicated systems. Hence, these systems should be
trained transformers in other summarization tasks, the designed as facilitators in a human-in-the-loop
environtask of producing massime is challenging for current ment [9].
extractive and abstractive summarization systems.
Unlike standard summaries, massime must follow rigorous
specifications in some courts. Extractive and abstractive 2. Diferent courts, Diferent
summarization datasets and relative systems, in contrast, Judgments, and Diferent
aim to reduce the size of a text while preserving its
overall meaning. However, this approach difers from the massime
specific requirements of creating massime.</p>
      <p>Additionally, legal texts are often extensive, further The data for the Italian legal domain have some
peculiarincreasing the summarization task’s complexity. Identi- ities that require careful consideration. Diferent courts,
such as the Constitutional and Supreme Court, produce
diferent judgments, leading to diferent massime.
Moreover, within the same court, there can be judgments with
varying numbers of related massime, ranging from one
to five or even more. In addition, the availability of data
depends on the presence or absence of sensitive
informaCLiC-it 2023: 9th Italian Conference on Computational Linguistics,
Nov 30 — Dec 02, 2023, Venice, Italy
$ [first-initial].[surname]@almawave.it (F. Achena);
[name].[surname]@uniroma2.it (D. Venditti)</p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAtstrhiboutpionP4.r0 oIncteerneadtioinnagl(sCC(BCYE4.U0).R-WS.org)
tion within the judgments, so the legal courts provide 2.2. Availability of judgments and
access to the data in diferent ways. massime in the two Courts</p>
      <p>Both Constitutional and Supreme Court courts share
that producing a massima requires a relevant cognitive A fundamental element that afects data availability
contask carried out by the "Massimario Ofice" body, as it cerns personal information and privacy. In cases where
involves identifying the principle of law present in the judgments contain sensitive personal information, access
judgment and satisfying some precise criteria for its writ- to such data is restricted due to privacy protection laws.
ing. The following are the characteristics of the two The Italian Constitutional Court, since it is central to
Italian Courts under consideration (Sec. 2.1), the nature the defense of the Constitution, prioritizes the
availabilof their judgments and massime, how a massima is struc- ity of data on its proceedings and decisions. The Court
tured (Sec. 2.3), and, finally, a comparative analysis of must ensure the integrity and adherence to constitutional
the two types of corpora that can be derived from these principles within the legal system. As a result, inquiries
two courts (Sec. 3) is provided. made to the Court generally focus on broad issues that
do not involve specific individuals. Consequently,
judgments do not contain any personal information and are
2.1. Two Italian Courts: Constitutional not subject to privacy-related restrictions. Data of the
Court and Supreme Court Italian Constitutional Court are thus open and accessible
through its portal1.</p>
      <p>On the other hand, the Supreme Court deals with
cases that may involve specific physical or juridic
people, which requires compliance with privacy regulations.</p>
      <p>Consequently, access to its data must be restricted.
Information on the proceedings and decisions of the Supreme
Court is only accessible through the Italgiure platform2,
which is exclusively available to professionals and legal
practitioners. Data cannot be shared, and accesses are
controlled and logged. Currently, the dataset selected
for the Supreme Court cannot be made public because
it would require an expensive anonymization process to
ensure privacy.</p>
      <p>The Italian Constitutional Court has the primary
responsibility to assess the constitutionality of the acts and
laws of the State and the Regions. Among other functions,
it assesses charges against the President of the Republic
in accordance with constitutional provisions. The Court
examines the admissibility of abrogative referendums. To
ensure impartiality and independence, the Constitutional
Court is composed of 15 lawyers, chosen from among
judges, law professors, or lawyers with at least 20 years
of experience.</p>
      <p>The Supreme Court - also known as the "Corte di
Cassazione" - is the highest authority in the Italian
judicial system. It serves as the court of final appeal and
has two main functions. Firstly, it resolves judicial
conlficts to determine which judge has jurisdiction over a
case. Secondly, it has a nomophylactic function, ensuring
that the law is interpreted uniformly. Within the Court,
there is the "Massimario Ofice", responsible for
identifying nomophylactic judgments and producing concise
summaries called massime. These summaries contain the
legal principles from the Court’s judgments, not just a
summary of the cases themselves. The primary objective
of the Massimario Ofice is to disseminate legal
knowledge and facilitate comprehension of past court decisions.</p>
      <p>To accomplish this, the ofice updates its collection by
incorporating new judgments, ensuring access to the
most current precedents. This results in a large number
of judgments and massime so in the vision of making
the judicial system more eficient by digitising court
proceedings, providing automatized support to the processes
can reduce the time and efort required to analyze and
summarize them.</p>
      <sec id="sec-1-1">
        <title>2.3. The shape of massime</title>
        <p>Each legal judgment (also called decision), despite its
individuality in terms of case and subject matter, has a
shared overall structure. This structure comprises the
following key components:
• Heading/Epigrafe: It is the initial part containing
the indication of the members of the court, the
details of the initiating document, the reporting
judge, and the attorneys heard by the Court.
• Statement of Facts: Summarizes the relevant facts
of the case, often introduced by "considered in
fact and in law.".
• Reasons: It is the section where the Court
provides an explanation or argumentation for the
conclusions reached in the judgment. This
section typically presents the legal principles, factual
analysis, and logical reasoning that support the
Court’s decision.
• Ratio Decidendi: Establishes the binding legal
principle or rule derived from the court’s decision.</p>
        <sec id="sec-1-1-1">
          <title>1https://dati.cortecostituzionale.it</title>
          <p>2https://www.italgiure.giustizia.it/
• Disposition: Concludes the decision with the
final ruling and any related orders or remedies. It
is often introduced by "P.Q.M."3: It contains the
determination of the judges.</p>
        </sec>
        <sec id="sec-1-1-2">
          <title>Similar to the decisions, the creation of summaries of</title>
          <p>legal principles, commonly known as massime, follows
well-defined summarization criteria. As outlined in [ 10],
these massime must contain explicit legal references and
embody the fundamental principles of law. This detailed
approach ensures the efective spread of legal knowledge.
Massime must meet the following requirements:
• Faithfulness to the decision.
• Conciseness in stating the legal principle.</p>
          <p>• Clarity and precision of the stated principle.</p>
        </sec>
        <sec id="sec-1-1-3">
          <title>Hence, massima represents the expression of the legal</title>
          <p>principle and must not be considered a summary of the
decision.</p>
        </sec>
        <sec id="sec-1-1-4">
          <title>According to our analysis, massime of the Constitutional</title>
          <p>3. The datasets of judgments and Court became more extractive after 2000. Indeed, in that
massime period, it seems that Constitutional Court Judges forced
the "Massimario Ofice" to avoid changing the text
ex3.1. Analysis of the massime of the two tracted from judgments because even a small change of
Courts a single word could significantly alter the overall
concept expressed in the judgment. As a result, since 2000,
To better understand how to develop a system for mas- the process of producing massime become an extractive
sime generation, we analyzed the correlation between summarization task guided by a topic presented in the
the judgments and the massime of both courts as they last part of the judgment.
have diferent roles and consequently deliver diferent
judgments. 3.2. Producing massime as a classification</p>
          <p>For the Supreme Court, we selected a subset of judg- task
ments, from 2010 to 2013, to build our dataset useful for
the extractive summarization task. During our analysis, Summarization is an inherently abstractive task.
Howwe noticed that some decisions may include a massima ever, it can be treated as an extractive classification task
without any text or expressed with an abbreviation, such once the target summary (i.e., a massima) is used to
seas "CONFORME A CASSAZIONE ASN: ...". For these cases, lect the relevant or irrelevant sentences from the starting
we interpret them as references to previous massime, but document (i.e., a judgment).
decline to use these specific examples. We started by
selecting only the judgments corresponding to at least one 3.2.1. Supreme Court Extractive Data-set
massima. Indeed, we observed that while most legal
judgments of the Supreme Court are tied to a single massima, As mentioned before, the first step of the extractive model
there are a sizable amount of cases in which multiple used to deal with the Supreme Court dataset (see Sec. 2.1)
massime refers to the same judgment (see Tab. 1). Details consists in rephrasing a generic abstractive
summarizaabout how we handled such cases are discussed in the tion dataset into something suitable for a (classical)
clasnext subsection. sification model. This is achieved via the introduction of</p>
          <p>In addition to analyzing judgments from the Supreme a Oracle meta-model[11].</p>
          <p>Court, we also conducted a systematic analysis of Italian For each pair (document, summary), all the sentences
Constitutional Court judgments from 1956 to 2021. We forming the set with the highest F1 Rouge [12]
combinaaligned sentences in massime with sentences in the judg- tion 1 + 2 concerning the summary are selected and
ments in order to understand how sentences in massime annotated as relevant, while all the others are
automatiare diferent from those in the judgments (see Figure 1). cally identified as irrelevant. This automatically frames</p>
        </sec>
        <sec id="sec-1-1-5">
          <title>3for these reasons</title>
          <p>massima per judgment
1
2
3
4
5+
65%
24%
6%
3%
2%
(, ) ⏞→⏟ (, )
process ensured a balanced distribution between positive
and negative examples, maintaining a 50/50 ratio. It is
important to note that the specific details and steps of the
method used to derive the larger dataset from the original
14,316 rows are not provided in this paper. However,
this method facilitated a focused analysis of the textual
components, shedding light on the connections between
phrases, device points, and the formation of massime.
with  = 0, 1. As shown in Tab. 1, when a
judgment is related with more than one massima, the
Oracle model acts independently on each judgment-maxima
couple, and then the annotated sentences are merged
together without repetitions. This is done because
otherwise, it can most likely happen that in a multiple massima 4. Models
scenario, the same sentence in a judgment is related only
with one massima, ending up with the same sentence Our main challenge was identifying the most relevant
annotated with opposite categories. parts of pronouncements to assist the massima producer</p>
          <p>Starting from a dataset corresponding to 12000 couples in crafting legal maxims.
of (, ) we decided to keep only the As mentioned before, extractive summarization
moddata corresponding to at an Oracle rouge of 1 + 2 ≥ els treat the task of automatic summarization generation
0.55, reducing our training data almost by half (6.849). as a straightforward sentence classification task. In this
We observed that, given the nature of the judgment, the vision, the summary of a given document emerges by the
number of relevant sentences in any judgment is a very concatenation of all the most relevant document
fragsmall portion, inevitably producing a highly unbalanced ments (i.e., sentences or sub-sentences) classified by the
dataset toward the irrelevant sentences. model, this could efectively provide the Massimario with
the essential subparts of pronouncements for massima</p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>3.2.2. Constitutional Court Extractive Data-set construction.</title>
        <p>Both models proposed in the current work are
essenGiven the analysis of the massime and the related judg- tially based on a deep encoder which maps the fragments
ments of the Constitutional Court, we decided to define to a vector representation in a high dimensional space
the task of producing massime as the classification task subsequently classified into two classes: relevant
senof selecting the appropriate sentences of the judgment tences (i.e., candidates for the summary) or irrelevant
given a target topic. The classification dataset is then sentences (i.e., not containing relevant information for
built as follows, starting from the judgments and the the summary).
related massime. For each judgment, we extracted the
points of its operative part (punti del dispositivo). For
each point, we selected the correlated massima. Then, 4.1. Supreme Court Model
we divided the judgments into sentences and produced a
set of triples:</p>
        <p>(, , _)
where  is a sentence of the judgment,  is
a point of its operative part, and _ is True if
the sentence overlaps for more than 90% with a sentence
in the massima related to the .</p>
        <p>For our experiments, we extracted a subset of 40,000
data points from this expanded dataset. The selection</p>
        <sec id="sec-1-2-1">
          <title>Data-sets with very long documents (as the one intro</title>
          <p>duced in Sec. 2.1) are usually dificult to handle using a
BERT-based [2] transformer encoder. The well known
self-attention (introduced in [13]) which characterizes
most of the transformer networks is plagued by a fast
scaling of computational and memory requirements with
the input sequence. Instead of proceeding with a more
memory-eficient attention implementation (for instance,
see [14]), we decided to act on data and restrict the
context length. In this perspective, we introduced a fixed
Court</p>
          <p>Supreme
Constitutional</p>
          <p>Prec</p>
        </sec>
        <sec id="sec-1-2-2">
          <title>We observe that the model does not seem to reach</title>
          <p>high performances both in terms of absolute and relative
scores, normalized with the Oracle Rouge scores, which
are the maximum scores such a model can achieve (see
Tab. 2). This is partially due to the violent class unbalance
present in the dataset, even if marginally mitigated by the
introduction of a weighted loss, with weights inversely
proportional to the category frequency in the train set. As
a baseline comparison, we decided to include the scores
normalized with the one of a random classifier to assess
that no random classifications are being performed.</p>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>4.2. Constitutional Court Model</title>
        <sec id="sec-1-3-1">
          <title>The challenge lay in selecting the most useful subparts</title>
          <p>of legal judgments for the purpose of the massima
prolength sliding window similar to the implementation de- ducer. We sought to leverage BERT[2] to provide the
scribed in [15]). best subparts of pronouncements to aid in massima
con</p>
          <p>Our goal was to optimize the context length and miti- struction. However, we soon realized that the task was
gate context-truncation efects. To achieve this, we de- exceptionally complex, requiring the ability to
summaifned the context window based on word-pieces and in- rize and generalize the text in a unique manner.
troduced the possibility of overlapping windows up to To address the above multifaceted challenge, we
foa maximum number of sentences. The latter, while still cused on using BERT to assist us in identifying the most
under investigation, ofers an intriguing tool to probe relevant subparts of legal judgment. Through this
apthe context efects on the model. Even in the simplest proach, we aimed to equip the massima producer with
implementation, with only one overlapping sentence (see essential tools for constructing the maxim more
efecFig. 2), it is interesting to see the efect of a preceding or tively. While our eforts resulted in the development
subsequent context on the same sentence. of a tool to assist the massima, producer, we must
ac</p>
          <p>The model used, with the aforementioned modifica- knowledge that the results achieved with BERT were
tion in the document pre-processing, is based on the one not as remarkable as initially hoped. The complexity of
proposed in [16, 3] referred to as BERTSUM4. It is worth the problem, combining the tasks of summarization and
mentioning that while in their works, the predicted prob- generalization in a unique manner, presented formidable
ability is used only as a ranking score and a fixed number hurdles.
of sentences are extracted neglecting the actual probabil- Nevertheless, we view this endeavor as a stepping
ities, we select the relevant sentences accordingly to a stone toward understanding and tackling the intricacies
cutof parameter. The latter is a necessary introduction of legal text processing. Our tool, despite its limitations,
since we observed that, in our data-set, it is common to serves as a valuable resource for the Massimario, aiding
have documents with a clear separation between sen- them in the maxim construction process. We recognize
tences that are very likely to be extracted compared to that further research and advancements in natural
lanothers whose predicted probability is minimal. Therefore, guage processing will be crucial in making substantial
ifxing the number of extracted sentences introduces a strides in this domain.
strong bias toward the extraction of irrelevant sentences. Even in this case, results are interesting but not yet
satisfactory (see Tab. 2). Indeed, R1, R2, and R3 are 0.32,</p>
        </sec>
        <sec id="sec-1-3-2">
          <title>4sometimes named BertSumExt in literature.</title>
          <p>0.29, and 0.24 respectively. This suggests that the task of
producing massime is indeed a challenging task.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>5. Conclusions</title>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgments</title>
      <sec id="sec-3-1">
        <title>This work was conducted within the DATALAKE Gius</title>
        <p>tizia project; we acknowledge the partners and the
scientific committee for their support.</p>
        <p>In conclusion, while we did not achieve outstanding
results, our eforts shed light on the intricacies of this
challenging problem. During our analysis, we also noticed
notable diferences between the two courts, which
further emphasizes the complexity of generating accurate
massime. We find it particularly intriguing to explore the
factors contributing to these variations and understand
how they impact the summarization process. Despite
the challenges, we remain committed to refining our
approach and exploring innovative techniques. Recent
advances in the field further motivate us to seek a proper
solution that addresses data privacy concerns and
significantly improves the task of summarization in the legal
ifeld for the Italian language.
based transformer architectures for long document
summarization, in: Proceedings of the 16th
Conference of the European Chapter of the Association
for Computational Linguistics: Main Volume,
Association for Computational Linguistics, Online, 2021,
pp. 1792–1810. URL: https://aclanthology.org/2021.
eacl-main.154. doi:10.18653/v1/2021.eacl-m
ain.154.
[16] Y. Liu, Fine-tune bert for extractive summarization,
2019. URL: https://arxiv.org/abs/1903.10318. doi:10
.48550/ARXIV.1903.10318.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>