<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Semantic Document Management System for Public Administration</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Carlo Batini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gaetano Santucci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Palmonari</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio Bellandi</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elisabetta Fersini</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Barbara Pernici</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Zanzotto</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giancarlo Vecchi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Ronchi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Consorzio Interuniversitario Nazionale di Informatica (CINI)</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Polytechnic University of Milan</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Università degli Studi di Milano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Milan-Bicocca</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Rome Tor Vergata</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>To deliver services to users, central and local Public Administrations (PA) make extensive use of data. Various qualitative estimates suggest that databases contain 10-20 This work has two objectives: to summarize the experiences carried out over the past four years by the National Interuniversity Consortium for Informatics (CINI) in the Datalake project funded by the CRUI in collaboration with the Directorate General of Automated Information Systems (DGSIA) of the Ministry of Justice, in synergy with other related projects of the Ministry; and to demonstrate how the experiences, Proof of Concepts, and functional specifications produced can serve as a repository of functionalities for a “semantic document management system for PA,” which aims to evolve the information systems of PAs into platforms where unstructured data can be exploited and integrated with structured data to enhance and add value to the digital services provided by the PA, and where governance processes can be conducted using all knowledge expressed in documents and other forms of unstructured data. The judicial organization, proceedings, processes, user needs, functional structure of the Datalake, and implementation architecture are described, aiming towards a design and production pathway directed at all PAs.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Semantic Document Management</kwd>
        <kwd>Data Lake</kwd>
        <kwd>Legal AI</kwd>
        <kwd>Civil Trials</kwd>
        <kwd>Criminal Trials</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Proceedings, Trials,</title>
    </sec>
    <sec id="sec-2">
      <title>Organization, Justice</title>
    </sec>
    <sec id="sec-3">
      <title>Information Systems</title>
      <p>Ital-IA 2024: 4th National Conference on Artificial Intelligence,
organized by CINI, May 29-30, 2024, Naples, Italy
* Corresponding author.
† ChatGPT was used to translate original content written by the
authors in Italian; the authors have read and revised the translation,
ultimately agreeing on the final content.
$ carlo.batini@unimib.it (C. Batini); matteo.palmonari@unimib.it
(M. Palmonari); valerio.bellandi@unimi.it (V. Bellandi)
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License "Casellario Giudiziale") and databases of the Department
Attribution 4.0 International (CC BY 4.0).
of Penitentiary Administration and the Department of
Juvenile and Community Justice being the main
realizations. The Datalake project was initiated in 2019 by the
DGSIA, which entrusted the CRUI and, subsequently, the
Consorzio Interuniversitario Nazionale per l’Informatica
(CINI), with a renewed line of research over the years,
investigating the adoption of technologies based on natural
language processing (NLP), knowledge graphs, machine
learning, and, recently, generative AI, which can be most
useful in carrying out the primary processes of civil and
criminal cognition and execution. In 2022, Justice
included Datalake among the projects funded with PNRR
funds, whose implementation was entrusted through a
tender to a group of three companies: Almaviva,
Almawave, and Accenture.</p>
    </sec>
    <sec id="sec-4">
      <title>2. User Needs in Civil and</title>
    </sec>
    <sec id="sec-5">
      <title>Criminal Trials</title>
      <p>The Datalake project focused on the preliminary
investigation phase of criminal proceedings, criminal
enforcement, and the cognition phase of civil proceedings.</p>
      <p>The functionalities developed originate from a user
requirements elicitation activity, which for the criminal
procedure involves Prosecutors and the Judicial Police,
and for the civil procedure involves Judges. The
requirements were collected in the preliminary investigation
phase, through the sharing of Proof of Concepts on
deifned procedures, and subsequently through
consultations on ongoing procedures. In the civil domain,
interviews were conducted with Judges of the Court of Appeal
of Milan, who will soon experiment with a first set of
functionalities developed by the supplier Almawave. The
outcome of the experiment will result in a new version
of the system that can be adopted in all Courts of Appeal.
The user needs are briefly described below.</p>
      <sec id="sec-5-1">
        <title>Primary Activities – Preliminary</title>
      </sec>
      <sec id="sec-5-2">
        <title>Investigations in Criminal Proceedings</title>
        <p>Primary Activities – Civil Proceedings
• Enrichment and exploration of legal knowledge
during the process.
• Semantic search for nominal entities, mentions,
concepts, terms, phrases, and simple sentences, to
analyze seriality and search for precedent cases.
• Search for relevant judgments and case law with
a single integrated access point.
• Selection of judgments concerning topics (e.g.,
damages for privacy violations) not included in
the system metadata.
• Decision support in the cognition phase of the
judgment and linking relevant acts and
documents with the judgment.
• Extraction from civil judgments of the outcome of
the process (e.g., damage compensation,
maintenance allowance in judicial divorce proceedings)
and correlated salient features.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Governance Activities in the PCT (On-line Civil Trial)</title>
        <p>• Predictive models of the expected durations and
variability of the proceedings based on their
characteristics.
• Assessment of the expected complexity of the
proceedings for the distribution of workload among
judges, identification of "bottlenecks",
identification of signals and events that significantly
impact the duration of the proceedings, and analysis
of the impact of changes in laws, regulations, and
practices.
• Descriptive statistics on structured data and
judg</p>
        <p>ments.
• Correlation analysis between salient
characteristics and outcomes in civil proceedings for
uniformity purposes (so-called "tabulation").
• Comparative analysis of trial durations in
diferent sections and districts.
• Specific searches and semantic aggregations for 3. Functionalities
the discovery and confirmation of clues and
evidence. The Proof of Concept developed and the functional
spec• Integrated analysis of relational knowledge with ifications produced within the Datalake project concern
visualization. the following macro-functionalities: Preparation,
Seman• Selection of node clusters in a semantic graph tic Enrichment and Knowledge Integration, Semantic
with certain properties. Search and Analysis, Knowledge Base Management, and
• Reconstruction of relationships maintained by Quality Control. The following are the detailed
functionsuspected individuals, composing the entire rela- alities.</p>
        <p>tional network of the suspect.
• Transcriptions and semantic enrichment of audio F1 - Preparation
messages.
1. Document pre-processing (removal of special
characters, correction of accented letters, removal
of headers, removal of stamps, punctuation
management)
2. OCR and generation of interpretable documents.
3. Identification of sections of the judgments:</p>
        <p>preamble, case description, and decision.</p>
        <p>4. Classification of texts within the files.</p>
      </sec>
      <sec id="sec-5-4">
        <title>F2 - Semantic Enrichment and Knowledge</title>
      </sec>
      <sec id="sec-5-5">
        <title>Integration</title>
        <p>
          Semantic enrichment is performed by extracting
information from documents, especially named entities and
terms, and persisting the result of this extraction process
into semantic annotations. This process is in use in
legal AI to a large extent. The peculiar characteristic of
the proposed approach lies in the efort to consolidate
the knowledge extracted by linking diferent mentions
that refer to the same entities (exploiting background
knowledge bases like Wikipedia and clustering mentions
of the entities - of course, the majority - that are not
present in Wikipedia) [
          <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
          ]. The impact of this
approach is particularly noticeable during document search
(see functionality F3).
        </p>
        <p>
          Civil Trials. Various NLP techniques have been
applied to extract, link, and consolidate entity mentions
from judgments and produce semantic annotations that
associate the extracted entities with specific token
sequences in the judgments. In particular, the current
pipeline combines the following techniques [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ]:
• Named Entity Recognition (NER): Utilizes
rulebased and neural approaches, tuned to the data
distribution in the domain (sequential classifiers
on features from a BERT-based encoding
transformer).
• Named Entity Linking (NEL): Based on the BLINK
entity retrieval algorithm trained on the Italian
Wikipedia within the project.
• NIL Prediction: Decides whether to link an entity
mention to the entity associated with it by NEL or
label it as a new entity not present in the
knowledge base (NIL); for this task, an internal classifier
based on features is used. To perform NEL and
NIL prediction at once, an extended named entity
disambiguation algorithm has also recently been
explored to predict NIL as a class.
• NIL Clustering: Groups entity mentions referring
to the same real-world entities (typically applied
to mentions labeled as NIL because entities linked
to a knowledge base are implicitly grouped).
• Entity Registry Construction: The Entity Registry
is a component where each entity, enriched with
attributes deduced during the linking phase,
corresponds to a unique entry, avoiding duplicates
and disambiguating homonyms and synonyms.
Text annotations are updated with entity
identiifers.
• Refinement of Decisions: Final decisions made at
the end of the pipeline are refined based on some
domain-specific rules (especially for the
classification of specific and fine-grained entities).
• Relation Extraction: Extraction of relationships
such as victim-ofender relationship, or based
on the expression “against”. We used
pretrained transformer models for text
representation, with training conducted according to a
cross-validation policy and an extraction model
based on the entity-relationship paradigm and
REBEL [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
• Features &amp; values Extraction: Aims to extract
values associated with features (e.g., the economic
value of maintenance payments). Two available
open-source models, Camoscio and Stambecco
(versions of LLAMA trained on the English
language and adapted for the Italian language), and
the pay-per-use model known as ChatGPT were
considered. Techniques based on prompt
engineering were experimented with, using the
following types of prompts: Direct Instruction
Prompts, Contextual Prompts, Bridging Prompts,
Socratic Prompts.
• Few-shot Fine-grained Entity Typing:
Assignment of specific types from taxonomies to entity
mentions. We used a neuro-symbolic method,
where the taxonomy is explicitly modeled, and a
method based on LLM with implicit prompts.
        </p>
        <p>
          Criminal Trials - Preliminary Investigations. For
documents related to preliminary investigations, a very
similar pipeline was applied for entity extraction and
subsequent document annotation, a similar semantic search
paradigm. A first discussion of the application of
entitycentric approaches to manage documents in preliminary
investigations can be found in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. However, other
functionalities and techniques were applied such as:
• Extraction of graph representations from instant
messaging applications (IMA) data, e.g.,
WhatsApp dumps, and storage in a graph DB (Neo4J);
messages can be queried using a structured
language that supports graph-based data analysis.
• Content enrichment with speech-to-text
technology; OpenAI’s Whisper was used to transcribe
audio messages and make these contents
searchable. All messages and chats are analyzed using
small adaptations of the NLP pipeline described
earlier, supporting semantic search powered by
entity-based annotations.
• Semantic enrichment and specialization of entity
annotation ontologies relative to specific
taxonomy (is-overlapping, is-within, ordering).
embeddings and retrieving the relevant ones for
a user’s question.
• Document explorer: Allows exploring a document,
such as a judgment, guiding the search within it
for specific entities or mentioned concepts.
• Annotation editor: Allows modifying annotations
to support a supervised annotation process where
users can correct wrong or imprecise annotations
and add new annotations.
• Concept search: Allows searching or exploring
concepts according to domain logic. This
module can be useful to help the user select specific
concepts of interest in an exploratory or search
refinement phase.
        </p>
        <p>
          Other developed functionalities include domain
concept extraction, text summarization, and georeferencing
of spatial entities. For all functionalities, accuracy
analyses were conducted based on scientific methodologies.
For the entity extraction pipeline, some results are
reported in [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ]. As examples of accuracy measured for
relation and feature extraction capability, we report
accuracy for the Relationship “against”, 83.5%, and for the
extraction of the maintenance payment in favor of
children in separation cases, 77.52%.
        </p>
      </sec>
      <sec id="sec-5-6">
        <title>F3 - Semantic Search and Data Analysis</title>
        <p>Common Search Functionalities for Preliminary
Investigations and Civil Proceedings. Search
functionalities are inspired by the well-known faceted and
semantic search paradigms, with additional and more
experimental Question Answering (QA) capabilities based
on the Retrieval Augmented Generation (RAG) paradigm.</p>
        <p>Based on the semantic enrichment functionalities shown
in the previous point, the entities that appear in the filters
during the search phase can refer to mentions present
in diferent documents; moreover, when a user explores,
for example, a judgment, they can find all mentions of
an entity throughout the document, a feature that can
become particularly relevant for long judgments or other
documents. The conceptual architecture for semantic
search is shown in Fig 1. The components are:
The above functionalities have all been demonstrated
using DAVE, a prototype open-source application for
semantic search developed in the context of this and the
PON Next Generation UPP 1 project. A video
demonstrating the proposed combination of semantic and
conversational search on judgments of criminal trials published
online is available at https://www.youtube.com/watch?
v=XG7RsI3t-2Q. However, the data enrichment process
developed in the project supports also other forms of
search, such as Advanced search. This functionally
supports advanced searches by combining various filters on
document attributes. This module is included in many
search applications on structured or semi-structured data,
to complement the modules based on Keyword search
and Faceted search; typically, the function of this
module is to construct precise queries based on structured
descriptions of documents.</p>
        <p>
          Analysis Functionalities for Governance
Activities - Civil. The semantic organization of documents
obtained through semantic enrichment and integration
functionalities enabled by the Entity registry allows for
multiple statistics and correlations on structured data
linked to annotated documents, e.g., the number of
documents involving natural legal entities, the number and
average value of minors involved in divorce decisions,
correlation for tabulation purposes of the compensation
value and related features in non-pecuniary damage cases.
• Keyword search: Allows simple keyword searches. Further analyses concern survival curves of processes
This module can be useful as a starting point for and explanatory variables of temporal duration and
prothe search, before activating the faceted search. cess complexity. Several analysis functionalities were
• Faceted search: Combines keyword searches and developed within the PON Next Generation UPP project
iflters based on the attributes of the judgments. and other CRUI-funded projects. The following research
The module uses known technologies for index- based on the SICID system registers for the PCT was
ing and querying document databases (e.g., Elas- conducted (see [
          <xref ref-type="bibr" rid="ref6">6, 7</xref>
          ]):
ticsearch).
• Variant Analysis: Clusters of proceedings with the
• LLM-QA: Implements a conversational search same structure and sequence of states and their
based on the RAG paradigm. A generative LLM evaluation for monitoring purposes. In particular,
manages the interaction with the user and the the factors that have the greatest impact on the
generation of responses; a neural retrieval
module allows indexing chunks of judgments using 1https://www.nextgenerationupp.unito.it/
duration of the processes were analyzed. For this
activity, the process mining tool Apromore2 was
used.
• Identification of Critical Events: The impact of
specific events on the duration of a process
execution is evaluated to identify events systematically
associated with anomalous situations. Both the
phases and the total duration of the proceedings
were examined.
• Predictive Approaches for Alerts: Predictors were
constructed from sequences of states or events
in the registers, based on machine learning
techniques with LSTM neural networks, to predict the
residual duration of processes and states during
their course.
        </p>
        <p>A management control dashboard was created for the
Court of Cassation. The adopted solution was to create a
dashboard directly fed by the underlying database of the
Court’s SIC register, with data updated four times a day.</p>
        <p>All data were identified for:
• Feeding the variables and indicators identified
as necessary to describe the file path in the
various phases and to calculate indices such as the</p>
        <p>Disposition Time and the turnover index;
• Building the historical series of such data from</p>
        <p>January 2019.</p>
        <p>Analysis Functionalities for Preliminary
Investigations. Relational knowledge analysis with
visualization (e.g., selection of clusters of nodes with certain
properties) and anomaly detection.</p>
        <p>Functionalities for Penal Execution. Integration
for the social analysis of data relating to liberty
restrictions/alternative penalties experienced by detainees
during their lives.</p>
      </sec>
      <sec id="sec-5-7">
        <title>F4 - Knowledge Base Management and Quality Control - Main Methodologies and Developed Functionalities</title>
        <p>
          3. Extraction of Lexicons: Involves extracting
lexicons of terms based on noun phrases from
judgments and organizing them into an ontology, with
specialization of the lexicon in the legal field
(finetuning).
4. Quality Assessment of NER and NEL: Evaluation
of the quality of Named Entity Recognition (NER)
and Named Entity Linking (NEL) [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ].
5. Benchmarking Extraction Models: Benchmarking
extraction models against various levels of
taxonomy depth, and annotation tools among diferent
relationship extraction models.
6. Introduction of Guardrails: Implementing
guardrails to prevent errors or unprocessable
judgments.
7. Quality Manual for Data, Documents, and
Diagnostic and Predictive Models: Covers aspects such
as accuracy, completeness, currency, fairness, and
explainability (see [
          <xref ref-type="bibr" rid="ref7">8</xref>
          ]). For accuracy and fairness,
the manual aligns with policy documents issued
by the EU (see [9]).
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>4. Ontologies/Taxonomies and</title>
    </sec>
    <sec id="sec-7">
      <title>Their Top-Down and Bottom-Up</title>
    </sec>
    <sec id="sec-8">
      <title>Generation</title>
      <p>In the functionalities of the Datalake, the following
ontologies are used:
• Top Ontology of Justice Procedures
(cognition and execution): Consists of about 400
classes, represented through approximately 40
schemas in the Entity-Relationship model at
different levels of integration/abstraction.
• Ontology for Penal Execution: Consists of
about 100 classes and 8 schemas in the
EntityRelationship model, including all the databases
related to penal execution.</p>
      <p>The following additional ontologies are represented
1. Manual of Pseudonymization Trial Policies: Dif- in the form of two-level taxonomies: i) Top ontology of
ferent types of pseudonymization are considered, preliminary investigations, ii) Top ontology of the civil
and various types of data and document process- trial, iii) Domain ontologies of the civil process: banking,
ing where pseudonymization is relevant (e.g., pub- labor, non-patrimonial damage from privacy violation,
lication, linking databases, etc.) are identified, judicial separation, iv) Ontology for penal-cognition
proalong with the properties that must be respected cedure: victim-perpetrator relationship. The top
ontolin each case. A general method is provided that ogy of Justice procedures and penal execution were
procan be followed for the diferent types of data duced through reverse engineering from logical schemas.
processing relevant to the Datalake project. The ontologies for the victim-perpetrator relationship
2. Entity Registry Management: Includes creation, and non-patrimonial damage were produced by domain
updating, deletion of entities, merging, and split- experts. The ontologies for banking and labor were
proting of entities. duced from lexicons built through the analysis of
judgments.
governance, aligning with the strategic path of digital
transformation of the country, currently being
implemented in the National Strategic Hub. Including services
for a semantic document system for the PA in the service
architecture of the Hub would require the production of a
common top ontology for the PA and high-level modeling
of primary and governance processes, with subsequent
customization by the individual PAs.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>V.</given-names>
            <surname>Bellandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lodi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pozzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ripamonti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Siccardi</surname>
          </string-name>
          ,
          <article-title>An entity-centric approach to manage court judgments based on natural language processing</article-title>
          ,
          <source>Computer Law &amp; Security Review</source>
          <volume>52</volume>
          (
          <year>2024</year>
          )
          <fpage>105904</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Pozzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rubini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <article-title>Named entity recognition and linking for entity extraction from italian civil judgements</article-title>
          ,
          <source>in: International Conference of the Italian Association for Artiifcial Intelligence</source>
          , Springer,
          <year>2023</year>
          , pp.
          <fpage>187</fpage>
          -
          <lpage>201</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Pozzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Moiraghi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lodi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <article-title>Evaluation of incremental entity extraction with background knowledge and entity linking</article-title>
          ,
          <source>in: Proceedings of the 11th International Joint Conference on Knowledge Graphs</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>30</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.-L. H.</given-names>
            <surname>Cabot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          , Rebel:
          <article-title>Relation extraction AFidgmuirneis3t:raAtigoenn. eral semantic document for the Italian Public by end-to-end language generation</article-title>
          ,
          <source>in: Findings of the Association for Computational Linguistics: EMNLP</source>
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>2370</fpage>
          -
          <lpage>2381</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Batini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bellandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ceravolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Moiraghi</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>Pal5</year>
          .
          <article-title>Service Architecture monari</article-title>
          , S. Siccardi,
          <article-title>Semantic data integration for The developed functionalities adopt a service architec- investigations: lessons learned and open challenges, ture for deployment. The multi-node macro functional in: 2021 IEEE International Conference on Smart architecture is shown in Fig. 2. The components of the Data Services (SMDS)</article-title>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>183</lpage>
          .
          <article-title>single node architecture (red frame) are the Multilayer [6]</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Campi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ceri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dilettis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pernici</surname>
          </string-name>
          , et al.,
          <source>VariIngestion Protocol</source>
          ,
          <article-title>Access Control &amp; User Management, ants analysis in judicial trials: Challenges and iniStorage Manager, Document Component, Metadata Man- tial results</article-title>
          ,
          <source>in: Proc. ECML PKDD Workshop on ager, Service Manager</source>
          ,
          <article-title>NLP Service Manager, Analysis, Knowledge Discovery and Process Mining for Law Front End, and Multilayer Export Protocol</article-title>
          .
          <source>(KDPM4LAW)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Pernici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Bono</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Piro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Del Treste</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Vecchi, Improving the analysis of the judiciary 6</article-title>
          . Conclusions:
          <article-title>Towards a performance-the use of data mining techniques to Semantic Document System for assess the timeliness of civil trials</article-title>
          ,
          <source>International Journal of Public Sector Management</source>
          <volume>37</volume>
          (
          <year>2024</year>
          )
          <fpage>59</fpage>
          - the
          <source>Public Administration</source>
          <volume>76</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Batini</surname>
          </string-name>
          ,
          <article-title>Manuale di qualità dei dati, documenti, The semantic document system described in the work modelli di giustizia, 2022. is potentially useful for all Public Administrations (PAs)</article-title>
          . [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Floridi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Holweg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Taddeo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Amaya</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>MökanA project to disseminate the system should involve two der</article-title>
          , Y. Wen,
          <article-title>Capai-a procedure for conducting conphases: an initial phase of parameterization, concern- formity assessment of ai systems in line with the eu ing the organizational structure, ontologies, and primary artificial intelligence act, Available at SSRN 4064091 and governance processes, and a second phase of cus-</article-title>
          (
          <year>2022</year>
          ).
          <source>tomization (see Fig. 3)</source>
          .
          <article-title>Such a project requires strong</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>