<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Generation of Common Procurement Vocabulary Codes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lucia Siciliani</string-name>
          <email>lucia.siciliani@uniba.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Tanzi</string-name>
          <email>e.tanzi2@studenti.uniba.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierpaolo Basile</string-name>
          <email>pierpaolo.basile@uniba.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pasquale Lops</string-name>
          <email>pasquale.lops@uniba.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Bari Aldo Moro, Department of Computer Science</institution>
          ,
          <addr-line>via E. Orabona, 70125, Bari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The role of tenders as means of investment of public funds and as vehicles of strategic development is nowadays crucial. For this reason, developing and enabling new solutions for e-procurement procedures can help to manage and invest funds. In e-procurement, the Common Procurement Vocabulary (CPV) allows assigning a code that classifies its subject to each tender. This study addresses the challenge of automatically assigning a CPV code to a tender. We tackle this problem in two diferent ways: as a classification problem and as a generative task. To develop and test our models, we build a dataset of 5M Italian tenders extracting them from the National Anti-Corruption Authority (Autorità nazionale anticorruzione - ANAC) website. Results show that text classifier approaches exhibit superior performance in this regard. However, they also reveal the potential of generative models in overcoming the limitations of existing classification methods for CPV code assignment in tender classification, providing valuable insights for improving procurement processes and enhancing eficiency in public sector operations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Natural Language Processing</kwd>
        <kwd>e-procurement</kwd>
        <kwd>e-tendering</kwd>
        <kwd>Text Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1. Introduction Vocabulary (CPV)1 that identifies the subject of a tender.
The adoption of the CPV also allows companies to find
Knowledge organization systems (KOS), such as thesauri, new public contracts easily, thus fostering
competitivegazetteers, lexical databases, ontologies, and classifica- ness.
tion systems, are used by institutions to organize large The CPV is structured as a tree of codes comprising
data collections, e.g. documents, web pages, and texts. Us- 9 digits, eight plus a check digit, and specifies whether
ing a standard format guarantees semantic interoperabil- the tender in question refers to supplies, works or
serity and allows for a faster exchange of information. Public vices covered by the contract. Each digit indicates
proprocurement represents a field where adopting such sys- gressively finer-grained classifications. More specifically,
tems can bring many advantages. On one hand, citizens each CPV is composed as follows:
can access data more easily, enabling more
straightforward communication with institutions, which can help
streamline many bureaucratic processes. Concurrently,
adopting KOS systems in procurement has profound
implications for professionals working within public
administrations. In fact, these systems can serve as invaluable
tools, ofering support in the day-to-day activities of
public sector employees. Integrating advanced technologies
will facilitate a paradigm shift towards higher
productivity and eficiency. Tasks that were once labor-intensive
and time-consuming can now be executed with greater
precision and speed, allowing public administrators to
focus on more strategic and value-driven aspects of their
roles. For this reason, in the field of public procurement,
the European Union developed a Common Procurement
• the first two digits identify the divisions (e.g.</p>
      <p>71000000-8 Servizi architettonici, di costruzione,
ingegneria e ispezione (Architectural,
construction, engineering and inspection services));
• the first three digits identify the groups (e.g.</p>
      <p>71300000-1 Servizi di ingegneria (Engineering
services));
• the first four digits identify the classes (e.g.</p>
      <p>71310000-4 Servizi di consulenza ingegneristica
e di costruzione (Consultative engineering and
construction services));
• the first five digits identify the categories (e.g.</p>
      <p>71311000-1 Servizi di consulenza in ingegneria
civile (Civil engineering consultancy services));
• each of the last three digits provides an
additional degree of precision within each category
(e.g. 71311210-6 Servizi di consulenza stradale
(Highways consultancy services));
• a ninth digit serves to verify the previous digits.
Examples of CPV codes are: 30200000-1 (Computer
equipment and supplies), 30230000-0 (Computer hardware), and</p>
    </sec>
    <sec id="sec-2">
      <title>1https://simap.ted.europa.eu/it/web/simap/cpv</title>
      <p>S30231000-7 (Computers and printers). demanding a level of expertise that even seasoned
pro</p>
      <p>
        The supplementary vocabulary can be used to com- fessionals in the field find daunting. Already existing
plete the description of the subject of a contract. The approaches for CPV classification have been condensed
items consist of an alphanumeric code corresponding to within the last few years. In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the author compared
a denomination that allows you to provide further de- diferent deep-learning models for single-label and
multitails on the specific nature or destination of the asset to label CPV classification. The best-performing model was
be purchased. The alphanumeric code is structured as represented by GRU [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] with attention mechanism [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
follows: The dataset used in this work comprises 30,000 Swedish
tenders provided by e-Avrop2, a Swedish company that
• a first level, consisting of a letter corresponding manages an e-procurement platform. Each document
to a section (e.g. A Materiali (Materials)); in the dataset comprises titles and descriptions of each
• a second level, consisting of a letter correspond- tender with their respective categories.
ing to a group (e.g. AA Metalli e leghe (Metal and In [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], the authors used a dataset of 40,000 tenders
alloy)); extracted from TED. The authors have used an LSTM
• a third level, consisting of two digits correspond- [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] architecture for sequence prediction and
classificaing to the attribute (AA02-4 Alluminio (Alu- tion. They also used a Support Vector Machine (SVM)
minium)); to classify the CPV main code category within the same
• the last digit is used to verify the previous ones. framework. The PhD thesis by [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] addressed the CPV
classification problem using a Linear SVM and a bag of
Examples of supplementary codes are the following: words representation from a random sample of 200,000
AA01-1 Metal, or UB05-6 Ofice items . documents extracted from the TED. An important aspect
      </p>
      <p>
        The main vocabulary comprises 9,454 terms and, more to notice is that this work focuses on both English and
specifically, 45 divisions, 272 groups, 1,002 classes, 2,379 French. Kaan Görgün (Mkaan)3 proposed a multilingual
categories, and 5,756 sub-categories. Assigning a CPV to approach which is a fine-tuned version of mBERT [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] on
a tender is a task which is accomplished by RUPs (Respon- tenders extracted from the TED. With their work, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] are
sabile Unico del Procedimento, i.e. Tender’s Managers), instead focused on the Spanish language. The proposed
however, given the high number of terms, it is really dif- method uses RoBERTa-base-bne [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], a RoBERTa model
ifcult even for human experts to identify the right CPV pre-trained on Spanish documents. The authors fine-tune
to use. For this reason, despite assuring a fine-grained this model on Spanish Public Procurement documents,
classification, the high number of labels frequently leads classifying the 45 CPV divisions. Next, they compare
to errors in the CPV assignment like typos or wrong several models, ranging from more classical ones (e.g.
interpretation of the description of each code. Another Naive-Bayes, SVM, KNN, etc.) to the one proposed by
phenomenon is represented by the skewed usage of the MKaan.
codes as there is a small number of CPVs which are more Data from ANAC are a valuable resource for building
known and thus used more frequently while a large num- data-driven systems in the public administration domain.
ber of CPVs are underused. Given these premises, in this In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], authors propose an information extraction
framework, we propose a method for automatically classifying work for Italian tenders [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] that leverage ANAC datasets
tenders to their CPV codes. and other information sources. Moreover, a decision
sup
      </p>
      <p>
        The paper is organized as follows: Section 2 provides port system that helps users during the entire course of
an overview of the approaches available at the state of investments and contracts in e-procurement is described
the art, Section 3 contains the details of the proposed in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
solution, Section 4 reports the results obtained by the
evaluation of our model, and finally Section 5 closes the
paper. 3. Methodology
The main idea is to support RUPs in assigning CPV to
a new tender. Specifically, the aim is to establish a
robust system capable of proficiently classifying a tender
based on its specified object, thereby facilitating an
accurate alignment with the comprehensive set of CPV codes
available. The classification task is dificult since the
number of codes (CPVs) is high. Therefore, we propose two
      </p>
      <sec id="sec-2-1">
        <title>2. Related Work</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The assignment of a Common Procurement Vocabulary</title>
      <p>(CPV) code to a tender is crucial for the accurate
identification and precise retrieval of analogous documents.</p>
      <p>This meticulous categorization process is indispensable
for ensuring the streamlined organization and efective
utilization of information. Given the nature of the CPV 2https://info.e-avrop.com
vocabulary, which encompasses a set of over nine thou- 3https://huggingface.co/MKaan/
sand terms, this task assumes a challenging dimension, multilingual-cpv-sector-classifier
methodologies: 1) a text classification approach based on
diferent classifiers; 2) a generative approach based on
Transformers with an encoder-decoder architecture.</p>
      <p>Both approaches work on the same data. In
particular, given a list of tuples (CPV, CPV description, tender
object), we split it into three sets: training, validation
and testing. More details on the dataset are reported in
Section 4. Then each approach is implemented, trained
and validated separately on the same data.
use only the divisions as labels to obtain reasonable
results. This strategic decision is driven by recognising that
achieving reasonable results across the entire spectrum
of CPV codes poses significant challenges. By
concentrating on this subset, we optimize the model’s capacity
to provide meaningful and accurate predictions within a
more manageable scope.</p>
      <sec id="sec-3-1">
        <title>3.2. Generative Approach</title>
      </sec>
      <sec id="sec-3-2">
        <title>3.1. Text Classification</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Regarding text classification, our approach aligns with a</title>
      <p>classical pipeline:
Since our approach aims to suggest a CPV given a
tender’s description, we decided to investigate the ability of
AI generative methods to automatically produce a text
given a textual input. Moreover, we want to test if a
generative approach can provide better results when a
• Preprocessing: the initial step involves the pre- large number of classes is involved, as in our domain. We
processing of the tender object, wherein trans- adopt a classical encoder-decoder architecture that has
formations such as lowercase conversion and to- proven to provide promising results in several NLP tasks.
kenization are applied. These essential prepro- Encoder-decoder architectures are well-suited for solving
cessing techniques lay the groundwork for sub- sequence-to-sequence problems like machine translation,
sequent stages by standardizing the textual data; as they can efectively process variable-length input and
• Feature Vector Generation: following the pre- output sequences. In this architecture, the encoder takes
processing step, we construct feature vectors for in a sequence of any length and converts it into a
fixedeach tender object. This involves the utiliza- shaped state. On the other hand, the decoder maps the
tion of both Bag-of-Words (BoW) and TF-IDF encoded state, which has a fixed shape, back to a sequence
(Term Frequency-Inverse Document Frequency) of variable length.
approaches. By encoding the textual information In our case, the encoder’s input is the tender’s
obinto numerical representations, we aim to cap- ject description and the decoder output is the CPV code
ture the salient features that contribute to the and its description. It is important to underline that this
classification task; method can produce a CPV code or a description that
• Classifier Training and Tuning : subsequently, is not present in the original list of CPVs, while a text
a classifier is trained using the feature vectors gen- classifier produces as output a CPV from the set of
predeerated in the previous step. The training process ifned CPVs. This poses problems in the evaluation phase
is complemented by the optimization of hyper- as comparing the output produced with the gold
stanparameters. This optimization is achieved with dard present in the test set becomes more complex. More
the use of a validation set, ensuring that the clas- details about the evaluation are reported in Section 4.
sifier retains generalization capabilities;
• Evaluation: the final stage of our classification
pipeline involves the evaluation of the trained 4. Evaluation
classifier on an independent test set. This allows
an assessment of the model’s ability to generalize
and classify unseen tender objects. It serves as a
critical benchmark to validate the efectiveness
and robustness of the entire text classification
system.</p>
    </sec>
    <sec id="sec-5">
      <title>For the implementation, we rely on spaCy for text</title>
      <p>processing and scikit-learn for classification. The
hyperparameters are found through the grid search. After a
ifrst evaluation, we select the following classifiers: Linear
SVC and Multinomial Naive Bayes. Moreover, we cast the
problem only to classify the divisions that are composed
of 45 classes since the model cannot provide reasonable
results when the whole set of CPVs is involved.</p>
      <p>Moreover, we choose to investigate a classifier based
on BERT using its tokenizer. Again, in this setting, we</p>
    </sec>
    <sec id="sec-6">
      <title>For building and testing, we extract data from the</title>
      <p>ANAC anticorruzione (anticorruption) website4. ANAC
-Autorità Nazionale AntiCorruzione National
AntiCorruption Authority is an Italian independent
administrative authority with the aim of combating corruption
in the country.</p>
      <p>In particular, we retrieve for each tender the CIG (the
tender identifier), the tender type, the description of the
tender object, the CPV assigned by the RUP and the CPV
description. The tender type identifies three kinds of
tender: 1) supplying, 2) service and 3) work. From the
original dataset, we remove CPV codes that occur less
than 20 times and store data in a CSV file for a total of
5 million tenders. We split the dataset in training, test</p>
    </sec>
    <sec id="sec-7">
      <title>4https://dati.anticorruzione.it/opendata/dataset/</title>
      <p>and validation according to the percentages reported in
Table 1.</p>
      <p>training
test
validation</p>
      <p>For training the encoder-decoder architectures, we
build a diferent version of the dataset in the JSONL
format. Each row in the dataset is a JSON object with two
elements: source and target. The source is the input text
of the encoder and the target is the output text of the
decoder. In our case, the source is the concatenation of
the type of the tender and the description of the tender’s
object, while the target is the concatenation of both the
code and the description of the CPV. The tender’s type
defines the nature of the object from a list of predefined
types: service, supply and work. Listings 1 shows an
example of a JSON object related to a tender.
1 {"source":"lavori lavori di pavimentazione delle
vie san martino e santa Maddalena",
2 "target":"45262321-7 - lavori di pavimentazione"
}</p>
    </sec>
    <sec id="sec-8">
      <title>Listing 1: An example of a JSON object for training the encoder-decoder architecture.</title>
      <p>The dataset is stored on Zenodo5, while the code is
available on GitHub6. The code for fine-tuning the IT5
model is available here7. The IT5 models fine-tuned on
the CPV generation task are on HuggingFace: the large
model8.</p>
      <sec id="sec-8-1">
        <title>4.1. Text Classification</title>
        <p>This sub-section reports information and results about
text categorization approaches. Regarding the
parameters’ optimization, we adopt a grid search for Linear SVC
and Multi-NB. For Linear SVC, we optimize the
parameter  in the set {0.5, 1, 2, 4, 8}, while for Multi-NB we
consider ℎ in {0.1, 0.3, 0.5, 0.7, 0.9} and  _ in
{True, False}. The best values selected after the grid search
are  = 0.5, ℎ = 0.3 and  _ =  .</p>
        <p>For BERT, we did not perform parameters
optimization since the required computational time is very high.
We fine-tuned a specific language model for Italian
called dbmdz/bert-base-italian-uncased9 using
5https://zenodo.org/records/10007545
6https://github.com/ematanzi/Valutazione-CPV/
7https://github.com/gsarti/it5
8https://huggingface.co/basilepp19/cpv-it5 and the base model
https://huggingface.co/basilepp19/cpv-it5-base
9https://huggingface.co/dbmdz/bert-base-italian-uncased
the Adam optimizer with a learning rate of 5e-05 and
a batch size of 16 trained for 5 epochs.</p>
        <p>Results of text classification approaches are reported
in Table 2. Generally, the results are very low due to
the large number of classes and BERT reports the worst
performance since, for some classes, there are very few
examples in training data. We observe a large accuracy
with respect to the F1 measure. This is due to the presence
of few classes with many examples. For these classes,
classifiers can achieve good performance. For example,
BERT achieves the 90% of F1 for the most frequent class10.</p>
      </sec>
      <sec id="sec-8-2">
        <title>4.2. Generative Approach</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>We used the Java library Lucene for text searching and</title>
      <p>indexing, with the aim of solving the task as a retrieval
task.</p>
      <p>In detail, we indexed CPV code description pairs
corresponding to target in the two fields code and description.
Afterwards, we ran the search for each source element
in JSON file, obtaining for each search the element
belonging to target with the most similar description to
the source string, and saved results in a file containing
the triple source, target and generated. Where target is
the expected description and generated is the retrieved
one. We adopt the same output format for the generative
approaches. The idea is to exploit the source as the query
for the search engine and retrieve the most similar code
descriptions using the search engine.</p>
      <p>We executed this experiment four times, implementing
variations in both the configuration of the text analyzer
and the choice of two distinct similarity measures. The
adopted configurations are the following:
• StandardAnalyzer + default similarity;
• StandardAnalyzer + LMDirichletSimilarity;
• ItalianAnalyzer + default similarity;
• ItalianAnalyzer + LMDirichletSimilarity.</p>
    </sec>
    <sec id="sec-10">
      <title>The ItalianAnalyzer performs a specific stemming al</title>
      <p>gorithm for Italian, while the StandardAnalyzer
implements a grammar-based tokenizer for several languages.</p>
      <p>The default similarity provided by Lucene is the BM25
model [12], while the LMDirichletSimilarity uses a
language model for information retrieval with the Bayesian
smoothing based on Dirichlet priors [13]. The evaluation
has been carried out on the test set. We decided to use
BLEU metric to measure matching between generated
and target text since this metric is based on the idea that
the nearer the predicted text is to the target one, the more
correct it is. Considering the small size of the compared
strings, we decided only to use 1-gram and 2-gram of</p>
      <p>1033000000-0 Apparecchiature mediche, prodotti farmaceutici
e per la cura personale (Medical equipment, pharmaceuticals and
personal care products)
consecutive words, attributing them to the same weight We evaluate the generative approach based on the
(0.5). In calculating the metric, a smoothing function has encoder-decoder transformer by fine-tuning it on
trainbeen used, increasing the score when there are partial ing data. We start from a pre-trained Italian model called
matches between the generated text and the target text. IT5 [14]. The IT5 model family is the initial endeavour</p>
      <p>In the evaluation, since the resolution of this classi- to pre-train extensive sequence-to-sequence transformer
ifcation task with utmost precision is arduous even for models specifically designed for the Italian language,
inhuman experts, we decided to consider every possible spired by the methodology employed in the original T5
correspondence between the code-description generated model [15]. We fine-tuned two diferent models with
couple and the target one, which are the following: diferent sizes: IT5-lager and IT-base.</p>
      <p>Results are reported in Table 4 for the large model and
in Table 5 for the base one.
• the full correspondence;
• the correspondence between codes, but not the</p>
      <p>description (the opposite case can never occur);
• the correspondence between categories;
• the correspondence between classes;
• the correspondence between groups;
• the correspondence between divisions;
• the case of no match.</p>
    </sec>
    <sec id="sec-11">
      <title>In this way, we also evaluate the cases in which the</title>
      <p>solution has been approached. For all of these cases, we
calculated the number of times they occurred along with
the relative average of the BLEU score metric. Eventually,
we also calculated the average BLEU score related to all
the tests performed.</p>
      <p>The best results obtained by the baseline are reported
in Table 3. The results are obtained using the
ItalianAnalyzer and the default similarity. The baseline based on
the search engine is able to correctly retrieve the correct
CPV with the correct description for only 11.34% of
testing data. In the 61.77% of cases is not able to retrieve the
correct CPV with very low BLEU (0.0515), this means
that the description of the first retrieved CPV is very
diferent from the correct one.</p>
      <p>Perfect match
Only CPV code
Only category
Only class
Only group
Only division
No match
All
# matchs
113,440
872
30,502
53,232
90,673
93,574
617,707</p>
    </sec>
    <sec id="sec-12">
      <title>To compare generative approaches with text catego</title>
      <p>rization ones, we consider from the generated output
only the first two digits of the CPV, i.e. the CPV divisions.
This choice allows us to compare the generative approach
with the ones based on text categorization since the latter
are trained to predict only the division of each tender.
Results of this analysis are reported in Table 6 and show
that if we consider generative approaches as a classifier,
they are below the simple Linear SVC. Performing a dual
evaluation in which a classifier is evaluated as a
generative approach is not possible since we train classifiers
only for predicting divisions. Considering only the code
of the division is not possible to generate a description
comparable to the text generated by an IT5 model.
model
Linear SVC
IT5-large
IT5-base</p>
    </sec>
    <sec id="sec-13">
      <title>Anyway, these outcomes are encouraging since the</title>
      <p>IT5 model is trained on all the possible descriptions of a
CPV, while the text classifier approaches handle only the
code of the division and cannot provide usable results
when they are trained on the whole set of possible classes.
Moreover, generative approaches provide a significant
improvement with respect to the baselines obtained by a
search engine.</p>
      <sec id="sec-13-1">
        <title>5. Conclusions</title>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>In this paper, we tackle the challenge of categorizing a</title>
      <p>tender by aligning it with the comprehensive Common
Procurement Vocabulary (CPV), i.e. a meticulously
curated European lexicon of codes designed to precisely
identify the subject matter of a tender. The complexity
of this task lies in the diverse nature of procurement
scenarios, where each tender has its own description
and requirements. The CPV emerges as a fundamental
tool in deciphering the procurement language, trying to
define a European dictionary allowing interoperability
among diferent countries. Our proposed methodologies
encompass two distinctive approaches: the former relies
on a conventional text classification paradigm, whereas
the latter leverages a generative strategy hinging on the
encoder-decoder architecture as conceptualized by the
T5 model.</p>
      <p>In our systematic exploration of the system’s
proficiency in discerning the accurate division of a tender,
specifically on the initial two digits of the CPV, it
becomes evident that text classifier approaches provide
the best results. Nevertheless, a noteworthy result
surfaces when we focus on the holistic identification of the
entire CPV through a descriptive context. In this
context, the generative approaches exhibit commendable
eficacy, demonstrating promising outcomes. Notably, these
generative techniques surpass established baselines
constructed through conventional keyword-centric search
engines, attesting to their heightened capabilities in
nuanced comprehension and contextual inference.</p>
      <sec id="sec-14-1">
        <title>Acknowledgments</title>
        <p>We acknowledge the support of the PNRR project FAIR
Future AI Research (PE00000013), Spoke 6 - Symbiotic AI
(CUP H97G22000210007) under the NRRP MUR program
funded by the NextGenerationEU.</p>
        <p>We thank Salvatore Tucci for his helpful support in
evaluating text classification approaches.
P. Lops, Ai-based decision support system for
public procurement, Information Systems 119
(2023) 102284. URL: https://www.sciencedirect.com/
science/article/pii/S0306437923001205. doi:https:
//doi.org/10.1016/j.is.2023.102284.
[12] S. Robertson, H. Zaragoza, The probabilistic
relevance model: Bm25 and beyond, in: the 30th
Annual International ACM SIGIR Conference, 2007,
pp. 23–27.
[13] C. Zhai, J. Laferty, A study of smoothing methods
for language models applied to ad hoc information
retrieval, in: ACM SIGIR Forum, volume 51, ACM
New York, NY, USA, 2017, pp. 268–276.
[14] G. Sarti, M. Nissim, It5: Large-scale text-to-text
pretraining for italian language understanding and
generation, arXiv preprint arXiv:2203.03759 (2022).
[15] C. Rafel, N. Shazeer, A. Roberts, K. Lee, S. Narang,
M. Matena, Y. Zhou, W. Li, P. J. Liu, Exploring the
limits of transfer learning with a unified
text-totext transformer, The Journal of Machine Learning
Research 21 (2020) 5485–5551.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Suta</surname>
          </string-name>
          ,
          <article-title>Multilabel text classification of public procurements using deep learning intent detection</article-title>
          ,
          <source>Degree project in mathematics (second cycle)</source>
          ,
          <source>Kth Royal Institute of Technology</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Van</given-names>
            <surname>Merriënboer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gulcehre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bougares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Learning phrase representations using rnn encoderdecoder for statistical machine translation</article-title>
          ,
          <source>arXiv preprint arXiv:1406.1078</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Bengio,</surname>
          </string-name>
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          ,
          <source>arXiv preprint arXiv:1409.0473</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kayte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schneider-Kamp</surname>
          </string-name>
          ,
          <article-title>A mixed neural network and support vector machine model for tender creation in the european union ted database</article-title>
          .,
          <source>in: KMIS</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>139</fpage>
          -
          <lpage>145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <article-title>Long short-term memory</article-title>
          ,
          <source>Neural computation 9</source>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>O.</given-names>
            <surname>Ahmia</surname>
          </string-name>
          ,
          <article-title>Assisted strategic monitoring on call for tender databases using natural language processing, text mining and deep learning</article-title>
          ,
          <source>Ph.D. thesis</source>
          , Université de Bretagne Sud,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Navas-Loro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garijo</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <article-title>Corcho, Multi-label text classification for public procurement in spanish</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>69</volume>
          (
          <year>2022</year>
          )
          <fpage>73</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gutiérrez-Fandiño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Armengol-Estapé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Llop-Palao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Silveira-Ocampo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Carrino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gonzalez-Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Armentano-Oller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rodriguez-Penagos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <article-title>Spanish language models</article-title>
          ,
          <source>arXiv preprint arXiv:2107.07253</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Siciliani</surname>
          </string-name>
          , E. Ghizzota,
          <string-name>
            <given-names>P.</given-names>
            <surname>Basile</surname>
          </string-name>
          , P. Lops,
          <article-title>Oie4pa: open information extraction for the public administration</article-title>
          ,
          <source>Journal of Intelligent Information Systems</source>
          (
          <year>2023</year>
          ). URL: https://link.springer.com/ article/10.1007/s10844-023-00814-z. doi:https:// doi.org/10.1007/s10844-023-00814-z.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Siciliani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Taccardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Basile</surname>
          </string-name>
          , M. Di Ciano,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>