<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Legal Education (2023).
[31] I. Chalkidis</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1016/j.procs.2020</article-id>
      <title-group>
        <article-title>LLM-based Approaches for Automatic Ticket Assignment: A Real-world Italian Application</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicola Arici</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Putelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alfonso E. Gerevini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Sigalini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ivan Serina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Brescia</institution>
          ,
          <addr-line>Via Branze 38, Brescia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mega Italia Media</institution>
          ,
          <addr-line>Via Roncadelle 70A, Castel Mella</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>176</volume>
      <fpage>16</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>IT service providers need to take care of errors, malfunctions, customizations and other issues every day. This is usually done through tickets: brief reports that describe a technical issue or a specific request sent by the users of the service. Tickets are often read by one or more human employees and then assigned to technicians or programmers in order to solve the raised issue. However, the increasing volume of such requests is leading the way to the automatization of this task. Since these tickets are written in natural language, in this paper we aim to exploit the new powerful pre-trained Large Language Model (LLM) GPT-4 and its knowledge in order to understand the problem described in the tickets and to assign them to the right employee. In particular, we focus our work on how to formulate the request to the LLM, which information is needed and the performance of diferent zero-shot learning, few-shot learning and ensemble learning approaches. Our study is based on a real-world ticket dataset provided by an Italian company which supplies IT solutions for creating and managing online courses.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Automatic Ticket Assignment</kwd>
        <kwd>Large Language Models</kwd>
        <kwd>Prompt Engineering</kwd>
        <kwd>Text Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Modern companies which supply IT solutions not only have to provide an efective software
environment, but they also need to maintain it during the software lifecycle, to fix errors and
malfunctions, to introduce new functionalities and satisfy the requests submitted by the users.
This task is usually done by programmers specialized in maintenance tasks which need to take
care of new issues every day.</p>
      <p>Such issues are usually submitted through ticketing systems. In these systems, the users can
write a brief report that describes a problem they encountered, or a specific service they need.
These reports, typically called tickets, are then distributed among the maintenance specialists
which have to satisfy the users’ requests. However, in large companies which provide complex
IT solutions or more than one product, diferent employees devoted to the maintenance can
have diferent expertise. Therefore, there is the need to assign a ticket to the right person, i.e.
an employee who has the necessary technical skills to solve the raised issue.</p>
      <p>
        In order to do that, typically one or more human employees have to read the ticket, understand
the request and assign it to a technician. Since this task is quite time consuming, bigger
companies are starting to implement automatic solutions. Although these solutions can be
based on ad hoc algorithms [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] or on fine-tuning generic pre-trained language models [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] such
as BERT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], they would require a considerable amount of training data and an expensive efort
(by programmers and machine learning specialists) to implement such models. On the other
hand, the outstanding results obtained by pre-trained large language models (LLMs) as few-shot
learners (i.e. with a minimal number of training examples) [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
        ] could make automatic ticket
assignment available to many companies even without any particular efort.
      </p>
      <p>
        In order to verify whether that is achievable, in this work we investigate how these models
can be applied to a real-world case scenario: the assignment of the tickets received by the
Italian company Mega Italia Media1, which provides IT solutions in the e-learning sector for the
occupational safety. In particular, in 2011 they released the DynDevice2 Learning Management
System (LMS), facilitating companies in standard corporate training, allowing them to create
specific courses, managing final exams [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ] (providing also the related certificates if the
exam has been passed) and the interaction with the users [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The company receives many
tickets related to this platform, which has to be maintained and updated constantly in order to
satisfy the users’ needs. Using these tickets, we verify the performance of OpenAI GPT-4 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ],
a state-of-the-art pre-trained LLM based on the Transformer architecture [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], for this task.
      </p>
      <p>
        However, it has been noted that the performance of such models can significantly vary
depending on how the task requested is formulated or, in more technical terms, which prompt
has been used [
        <xref ref-type="bibr" rid="ref14">14, 15</xref>
        ]. Therefore, we study diferent configurations and prompts into which
more or less information is available to the LLM and in terms of how many examples we provide.
We compare these results with a baseline into which a BERT model is fine-tuned on this task
with 1000 labeled tickets.
      </p>
      <p>The rest of the paper is organized as follows. In Section 2, we provide the background and an
overview of the state-of-the-art and the related works. In Section 3, we describe the dataset
of our application. In Section 4, we describe our approaches for solving the automatic ticket
assignment task, which are evaluated and discussed in Section 5. Finally, in Section 6 we propose
some conclusions and future developments. The code and the datasets can be found on GitHub3</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        In recent years, several researchers have approached the support ticket domain, solving problems
such as ticket categorization [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4, 16</xref>
        ], ticket assignment and ticket resolution [17]; our work
falls in the second category.
      </p>
      <p>In 2018, Uber, the famous private car transport company, proposed COTA [18] (Customer
Obsession Ticket Assistant), a framework to take care of customer issues. They proposed two
versions of their system: the first one combines several features, such as user information, trip
information and ticket metadata, with a Random Forest algorithm for predicting the correct</p>
      <sec id="sec-2-1">
        <title>1https://www.megaitaliamedia.com/en/</title>
        <p>2https://www.dyndevice.com/en/
3https://github.com/nicolarici/AI-TS
operator of each ticket; the second version leverages a Encoder-Combiner-Decoder approach,
based on CNN and RNNs over diferent types of features (such as categorical, numerical, binary
and text features) and a multi-classification layer.</p>
        <p>A similar approach has been developed by DeLucia and Moore [19]; the authors implemented
a Random Forest model fed with features created with latent Dirichlet allocation topic modeling,
latent semantic analysis and Doc2Vec [20] starting from the ticket subject and message.</p>
        <p>Han and Sun [21] proposed in 2020 DeepRouting, an intelligent system for assigning tickets
to operators in an expert network. It contains two modules: one for text matching, based on a
convolutional neural network trained over tri-grams derived by the ticket description, and one
for graph matching, based on a Graph Convolutional Network fed with the experts graph.</p>
        <p>With Feng et al. [22], in 2021 Apple developed its personal ticket assignment system, TaDaa
(Ticket Assignment Deep learning Auto Advisor). This system is based on the
state-of-theart Transformer architecture, in particular a pretrained BERT model fine-tuned to solve two
classification tasks. The model has two diferent classifiers: the first one to assign the ticket to
one of the 3000 groups, and the second one to identify the expert (that belongs to that group)
that is going to solve the issue. We used a similar idea, in our much simpler context, with the
BERT baseline described in Section 5. However, in this paper we show how also with a limited
number of examples, pre-trained LLMs can achieve similar performance.</p>
        <p>Diferently from these works, which are based on custom algorithms and models (which
require a considerable efort for designing, implementing and testing), in our work we verify
whether pre-trained LLMs can be used for this kind of task, even without fine-tuning. More
generally, we exploit prompt engineering, which was designed precisely for obtaining the best
results from these pre-trained models. Regarding this line of work, White et al. [23] provide
a pattern catalog to solve common problems when conversing with a LLM. They propose 17
patterns which allow the users to better handle the input to give to the model, the output
structure and format, possible errors in content, i.e. invented answers based on unverified facts,
the prompt and how it can be improved to receive better responses and the interaction between
the user and the model and the context needed by the model to generate a better response.</p>
        <p>Moreover, Reynolds at el. [24] showed that zero shot learning (i.e. without any example
provided to the model) with a good prompt can outperform a standard few shots approach
(i.e. with some examples); to do so they introduced the concept of meta-prompt, that seeds
the model to generate its own natural language prompt to solve the task. A similar result has
been achieved by Zhou et el. [25], where the authors propose Automatic Prompt Engineer, a
framework for automatic instruction generation and selection. In their method the authors
optimized the prompt by searching over a pool of instruction candidates proposed by an LLM
in order to maximize a chosen score function.</p>
        <p>The LLMs, in particular ChatGPT, have been proven very efective in solving specific NLP tasks
on general domains. Even in specific domains, such as public health [ 26, 27, 28], environmental
problems [29] or legal rulings and laws [30, 31], the LLMs achieve acceptable performance.
To the best of our knowledge, there are currently no applications of GPT-based models in the
context of workplace security.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Available data</title>
      <p>Since the release of DynDevice in 2011, Mega Italia Media started facing the problem of assisting
and supporting end users. Originally, this service was provided by phone calls or emails but,
with the strong spread of the platform over the years, these channels were soon saturated.
To help the company operators to solve the users problem, in 2013 the company developed a
ticketing system; this new feature allowed the company to keep track of all the tickets opened,
memorizing the status of the user’s request and who was in charge to solve the problem.
Moreover, this system kept record of all the conversations between users and operators.</p>
      <p>Overall, the company has received more than 10000 tickets in the last 10 years. However, in
the last period several new features and services (such as multiple interface changes, a videocall
system and several AI applications) were introduced, and therefore we decided to consider in
our dataset only the tickets received in the last months, which are about 1300.</p>
      <p>Furthermore, to build our dataset we decided only to keep the significant information: the
ticket category, object and description and the area who solved the ticket; other information such
as dates and identifiers has been removed. In the following we provide an example of a ticket;
please note that this example has been translated, since all our tickets are written in Italian.
Example 1.</p>
      <p>CATEGORY: G. More on e-Learning/training solutions
SUBJECT: Certificate with exam in presence HTML5
DESCRIPTION: "Certificate with Examination in Attendance" I turned it into HTML 5. Everything
is fine except for the column rows that do not appear. What could be the problem? Thank you
AREA: SW</p>
      <p>A. Problems in using a course</p>
      <p>B. Inability to access a course
C. Clarification of the didactics of a course</p>
      <p>D. Problems in generating reports
E. Problem in generating/recovering documents</p>
      <p>F. User registration problem
G. Other on e-Learning/training solutions
I. e-Commerce and website management</p>
      <p>H. HR management
M. School of Security</p>
      <p>N. Commercial
28
7
3
29
38
36
145
12
86
2
3
SW
63
37
16
23
41
43
8
27
5
9
129
TECH
eLEARNING
173
28
43
14
75
36
39
5
0
30
10
180
160
140
120
100
80
60
40
20
0</p>
      <p>The category field contains one of the 11 predetermined categories decided by the company.
A complete categories list is reported in Figure 1. The object and the description are text fields
that require a slight pre-process in order to remove the HTML tags, the URLs and other special
characters that can harm the classification process. Each ticket is assigned to a specific operator
who is in charge of solving the raised issue. However, all the operators can be aggregated in
three main macro areas:
• eLEARNING: which handles the problems on the courses provided to the end users;
• TECH: which solves the technical issues about the platform;
• SW: that removes bugs and other software issues.</p>
      <p>Each area manager assigns the ticket to a single operator suited to handle the case.</p>
      <p>In the left part of the Figure 2, we reported the ticket distribution among the macro areas; as
we can see the three classes are approximately equally distributed. On the contrary, as shown
in the right part of Figure 2, the categories follow an unbalanced distribution, with category G
being the most frequent, with more than 300 tickets compared to the least frequent. Another
statistic that we extrapolate from the dataset is the cross tabulation between the categories
and the area fields. As we can see in Figure 1, there is a strong correlation between some
categories and some areas: the category A has a strong correlation with the eLEARNING area
(and vice-versa), whereas both TECH and SW have a strong correlation with the category G. We
expect tickets in these categories to be best assigned with a well constructed prompt containing
this information. Other categories do not present any correlation, in some cases due to the low
number of tickets, such as categories N and H. In other cases (such as categories F and E) the
tickets are equally distributed between all the areas.</p>
      <p>Finally, we decided to sample with stratification, following the area distribution,
approximately 250 tickets to build 5 diferent test sets; this way, each test set contains 18 tickets for
eLEARNING, 15 tickets for TECH e 14 tickets for SW. The tickets sampled with this strategy
retain also the categories distribution and the cross correlation discussed above.</p>
      <p>G A E F I B D C M H N</p>
    </sec>
    <sec id="sec-4">
      <title>4. Prompts and methods for automatic ticket assignment</title>
      <p>The base task to solve for the automatic ticket assignment is a multi-class classification, into
which we have to choose which group (eLEARNING, TECH or SW) will receive the ticket.
To solve this task, we decided to exploit the Python version of the OpenAI Chat Completion
API4, which allows interaction with pre-trained LLMs. For each call, the API requires several
parameters. The most relevant ones for our work are: the model we want to query, which in our
case can be GPT-4 or GPT-3.5-turbo, and the temperature, a decimal number between 0.0 and 2.0
(default 1.0), which controls randomness (higher values) or determinism (lower values) in the
response generated. To have as much determinism as possible, in all trials we set temperature
to 0.0. The corpus of the API request is composed by two messages: the system prompt, that
contains the description of the task and the information to solve it, and the user prompt, i.e. the
ticket to classify.</p>
      <p>We tested GPT in three ways: zero shot learning, few shots learning, and ensemble learning.
For the zero shot learning scenario, we provided to the model insightful information in the
system prompt and no examples; all the information was extracted from the database described
in Section 3. These are the zero shot prompts we implemented:
• Baseline: it describes the task, provides the basic information and imposes to GPT to
answer only with the macro area name. From here on, each subsequent prompt is to be
considered concatenated with this one.</p>
      <p>Example: You are the manager of a service center whose task is to divide tickets between
the various human operators. The available operators, contained in square brackets are as
follows: [eLEARNING, TECH, SW]. The tickets are divided into categories, listed by letters
of the alphabet, contained in round brackets, which are as follows: (A. Problems in using a
course, B. ...). Each ticket consists of a subject and a description. Your task is to assign the
ticket to the most suitable operator, answering only with the operator’s name.
• Human: it contains insightful information, provided by a human employee, on the role
of the three areas and the problems solved.</p>
      <p>Example: SW handles technical problems with software code, new customisation and
developments and ICT (Information and Communication Technologies) and SEO (Search Engine
Optimisation) issues.</p>
      <p>The same information is provided for the other 2 areas.
• Categories: it contains information related to the assignment of the tickets aggregated
by category. For each category, we extract the percentage of tickets solved by each area
belonging to that category (which can be seen in Figure 1). For instance, for the category
A we wrote:
Example: 66% of the tickets in the category "A. Problems in using a course" are assigned to
eLEARNING, 24% to TECH and 11% to SW.</p>
      <p>The same information is provided for the other 10 categories.
• Areas: it contains information about the assignment of tickets aggregated by areas. From
the Figure 1, for each area, we extract the percentage of tickets belonging to each category</p>
      <sec id="sec-4-1">
        <title>4https://platform.openai.com/docs/api-reference/chat</title>
        <p>solved by the area. For instance, for the area eLEARNING we have:
Example: To eLEARNING is assigned 38% of the tickets in category A, 8% in B, 9% in C, 3%
in D, 17% in E, 8% in F, 8% in G, 1% in H, 0% in I, 7% in M and 2% in N.</p>
        <p>The same information is provided for the other 2 areas.
• Summaries: for this prompt we asked GPT-4 to summarize the issues raised by the users
in 10 tickets for each area; these tickets were randomly sampled from the training dataset.
For instance, for the area eLEARNING we have the following result:
Example: eLEARNING solves problems related to the activation of courses, changes of
certificates, platform access problems, cancellation of courses, issuing of certificates, downloading
of certificates and approval of enrolment forms.</p>
        <p>The same summary is generated by GPT-4 for the other two areas and used by GPT in
the system prompt. Please note that these summaries are generated separately by the
model. At the moment of the ticket assignment, the LLM receives the summary without
any knowledge of the 10 examples used for generating it.</p>
        <p>
          As a second approach we implemented the few shots learning. The basic idea is to give to
the model some examples of classification so it can understand the pattern, generalize it and
then apply what it has learned to test instances. This approach has achieved outstanding results
in a lot of NLP tasks [
          <xref ref-type="bibr" rid="ref6 ref8">6, 8</xref>
          ]. All our few shots approaches use the Baseline Prompt, without
additional information. Instead, we provided to the model some examples in the same format
expressed in the Example 1, with the correct macro area assigned.
        </p>
        <p>The last approach, the ensemble learning, leverages the best results obtained in the past
trials to try to use all the information at our disposal. Thus, for each ticket, we made n calls
to the OpenAI API with diferent prompts, keeping model and temperature unchanged. Each
result is counted as a vote, with no specific weight assigned, and the area with the most votes
has the ticket assigned. In the event of a tie, the ticket is randomly assigned to one of the areas
with the most votes. In our experimental evaluation we use an ensemble made by three and
four prompts.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental results</title>
      <p>In this section, we report the results of our experiments. For each trial, in Table 1 we report the
mean and the standard deviation over the five test sets for two metrics: the accuracy and the
macro F1 score, both expressed between 0.0, the worst, and 1.0, the best.</p>
      <p>In addition to the approaches described in the previous section, to provide a baseline for
comparison with state-of-the-art approaches, we tested how a classical approach with BERT
solves the task. Starting from a pre-trained Italian version of BERT available on HuggingFace5,
we fine-tuned the model on this specific task on 907 samples for 20 epochs and we took the best
performing model on a validation set made by 101 examples. To perform the classification, we
add to BERT a simple feedforward layer with three neurons preceded by a 0.1 dropout layer. In
training we used the Binary Cross Entropy loss function, the AdamW optimizer, with learning
rate set to 2e− 5; we set decay to 0.01 and batch size to 32. We fed this classification model with</p>
      <sec id="sec-5-1">
        <title>Approach</title>
        <p>ZS
ZS
ZS
ZS
ZS
ZS
FS
FS
FS
FS
BERT
BERT
EL
EL</p>
      </sec>
      <sec id="sec-5-2">
        <title>Prompt</title>
        <p>Baseline
Summaries
Categories
All Information
Areas
Human
One Example
Three Examples
Five Examples
Ten Examples
Only Ticket
Full
Three Prompts
Four Prompts</p>
        <p>Accuracy
two types of input: one containing just the description of the ticket and one with all its three
properties: category, object and description. The results obtained by the zero shot learning with
no additional information (0.34 accuracy and 0.29 macro F1 score), and by BERT (0.65 accuracy
and 0.65 macro F1 score), constitute the baselines of our framework.</p>
        <p>As regarding the zero shot approach, when we start adding information to the base prompt,
the accuracy improves. The two trials that leverage the correlation between the area and the
category reach accuracy 0.53 for the Categories Prompt and 0.56 for the Areas Prompt. The
reason for this behavior could be due to the length of the prompts: the first prompt is, more
or less, 200 tokens long and some information is forgotten or ignored by the LLM. Also, in
these two cases we provided only information about the categories and no information about
the ticket description. This information is contained in the other two prompts: the Human
and Summaries prompts. For the first case, the model reaches about 0.57 accuracy and an
higher macro F1 score (0.54) and with the lowest standard deviation. For the second prompt, the
accuracy drops to 0.50 with an higher standard deviation (0.07), probably due to the fact that the
summaries are obtained with 10 tickets only, and the accuracy changes depending on whether
the tickets in the dataset are more or less similar than those used for the summaries. When
we pass all the information to the model the performance is lower (0.54 accuracy), probably
because the prompt reaches a critical length.</p>
        <p>For the few shots approaches, the results are not particularly good. The single example case
manages to perform even worse with respect to the baseline, with an accuracy of approximately
0.19. However, as it can be seen from the Table 1, increasing the number of examples improves
the accuracy; with 3 and 5 examples, we reach 0.44 accuracy. This is probably due to the fact
that the tickets are very varied and a small part of them does not represent the totality of issues
addressed by a macro area. Only with 10 examples we get 0.56 (with a standard deviation of
0.04), approximately the same results obtained by the zero shot approaches.</p>
        <p>The only approach that almost reaches the BERT baseline is the ensemble learning. In this
case, using the best single information zero shot trials, combined with a voting system, helps
the performances to get 0.57 with three prompts (Human, Areas and Categories) and 0.61
with all four prompts; in these cases the standard deviation is lower w.r.t. the BERT baseline.
Also, in these scenarios, we can use all the information described in the zero shot approach
without increasing the prompt length, unlike the full zero-shot case. Probably, with diferent tie
breaking strategies (perhaps based on prediction probabilities) even better performances can be
achieved; unfortunately, at the moment the Chat Competition OpenAI API does not provide
any probabilistic information.</p>
        <p>The second experiment focused on comparing the performance of the two main GPT models
made available by OpenAI: GPT-4, and its predecessor GPT-3.5. Although this can be interesting
from the researcher’s perspective, this comparison has also an economic motivation. In fact,
AREAS</p>
        <p>SUMMARIES</p>
        <p>HUMAN</p>
        <p>CATEGORIES
according to the API pricing, invoking GPT-3.5 costs more than 20 times less than GPT-4. In
this experiment, we keep the same prompts and the other hyperparameters, modifying just the
model. As we can see in Figure 3, the diference in performance is noticeable. The best result
obtained by GPT-3.5 is with the Categories Prompt, with an accuracy of 0.50, less than 5 points
lower than GPT-4; the Human prompt loses approximately 10 points, while the Summaries
prompt drops 15 points. The worst results is obtained by the Areas prompt, with an accuracy of
0.37 and a drop of 20 points. This behaviour is probably due to the long list of percentages for
each single areas and which can confuse the model. No diference was found in the standard
deviations. An average loss of 6 accuracy points between the 3 best performing prompts (thus
excluding Areas) could still justify the cost of using GPT-4 over its older but cheaper counterpart.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions and Future work</title>
      <p>In this work, we have shown an application of pre-trained Large Language Models for the
automatic ticket assignment, based on the real-world data provided by Mega Italia Media, a
company that provides IT solutions in the e-learning context.</p>
      <p>In particular, we analyzed how diferent prompts, with more or less information and examples,
influence the performance on this task. The experimental evaluation shows that in our context
the classical few-shots approach does not provide a considerable improvement. This is probably
because a limited number of examples can’t capture the overall variety of tickets that the
company receives. Instead, a zero-shot approach definitely improves the performance since
it considers more information related to the overall context of the application (for instance a
human description of the diferent classes or the percentage of instances belonging to each
category). The best results are obtained by the ensemble methods involving three and four
prompts. Their results almost reach the ones obtained by a BERT model specifically fine-tuned
for this task. In contrast to BERT, which uses more than 1000 tickets between training and
validation, this approach uses none. This approach can certainly be more eficient in scenarios
where there is a little amount of data available.</p>
      <p>As future work, we aim to study other prompting techniques, such as the chain-of-thought
prompting [32] or prompts automatically generated [15]. Moreover, we would like to explore
the performance of other LLMs, especially testing the open source ones, such as OpenAssistant,
Dolly, GPT-J or GPT-NeoX. This could lead to a more detailed study, not only in terms of
measuring the performance of each model, but also trying to understand what the models know
in this field (which involves information technology, programming languages, e-learning, etc.)
how this knowledge is stored in such models [33, 34], and focus on their explainability, analysing
the behaviour of the attention mechanisms [35, 36] and prevent unwanted or discriminatory
behaviour [37].
[15] T. Shin, Y. Razeghi, R. L. L. IV, E. Wallace, S. Singh, Autoprompt: Eliciting knowledge
from language models with automatically generated prompts, in: Proceedings of the 2020
Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online,
November 16-20, 2020, Association for Computational Linguistics, 2020, pp. 4222–4235.
[16] P. Zicari, G. Folino, M. Guarascio, L. Pontieri, Discovering accurate deep learning based
predictive models for automatic customer support ticket classification, in: Proceedings
of the 36th Annual ACM Symposium on Applied Computing, SAC ’21, Association for
Computing Machinery, New York, NY, USA, 2021, p. 1098–1101.
[17] W. Zhou, W. Xue, R. Baral, Q. Wang, C. Zeng, T. Li, J. Xu, Z. Liu, L. Shwartz, G. Y.</p>
      <p>Grabarnik, STAR: A system for ticket analysis and resolution, in: Proceedings of the
23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,
Halifax, NS, Canada, August 13 - 17, 2017, ACM, 2017, pp. 2181–2190.
[18] P. Molino, H. Zheng, Y. Wang, COTA: improving the speed and accuracy of customer
support through ranking and deep networks, in: Y. Guo, F. Farooq (Eds.), Proceedings of
the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining,
KDD 2018, London, UK, August 19-23, 2018, ACM, 2018, pp. 586–595.
[19] A. DeLucia, E. Moore, Analyzing HPC support tickets: Experience and recommendations,</p>
      <p>CoRR abs/2010.04321 (2020). arXiv:2010.04321.
[20] Q. V. Le, T. Mikolov, Distributed representations of sentences and documents, in:
Proceedings of the 31th International Conference on Machine Learning, ICML 2014, Beijing,
China, 21-26 June 2014, volume 32, 2014, pp. 1188–1196.
[21] J. Han, A. Sun, Deeprouting: A deep neural network approach for ticket routing in expert
network, in: 2020 IEEE International Conference on Services Computing, SCC 2020,
Beijing, China, November 7-11, 2020, IEEE, 2020, pp. 386–393.
[22] L. Feng, J. Senapati, B. Liu, Tadaa: real time ticket assignment deep learning auto advisor
for customer support, help desk, and issue ticketing systems, CoRR abs/2207.11187 (2022).
[23] J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith,
D. C. Schmidt, A prompt pattern catalog to enhance prompt engineering with chatgpt,
CoRR abs/2302.11382 (2023).
[24] L. Reynolds, K. McDonell, Prompt programming for large language models: Beyond the
few-shot paradigm, in: Y. Kitamura, A. Quigley, K. Isbister, T. Igarashi (Eds.), CHI ’21: CHI
Conference on Human Factors in Computing Systems, Virtual Event / Yokohama Japan,
May 8-13, 2021, Extended Abstracts, ACM, 2021, pp. 314:1–314:7.
[25] Y. Zhou, A. I. Muresanu, Z. Han, K. Paster, S. Pitis, H. Chan, J. Ba, Large language models
are human-level prompt engineers, in: The Eleventh International Conference on Learning
Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, OpenReview.net, 2023.
[26] S. S. Biswas, Role of chat GPT in public health, Ann Biomed Eng 51 (2023) 868–869.
[27] T. Mehmood, A. Gerevini, A. Lavelli, I. Serina, Leveraging multi-task learning for
biomedical named entity recognition, in: M. Alviano, G. Greco, F. Scarcello (Eds.), AI*IA 2019
- Advances in Artificial Intelligence - XVIIIth International Conference of the Italian
Association for Artificial Intelligence, Rende, Italy, November 19-22, 2019, Proceedings,
volume 11946 of Lecture Notes in Computer Science, Springer, 2019, pp. 431–444. URL: https:
//doi.org/10.1007/978-3-030-35166-3_31. doi:10.1007/978-3-030-35166-3\_31.
[28] T. Mehmood, A. E. Gerevini, A. Lavelli, I. Serina, Combining multi-task learning with</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bassignana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Polignano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramponi</surname>
          </string-name>
          , Preface to the
          <source>Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI)</source>
          ,
          <source>in: Proceedings of the Seventh Workshop on Natural Language for Artificial Intelligence (NL4AI</source>
          <year>2023</year>
          )
          <article-title>co-located with 22th International Conference of the Italian Association for Artificial Intelligence (AI* IA</article-title>
          <year>2023</year>
          ),
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Revina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Búza</surname>
          </string-name>
          ,
          <string-name>
            <surname>V. G.</surname>
          </string-name>
          <article-title>Meister, IT ticket classification: The simpler, the better</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>193380</fpage>
          -
          <lpage>193395</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Khowongprasoed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Titijaroonroj</surname>
          </string-name>
          ,
          <article-title>Automatic thai ticket classification by using machine learning for IT infrastructure company</article-title>
          ,
          <source>in: 19th International Joint Conference on Computer Science and Software Engineering</source>
          ,
          <string-name>
            <surname>JCSSE</surname>
          </string-name>
          <year>2022</year>
          , Bangkok, Thailand, June 22-25,
          <year>2022</year>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Marcuzzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zangari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schiavinato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Giudice</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gasparetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Albarelli</surname>
          </string-name>
          ,
          <article-title>A multi-level approach for hierarchical ticket classification</article-title>
          , in: Proceedings of the Eighth Workshop on Noisy User-generated
          <string-name>
            <surname>Text</surname>
          </string-name>
          (
          <article-title>W-NUT</article-title>
          <year>2022</year>
          ),
          <article-title>Association for Computational Linguistics</article-title>
          , Gyeongju, Republic of Korea,
          <year>2022</year>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis</article-title>
          , MN, USA, June 2-7,
          <year>2019</year>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <source>Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , et al.,
          <article-title>Language models are unsupervised multitask learners</article-title>
          ,
          <source>OpenAI blog 1</source>
          (
          <year>2019</year>
          )
          <article-title>9</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bosma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Guu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. W.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>Finetuned language models are zero-shot learners</article-title>
          ,
          <source>in: The Tenth International Conference on Learning Representations, ICLR</source>
          <year>2022</year>
          ,
          <string-name>
            <given-names>Virtual</given-names>
            <surname>Event</surname>
          </string-name>
          ,
          <source>April 25-29</source>
          ,
          <year>2022</year>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Kwok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Ni</surname>
          </string-name>
          ,
          <article-title>Generalizing from a few examples: A survey on few-shot learning</article-title>
          ,
          <source>ACM computing surveys (csur) 53</source>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Arici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Gerevini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Putelli</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Serina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sigalini</surname>
          </string-name>
          ,
          <article-title>A bert-based scoring system for workplace safety courses in italian</article-title>
          ,
          <source>in: AIxIA 2022 - Advances in Artificial Intelligence - XXIst International Conference of the Italian Association for Artificial Intelligence</source>
          ,
          <source>AIxIA</source>
          <year>2022</year>
          , Udine, Italy,
          <source>November 28 - December 2</source>
          ,
          <year>2022</year>
          , Proceedings, volume
          <volume>13796</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2022</year>
          , pp.
          <fpage>457</fpage>
          -
          <lpage>471</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Arici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Gerevini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Olivato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Putelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sigalini</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Serina</surname>
          </string-name>
          ,
          <article-title>Real-world implementation and integration of an automatic scoring system for workplace safety courses in italian</article-title>
          ,
          <source>Future Internet</source>
          <volume>15</volume>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .3390/fi15080268.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zubani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sigalini</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Serina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Putelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Gerevini</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Chiari,</surname>
          </string-name>
          <article-title>A performance comparison of diferent cloud-based natural language understanding services for an italian e-learning platform</article-title>
          ,
          <source>Future Internet</source>
          <volume>14</volume>
          (
          <year>2022</year>
          )
          <fpage>62</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12] OpenAI, GPT-4
          <source>technical report, CoRR abs/2303</source>
          .08774 (
          <year>2023</year>
          ). URL: https://doi.org/10. 48550/arXiv.2303.08774. doi:
          <volume>10</volume>
          .48550/arXiv.2303.08774. arXiv:
          <volume>2303</volume>
          .
          <fpage>08774</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          , in: I. Guyon, U. von Luxburg, S. Bengio,
          <string-name>
            <given-names>H. M.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. V. N.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9</source>
          ,
          <year>2017</year>
          , Long Beach, CA, USA,
          <year>2017</year>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. F.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Araki</surname>
          </string-name>
          , G. Neubig,
          <article-title>How can we know what language models know</article-title>
          ,
          <source>Trans. Assoc. Comput. Linguistics</source>
          <volume>8</volume>
          (
          <year>2020</year>
          )
          <fpage>423</fpage>
          -
          <lpage>438</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>