<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Recommender Systems, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Content Extraction in Question Answering Systems through T5 Model Variants</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuxin Luo</string-name>
          <email>yuxin.luo@student.uva.nl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Feng Lu</string-name>
          <email>feng.lu@randstadgroep.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vaishali Pal</string-name>
          <email>v.pal@uva.nl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Graus</string-name>
          <email>david.graus@randstadgroep.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Natural Language Processing, QA system, Domain Information Extraction, Transformer models</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Randstad Groep Nederland</institution>
          ,
          <addr-line>Diemen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Randstad</institution>
          ,
          <addr-line>Diemen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Amsterdam</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>22</lpage>
      <abstract>
        <p>Summarizing usable information from a large number of resumes is a tedious efort for all recruiters. The aim of this study is to explore the performance of the T5 model and its variants for automatic extraction of CV information by combining augmenting manual questions with a paraphraser under the same architecture, and fine-tuning a question and answering system using Dutch and English resumes in a multilingual version of the T5 model (mT5). Our results show that the quality of the generated answers varies considerably between information types, with superior performance for attributes such as basic information that rely on text extraction. However, there is more room for improvement in processing date-based information with multiple inputs, and inferring of multiple standardised answer choices.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>T5
1. Introduction
ing CVs can be seen as a recruiter (internally) posing
questions to a resume, and noting relevant information
related to skills, work experience, and personalia. In this
scenario, unstructured curriculum vitaes (CVs) could be
treated as contextual background for machine
learningpowered question-answering methods. In this paper,
we explore the application of pre-trained large language
models (LLMs) for Question Answering over CVs.</p>
      <p>
        LLMs for QA typically learn to map questions and
contexts to answers. In this paper, we rely on a dataset
tured job seeker data on the other hand. With these two
sources, that respectively represent the context and
”answers” in the QA task, we have but one element missing
to train an LLM for QA: the question, which typically is
hand-engineered or based on templates [
        <xref ref-type="bibr" rid="ref3">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>One of the most popular and advanced models in the ifeld of natural language processing is the Transformer which is a type of neural network architecture based</title>
      <p>nEvelop-O
RecSys in HR’23: The 3rd Workshop on Recommender Systems for
Human Resources, in conjunction with the 17th ACM Conference on
†Work done while on internship at Randstad Groep Nederland.
(D. Graus)</p>
    </sec>
    <sec id="sec-3">
      <title>This paper studies how this mT5 model can perform</title>
      <p>QA to extract information from diferent segments ( Basic
information, Education, Skills) of CVs. Our method
follows two steps: First, we enrich our dataset that consists
of (i) structured job seeker data, and (ii) unstructured job
seeker CVs, by hand-engineering (iii) a small set of
samety of NLP tasks by way of pre-training and fine-tuning.
on self-attention mechanisms. It typically solves a vari- a suficient number of questions for training, we apply a
that consists of job seekers’ CVs on one hand, and struc- of questions to answers. A typical approach is to use
tem© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License ple questions which we (iv) extend through
transformerAttribution 4.0 International (CC BY 4.0).</p>
      <sec id="sec-3-1">
        <title>2.1. Transformer-based</title>
      </sec>
      <sec id="sec-3-2">
        <title>Question-Answering</title>
        <p>based paraphrasing. In doing so, we ensure a large and
diverse enough set of &lt;question, context, answer&gt;-triples.
Second, we use these expanded QA training sets for
several downstream QA tasks.</p>
        <p>In this paper, we aim to answer the following research
question:</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Transformer-based Question-Answering (QA) systems have achieved state-of-the-art performance in a wide range of domains, including open-domain QA, closeddomain QA, and factoid QA.</title>
      <p>RQ1 To what extent can transformer-powered QA be In open-domain QA, systems are required to
anused to meet the basic screening needs in a multi- swer questions from a broad range of topics. Here,
lingual recruitment context? transformer-based QA systems have shown to be
efective in leveraging large-scale pre-training techniques and</p>
      <p>We aim to answer this research question by answering multi-task learning to improve their performance. For
the following sub questions: instance, models like T5, mT5, and ELECTRA [7] have
shown to achieve high accuracy on the open-domain
RQ1.1 Can we apply transformers for paraphrasing man- QA benchmark dataset SQuAD. For specific or
wellually generated questions, to increase the variety documented domains such as wiki or medical data [8, 9],
and volumes of training data for our QA models? they have also demonstrated strong performance.
RQ1.2 How do transformer models deal with diferent In addition, transfer learning has been applied to
facCV segments with diferent types of answers? toid QA, which involves answering questions that require
(e.g. basic information, skills stack, education, a factual answer. For example, models like GPT-3 and
etc.) XLNet have been applied to tasks like reading
compreRQ1.3 How accurate is the QA model in dealing with hension, summarization, and dialog generation, showing
diferent languages of resumes? the ability to generate human-like responses and engage
in natural language interactions.</p>
      <p>Li et al. [10] constructed a CV-related database and
trained it on a multi-turn question-answering system
to obtain valid information. However, it was based on
entity attribution relations and focused only on a limited
number of four types of structured data (name, place of
work, work duration, position), doing experiments on a
dataset of under one thousand resumes. To the best of
our knowledge, there is a few research work conducted in
applying the QA systems in the human resource domain.</p>
    </sec>
    <sec id="sec-5">
      <title>The rest of paper is organized as follows: we discuss</title>
      <p>prior work in transformer-based Question Answering
models, and CV-related datasets in Section 2. Next, in
Section 3 we detail the methodology and overall
experimental design. Then, Section 4 presents the results of
our question paraphrasing approach, and downstream
QA models separately. In Section 5 we reflect on the
experimental results, and discuss the limitations of our
experiments. Finally, in Section 6 we answer our
research questions, and propose possible future research
directions.</p>
      <sec id="sec-5-1">
        <title>2. Related Work</title>
        <p>In recent years, large pre-trained language models such
as BERT [3], GPT-2 [6], and T5 [4] have achieved
stateof-the-art results on many QA benchmarks. T5 employs
a text-to-text approach, taking questions and contextual
text as input and generating answers as output. It
extends transfer learning boundaries by modeling human
language use, understanding the importance of words
in sentences. When comprehending text, attention is
directed to specific words and their meanings. The novel
innovation in T5 includes using a ’prefix’ to specify the
task, which is crucial for this study with two diferent
downstream tasks. The tunability guaranteed by this
feature is particularly important in the context of this
study, where we aim to train two diferent downstream
tasks.</p>
        <sec id="sec-5-1-1">
          <title>2.2. QA in specific domains</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Applying transformers for QA in specific domains often</title>
      <p>faces the challenge of lacking large datasets of
annotated question-answer pairs. Existing solutions to
acquire such annotated datasets include crowdsourcing,
textbook problem repositories, transfer learning, and
domain-specific ontologies.</p>
      <p>Crowdsourcing practices are often used, i.e. manual
annotation of data with a large workforce, which usually
results in a high volume of training data in a short period
of time and a high accuracy rate. However, successful
crowd-sourcing is often run by large companies at a
significant cost and its superiority is currently only proven
in consumer products [11]. In the field we are applying
it to, the need for information from the perspective of
the professional recruiter may not match that of outside
volunteers. And there is no question banks in place that
can meet the demand.</p>
      <p>
        Transfer learning is pre-trained on a large dataset
(open source corpus), and once the model has a better
understanding of the language, it no longer needs to be As shown in the table, the education data represents the
heavily annotated with questions and then put into a level of completed education (as chosen from a list of five
domain-specific dataset after fine-tuning [ 8, 9]. In the education levels), and not e.g., the name of educational
other way, problems are generated in an ontological struc- institute or program. Educational levels are numbered
ture by identifying corresponding concepts and relation- from one to five.
ships through entities in the learning domain. That is, Finally, structured skills data is provided in a
simiconducting a graph-based question generation or rule- lar format. In the process of providing structured data
based question generation [
        <xref ref-type="bibr" rid="ref7">12, 13</xref>
        ]. on skills, job seekers select the appropriate description
      </p>
      <p>Both approaches have domain and coverage limita- from a multiple choice box containing 270 options to
tions. Re-matching question-answer pairs after question demonstrate the skills they possess.
generation is required for QA system input and the data
used in this experiment is bilingual. To address these lim- 3.2. Baseline QA approaches
itations, we apply data augmentation to obtain training
questions, including paraphrased questions generated
using a transformer model after key-value interrogation for
structured data, ensuring question-answer pair matching
for each input question.</p>
      <sec id="sec-6-1">
        <title>3. Methodology</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>In this section, we introduce each step of our QA ap</title>
      <p>proach. Initially, we preprocess the original dataset,
organizing attributes into three subsets. Next, we create
manual questions and apply the template approach to
establish baseline results. Finally, we proceed with question
augmentation as the upstream task for T5 architecture.
After evaluating the paraphrased questions quality, the
QA system is trained in the downstream and the results
are evaluated with our baselines.
3.1. Data</p>
    </sec>
    <sec id="sec-8">
      <title>Our proprietary dataset contains two sets of job seeker</title>
      <p>data, that can be mapped to each other: first, we have
parsed CVs, i.e., textual data extracted from CV source
ifles (e.g., PDF or DOC) through the Amazon Textract
service [14]. It should be noted that these texts are available
in English and Dutch.</p>
      <p>Next, we have structured data, either provided by the
job seekers themselves upon registration or submitted
by recruiters. The structured data includes basic
information such as name, address, contact details, but also
education experience, in addition to skills.</p>
      <p>In this paper, we focus on three diferent information
types: basic information, education, and skills. Table 3.1
shows the features in each module and the key attribute
to link them (candidate_id).</p>
      <p>First, basic information is mostly of textual nature
(i.e., string type data), and spans personal identifiable
information (PII) data such as name, address, and contact
information such as email address.</p>
      <p>The second information type, education data,
contains both textual (string) data (education level
description), and datetime data (education start and end dates).
Microsoft’s Presidio 1 serves as the initial baseline for our
experiment, primarily designed for detection of
personally identifiable information (PII) in text. In this study, we
employ Presidio to identify the basic information in each
CV, namely ’PERSON’ for names, ’DATE_TIME’ for birth
dates, ’EMAIL_ADDRESS’ for email addresses,
’LOCATION’ for addresses, and ’PHONE_NUMBER’ for mobile
phone numbers. It is essential to note that Presidio may
detect multiple entities for a single attribute. To ensure
consistency, we concatenate all detected entities into a
single string-format entity, which serves as the final
extracted entity for comparison with the target text.</p>
      <p>Presidio does not support extracting education and
skills data out of the box. For this reason, we introduce
a keyword-based textual segmentation method as the
second baseline in the experiment. In this approach,
we segment CV texts into diferent sections based on
keywords, and use the underlying sections as extracted
answers. We created a dictionary with section names as
keys, and six keywords lists (three in English and three
in Dutch) as values. For example, for the attribute mobile,
we have keywords telephone, mobile, phone in English
and keywords telefoonnummer, mobiel, contactgegevens
in Dutch.</p>
      <p>We perform a search for these keywords over the
parsed CV texts, marking their starting positions. Once a
ifrst keyword is found, the subsequent keyword is marked
as the first section’s boundary. Using these keyword
positions, we segment the CV text into diferent parts
between consecutive keywords. The text between two
keywords represents the first keyword’s corresponding
section’s content. For specific attributes, we used the
extracted section text as the answer. It’s worth noting that
to enable concise segmentation, we included some
sections that did not cover the specific attributes we focused
on for information extraction (e.g., the “work” keyword
to segment work experiences).</p>
    </sec>
    <sec id="sec-9">
      <title>1https://microsoft.github.io/presidio/</title>
      <sec id="sec-9-1">
        <title>3.3. Data augmentation for question reformulation</title>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>We utilize the PyTorch implementation from Hugging</title>
      <p>Face for fine-tuning, initially employing the pre-trained
mT5 small model. Our training process involves a triplet
of data, namely questions, answers, and contextual text
(parsed CV text).</p>
      <p>The contextual text and questions are input sequences
to the mT5 model, generating implicit representations.</p>
      <p>The answers serve as the target sequence for training.</p>
      <p>Special tokens are added to input sequences, and we
explore two diferent prompting methods:
For training QA approaches, larger numbers of question
and answer pairs typiucally yield better results, even if
there is a saturation efect [ 15], which occurs at around
1,000 samples for transfer learning according to Agrawal
et al. [16]. Therefore, we generate additional questions by
applying transformers for rephrasing manually written
questions.</p>
      <p>The three primary information types, namely basic
information, education and skills, have diferent attributes.</p>
      <p>Basic information comprises five attributes ( name, email,
mobile, birth_date, address), education consists of three at- A The question and context are combined and
tributes (level, start_date, end_date), and the skills module passed to the model. Special tokens, &lt;CLS&gt;
includes a single attribute (skill). (start), &lt;SEP&gt; (split), and another &lt;SEP&gt; (end),</p>
      <p>Attribute values are taken from the structured data, are used to separate the inputs.
and we apply template-based question generation for get- B Context and question are encoded with distinct
ting the questions per attribute. Ten manual questions prefixes. A &lt;RESUME&gt; token marks the start of
are initially generated for each attribute, followed by the context, followed by the CV text, and
&lt;QUESchanging formulations from first to third person to dou- TION&gt; tokens denote the question’s beginning.
ble the number to twenty, adding diversity (e.g., ’What In addition, for contexts, we add addtional prefixes
is your name?’ becomes ’What is the name of the candi- that indicate their languages: &lt;nl&gt; for Dutch, and &lt;en&gt;
date?’). Manual questions undergo review by academic for English. The token &lt;SEP&gt; is also used to divide two
and human resource industry experts. questions when multiple questions are input, which only</p>
      <p>We paraphrase each question using a linguistic happens for education information, where next to level,
rephrasing framework called Parrot [17], which is built we aim to extract start and end dates.
on the fine-tuned paraphraser model in the T5
architecture, and is supposed to be eficient for data augmenta- Example:
tion [18]. Context: ’Name: Mike, age: 30, gender: male’</p>
      <p>For each of the twenty original questions per attribute, Question: ’What is your name? ’
Parrot paraphrases are generated with the ’diversity’ op- Input A: &lt;en&gt; &lt;CLS&gt; Name: Mike, age: 30, gender: male
tion and an ’adequacy’ setting of 0.9, ensuring semantic &lt;SEP&gt; What is your name? &lt;SEP&gt;
similarity. We deduplicate (paraphrased and original) Input B: &lt;en&gt; &lt;RESUME&gt; Name: Mike, age: 30, gender:
questions, and repeat the process three times, yielding male &lt;QUESTION&gt; What is your name? &lt;SEP&gt;
an average of 1,371 questions per attribute.</p>
      <p>For the background input (i.e., CV text) we need to take
into account mT5’s limit on the maximum length of the
3.6.2. QA evaluation
input text, which is 2,048 tokens. The average length of BERTScore relies on a pre-trained BERT language
our CV text is around four 400 tokens, with seventy-five model, using word embeddings to calculate sentence
percent of the data being less than 573 tokens. But there similarity. It emphasizes semantic relevance over
lexare extreme data points with more than 10,000 tokens. ical overlap and selects the BERT-base model based on</p>
      <p>For Basic information and Education information, evaluation requirements.
which we found to usually appear at the beginning of the The ideal solution would exhibit low BLEU, and high
CV, we use a maximum input limit of 512 tokens, i.e., we METEOR and BERTScore, which indicates that the
genuse only the first 512 tokens and truncate the remaining. erated questions difer in syntax and surface form from</p>
      <p>For the Skills module, it is dificult to judge where they the standard manual questions, but remain semantically
are more likely to appear in the CV, so in this case we similar.
use a maximum input limit of 1,000 tokens after which In addition to these performance metrics, we also
calwe truncate. culate the number of questions augmented, and the
vocabulary size diference between the manual and
gen3.5. Experimental setup erated questions, i.e., nQuestions indicates the number
of generated questions, and nWords indicates the
number of unique words added to the question set after the
paraphrasing for each purpose.</p>
    </sec>
    <sec id="sec-11">
      <title>The dataset is divided into a training set, a validation set and a test set in a ratio of 6:2:2. We choose the Adafactor as optimizer under setting suggested by Shazeer and Stern [19].</title>
      <p>Considering the computational costs of resources for
three separate QA models (trained per information type),
unless otherwise stated, each single QA model
trainingvalidation-testing process will be performed on a random
set of 20,000 inputs.</p>
    </sec>
    <sec id="sec-12">
      <title>The three main evaluation metrics, BLEU, METEOR and</title>
      <p>BERTScore, are also used to evaluate the quality of
answers in the QA system in both lexical and semantic
perspectives. The base model of the BERTScore is the
BERT multilingual base model (cased).</p>
      <p>Moreover, there are specific metrics for each
informa3.6. Evaluation tion type. We define EM (Exact Match), for which the
As the question dataset is derived from our our annota- basic information should be most sensitive to, as there
tions, the evaluation is divided into two parts. First, we is only one standard and specific answer for each single
evaluate the question paraphraser, and next, we evaluate question.
the quality of the answers generated by the QA system. In the Education module case, we need to take into
account that a CV may contain information of
multi3.6.1. Question paraphraser evaluation ple educational experiences from diferent periods, and
hence it may take multiple questions at once. When
asThe expectation for the generated questions is that they sessing this, in addition to three metrics above, we define
use a diversity of forms (linguistically diferent) com- PM (Partial Match) to evaluate each separate answer’s
pared to the original questions in terms of wording and quality; we distinguish PM , which represents the level
grammar, but that they are identical in terms of meaning, of the education is extracted exactly, PM , which
indiin the sense that they correspond to the answers to the cates the start date of this education entry is answered
original questions (semantic agreement). In order to ver- correctly, and PM , which denotes the response to the
ify that the generated questions meet the requirements, end date.
we use BLEU [20] (BiLingual Evaluation Understudy), In the skill information, the attention is focused on the
METEOR [21] (Metric for Evaluation of Translation with METEOR score as the structured skill answer is provided
Explicit ORdering), and BERTScore [22] to evaluate the in phrase format, which means that there’s a possibility
generated questions from a linguistic and semantic per- that it is not in the resume text but come out with the
spective respectively. candidate’s former experience or minds.</p>
      <p>BLEU measures sentence similarity via word overlap
between generated and reference sentences. Scores range
from 0 to 1, with higher scores indicating greater word 4. Results
usage similarity. We use n-grams of size 1 and 2 due to
short question lengths (5-15 words).</p>
      <p>METEOR assesses sentence similarity by considering
word matching, word order correctness, and word sense
matching, capturing semantic similarity and relatedness.
input format (Input), training configuration (batch size
(ba), and number of epochs (ep)), lexical and semantic
matching quality of answers (BLEU, METEOR, BERTScore,
EM, PM ).</p>
    </sec>
    <sec id="sec-13">
      <title>For the basic information part, we evaluated the per</title>
      <p>formance of three methods: rule-based, Presidio, and
our finetuned model, using diferent evaluation metrics:
BLEU, METEOR, BERTScore, and Exact Match (EM).
4.1. Question paraphraser First, we turn to top rows in Table 3 for the answer
exThe results under attribute All represent all questions as traction performance comparison between the diferent
a whole, regardless of information type or attribute, and methods .
compares and evaluates the rephrased question set with Starting with the rule-based method, it achieved
modthe original question set. est results across the evaluation metrics. It attained low</p>
      <p>The overall quality of the generated questions pre- BLEU and METEOR scores of 15.71%, 9.26% and 28.45%.
sented in the Row 1 of the table 2 suggests relatively low The BERTScore of 69.67% is acceptable, and an Exact
similarity to the manual questions in terms of words and Match (EM) score of 19.6%. Although the rule-based
syntax, with BLEU scores below 50, while exhibiting high method served as a foundational baseline, its
perforsemantic similarity. mance revealed limitations in capturing the intricacies of</p>
      <p>Specifically, we observed results for the following eval- language and extracting information with high precision.
uation metrics; we find an inevitable positive correlation The Presidio method, on the other hand, exhibited
between lexical overlap and semantic similarity, with noticeable improvements over the rule-based technique.
Skills having the lowest BLEU score, and also bearing the Especially, it garnered higher BLEU scores of 36.99% and
only BERTScore under 50. We find that several other at- 27.31%, and 31.4% of the answers were given accurately
tributes (e.g., Mobile, Start date, End date) exhibit slightly (EM). Presidio’s ability to incorporate context and
contexhigher word overlap (i.e., BLEU scores) than others. For tually aware patterns enabled it to surpass the rule-based
the number of questions generated (nQuestions), we find method in all respects.
that the number of Birth attribute questions is just be- However, it is important to note that despite the
imlow 1,000, despite taking exactly the same settings and proved performance of Presidio compared to the
rulesteps as the other attributes, which yield roughly be- based approach, the fine-tuned model with suitable
contween 1,100 and 1,600 additional questions. In addition, figuration outperformed both of them in all metrics. In
for each attribute, we find that the act of expanding the particular, the best-performing results in the fourth row
question sets by paraphrasing, substantially increases the are a BLEU-1 score of 90.64 and a BERTScore of a
remarkvocabulary of the questions (nWords). able 98.38 percent. These scores were significantly higher
than those obtained by both the rule-based method and
Presidio, indicating the clear superiority of the
trans4.2. Question Answering former approach in text processing and information
extraction tasks.</p>
      <p>Rows three to ten in the table 3 show how the
trans</p>
    </sec>
    <sec id="sec-14">
      <title>In this section, we proceed to describe the results of our methods per information type.</title>
      <p>former model behaves with diferent hyperparameter ing the CVs and obtaining a small whole snippet of
educasettings. For our mT5 model, the best input format is tional information without dedicated rules for linking the
input A (with &lt;CLS&gt;&lt;SEP&gt; token encoding), which gen- time and education level within it, so no corresponding
erally performs better than input B for the same training matching scores about time can be derived.
settings. By fine-tuning the training configuration of the For the transformer experiments of the education
quesmodel and allocating computational resources, we were tions, the model configuration of four batches and two
able to achieve optimal answer quality in all scores by epochs still outperformed others. However, with
edusetting the batch size and number of epochs to four and cation, the diferent input formats resulted in smaller
two, respectively, with over 90% overlap of individual diferences than with the basic information QA model.
words and 87.02% accuracy in exact answer matching. Input B (&lt;RESUME&gt;&lt;QUESTION&gt; token encoding) only</p>
      <p>In summary, while Presidio showed an improvement performed slightly better than input A when we look at
over the rule-based method, the transformer method with the first partial match, meaning that input B may have
the trained hyperparameters demonstrated the best per- helped to better capture information on the level of
educaformance overall, making it the most suitable choice for tion in the CV. For either input method and model setting,
answering questions about the basic information. the accuracy of answers for education level (PM ) was
significantly higher than for start date (PM  ) and end
4.2.2. Education date (PM ). Moreover, in the training of this module,
the increase in batch size was to some extent detrimental
In this education module analysis, we compared two dis- to the scores of each response quality.
tinct methods: the traditional rule-based approach and
the modern transformer method, using BLEU, METEOR, 4.2.3. Skills
BERTScore, and Partial Match (PM) evaluation metrics.</p>
      <p>In the first row of this module, The rule-based method In the skill module, we evaluated the performance of
yielded results with low BLEU scores of 17.99 and 8.89, two methods: the rule-based approach and the advanced
a METEOR score of 23.92, a BERTScore of 67.87, as it transformer method, utilizing key evaluation metrics:
demonstrates similar capabilities in the basic informa- BLEU, METEOR, and BERTScore.
tion module. PM indicates that the answer obtained Both methods, unfortunately, fell short of our
expeccorrectly matches the level of education currently en- tations in terms of overall performance. The rule-based
tered. The rule-based method ofered a foundation for approach has a much lower eficiency in extracting skill
text processing to achieve accuracy at 34.87%. However, information than the above two modules. BLEU-1 and
the main idea of the rule-based approach lies in segment- METEOR scores drop to single digits, 1.63 and 1.75
respectively. The transformer method, while showcasing better the data will be useful in increasing the diversity and
results, still left room for improvement, with scores of coverage of the training sample. At the same time, the
BLEU-1 12.38%, METEOR 12.71%, and BERTScore 70.58%. high semantic similarity of the rephrased questions to the</p>
      <p>For the fine-tuning of the transformer model in the human questions will help to improve the performance
skills section, input A (&lt;CLS&gt;&lt;SEP&gt; token encoding) con- of the QA system in terms of semantic understanding
tinued to show its strengths, although the overall lexical and answering questions.
matching scores were not as good as the two modules See an example of a paraphrased question in Table 4
above, i.e., the task is harder. When training the model, in diferent similarity, given as input question: What’s
this module took much longer to run on the same amount the candidate’s name?. With this example, we can more
of data. intuitively feel that only high semantic similarity
paraphrasing can satisfy the downstream training needs, i.e.,
4.2.4. Multilingual performance the left column of the table. At the same time, low
linguistic similarity, paired with high semantic similarity,
will efectively diversify the question set, thus ensuring
the robustness of our Q&amp;A system.</p>
      <sec id="sec-14-1">
        <title>5.2. Overview of transformer models performance</title>
        <p>There were diferences in the quality and accuracy of
answer generation between languages, with Dutch CVs
scoring slightly higher than English in both the basic
information (with 90.98 vs. 85.71 BLEU-1 respectively) and
education modules (with 59.78 vs. 50.56 BLEU-1
respectively) under the same model. For the three attributes
in the education module, however, the model showed
homogeneity across languages for the questions related
to education start and end dates.</p>
        <p>As we mentioned in Section 3, all prefix tokens and
questions were in English. However, the overall quality
of the responses is still slightly better in Dutch than in
English, with the Dutch CV even scoring 9.36% higher
than the English CV on the measure of exact match in
basic module.</p>
        <p>We used a suficient number (over 1,000 for each
attribute) of questions to fine-tune the mT5 model.
Initially, we expected transformer models to outperform
baseline methods due to their ability to learn complex
patterns in textual data. However, experimental results
varied. Tweaking model settings and input forms
significantly improved the answer quality scores for basic
information. Under the 4-batch and 1-epoch
configuration, transformer models performed worse than baseline
methods, possibly due to limited exposure to diverse CV
5. Discussion data. Increasing the batch size improved transformer
model performance, allowing for more eficient parallel
In this section, we discuss question augmentation first, processing during training. Increasing the number of
and then the performance comparison between the mT5 epochs to 2 with a batch size of 4 had a positive impact,
transformer model, and our two baselines: Presidio for enabling the models to refine their learned
representabasic information, and the rule-based method, following tions and capture finer patterns in resume data. This
by the diferent features of transformer model across the resulted in improved the exact match accuracy to 87%
modules and language. under input 1 format (&lt;CLS&gt;&lt;SEP&gt; token encoding).
Interestingly, when comparing the performance of
5.1. Question augmentation usage models trained with 4 batches and 2 epochs to those
trained with 4 batches and 4 epochs, the former
perTable 2 shows the result of the comparison between man- formed better. This suggests that after a certain point,
ual and paraphrased questions. The expanded data gen- increasing the number of epochs may not yield
signifierated by the paraphraser is able to meet the needs of cant performance improvements and may even lead to
downstream question and answer system training. Man- over-fitting on the training data in basic information
ual inspection revealed that the paraphrased questions module.
vary, but largely maintain the meaning of the input ques- For both the education and skills information types,
tion. The relatively large number of rephrased questions the underlying transformer models went beyond the
ruleis due to them in some cases not being a complete change based baseline. The diferent input forms and the
configof questioning style to the original question, but a simple uration of model training all had no significant efect on
modification of the word abbreviation or tone. In other the acquisition of educational information.
cases, the question departs completely from the input Under the skills module, input A performs significantly
question in textual overlap (e.g., ”how should I call you?” better than input B.
given as input: ”what’s the candidate’s name?”). Furthermore, the potential of the models that we have</p>
        <p>As the rephrased questions have low similarity in word fine-tuned may be wider than the evaluation results
sugsyntax to the manual questions, it can be expected that gest. For example, a generated answer to a question about
tell me the name of the candidate
tell me the email of the candidate
address is Heilige Geeststraat, while the target answer (i.e., The CV’s Skills information posed the most
signifiground truth) is Eilige Geeststraat, which means the gen- cant challenge in comparing diferent information types
erated answer does not yield an exact match. In fact, we vertically.
found that there is no street named Eilige Geeststraat but Semantic and lexical similarity scores between
modelthe street Heilige Geeststraat is the correct street name. generated and target answers diverge widely, despite
In certain instances, individuals may include only partial having relatively similar semantics and low lexical
overaddress information in their CVs. Through the utilization lap. The diversity in skill descriptions likely contributes
of the fine-tuned mT5 model, it becomes evident that this to this outcome. During structured data entry, candidates
technology adeptly facilitates the precise extraction of choose skill descriptions from 270 options with diferent
said partial address information. granularities, e.g., ”Microsoft Word” or ”Ofice Suite”
be</p>
        <p>The original CV mentions the existent one, which ing skills that partly overlap, which leads to variations
means there is a possibility that the ground truth pro- in how similar skills are expressed.
vided by the candidate or filled by the recruiter are in- Personal preferences for customizing descriptions also
correct due to, e.g., a typo, but our fine-tuned mT5 QA pose a challenge. The low word overlap score in the
model demonstrates the ability to extract the correct in- segmentation method suggests a wide range of
strucformation from the original text. tured options, allowing room for personal preferences.
Candidates select descriptions that match their
experi5.3. Modules and Language-specific ence and understanding, which may not align with the
model’s training data. As a result, vocabulary similarity</p>
      </sec>
      <sec id="sec-14-2">
        <title>Adaptations</title>
        <p>in generated answers varies significantly.</p>
        <p>Our experiments focused on a question and answer sys- In such cases, the model must select the most
approtem for three CV information types: Basic information, priate skill description based on context and question,
Education, and Skills. The transformer model excelled at requiring additional reasoning and comprehension.
obtaining Basic information, demonstrating high accu- The performance of the fine-tuned model in terms of
racy and strong lexical and semantic similarity to original answer generation for both languages showed little
difanswers, likely due to direct extraction without complex ference in semantic similarity scores in the same module,
reasoning. For Education information, the model’s ac- with Dutch slightly outperforming English. However, for
curacy varied; it performed better for education level text-based question answer accuracy, basic information
questions than for date-related queries. Date data’s vari- and education level, the model outperformed English in
ous formats contributed to this discrepancy. Obtaining terms of extracting and reasoning about information in
educational information involved both transforming date Dutch (9.36% and 14.22% higher respectively). This has
data and mapping educational descriptions to levels, with to be attributed to the data itself, where the proportion of
CVs containing multiple instances of educational infor- raw data in the dataset is 90% for Dutch, and around 10%
mation. The model’s self-attentive mechanism might for English; and where the description of education level
not fully capture the sequential nature of numeric data, is a common expression under the Dutch education
sysas it was originally designed for natural language text. tem. This results in our model being better at capturing
The uniform date format in structured data difered from the textual information in Dutch CVs.
real CVs, where dates are often imprecise, leading to a
large amount of data ending on the first or last day of the 5.4. Limitations
month. The Education module achieved high semantic
similarity but low lexical similarity to original answers, Despite the results achieved in this study in the task of
using diferent expressions to understand and respond extracting CV information in the form of questions and
to questions rather than directly copying phrases from answers, there are still some limitations that need to be
the original text [23]. considered.</p>
        <sec id="sec-14-2-1">
          <title>6. Conclusion</title>
          <p>Firstly, there was a data imbalance in the training performed superior in question answering of basic
infordata in terms of language, as described above, where our mation, with high accuracy and high semantic similarity
dataset had an over-representation of Dutch CVs. This to the original answers.
may have some impact on the model’s performance in However, for education information, the model was
the case of English CVs or when English skill information more accurate for education level-related questions than
is included. The dominance of structured information in for education start or end date-related questions, and
Dutch phrases in the skills module may lead to dificul- scored high semantic similarity but relatively low lexical
ties and reduced performance of the model when dealing similarity. This may be due to the fact that the
acquisiwith English skill descriptions. tion of educational information involves reasoning about</p>
          <p>Secondly, annotation limitations also need to be con- educational descriptions (institution names or degrees
sidered. The annotation process for the skills module with levels) and the diversity of date data in formats may
involves candidates selecting appropriate descriptions pose a challenge to the model.
through multiple choice boxes, and this subjective na- In addition, in skills information, the semantic
simiture of the labelling process may lead to diferences in larity scores and lexical similarity scores of the
modelpersonal preferences and expressions. Inconsistencies in generated answers difered significantly from the target
the labelling criteria may afect the lexical overlap and answers, with similar semantics but low lexical overlap,
semantic similarity scores of the answers generated by and the overall performance was not as good as the first
the model. two modules. This may be due to the subjective nature of
the annotation process, resulting in diversity and lexical
variation in the generated answers.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-15">
      <title>In this study, we explored an approach combining up</title>
      <p>stream and downstream tasks, augmented by a
paraphraser of the T5 model for manual questions, and
finetuned using the mT5 model. We experimented with the
CV information extraction task in question-and-answer
format for three diferent information types, and
evaluated the performance of the model. Based on the results
and discussions we obtained, we are able to provide
answers to the research questions.</p>
      <sec id="sec-15-1">
        <title>6.1. RQ1.1: Transformers for dual tasks</title>
        <p>The questions generated by the paraphraser under the T5
architecture are suficient in number, have low overlap
with the manual question vocabulary and are
semantically similar. Thus the need for fine-tuning the question
and answer system for the mT5 model can be met. By
tuning the hyperparameter settings and input forms of
the training model, we found that for intercepted
answers, the tuned transformer model was able to exploit
its ability to learn complex textual contexts and thus
far outperform the keyword-segmentation approach. In
the hyperparameter setting, increasing the number of
epochs trained outperformed a larger batch of models,
for the same conditions. Labeling the
context-questionends with the same token helps the model to understand
the context better.</p>
      </sec>
      <sec id="sec-15-2">
        <title>6.2. RQ1.2: Information types</title>
      </sec>
    </sec>
    <sec id="sec-16">
      <title>Through comparison and analysis between the diferent information types, we found that the fine-tuned model</title>
      <sec id="sec-16-1">
        <title>6.3. RQ1.3: Cross-lingual performance</title>
      </sec>
    </sec>
    <sec id="sec-17">
      <title>In terms of language, the model maintained a relatively</title>
      <p>good semantic similarity, although the fact that there
was a data imbalance in our dataset prevented us from
drawing absolute conclusions and the accuracy of the
model’s extractive answers on English CVs decreased.</p>
      <sec id="sec-17-1">
        <title>6.4. Future work</title>
      </sec>
    </sec>
    <sec id="sec-18">
      <title>The possibilities for future work are varied. One is multi</title>
      <p>lingual support, where future research could expand the
dataset to include more samples in multiple languages to
improve the adaptability and generalisation of the model
to more linguistic contexts.</p>
      <p>Another direction is to address the performance
differences between diferent modules in the CV
information extraction task. For information extraction on skills,
consider trying to manipulate their terms as mentioned
by Smith et al. [24] or drawing on external references
(Wikipedia, LinkedIn) as Kivimäki et al. [25] did to make
skill annotation standards consistent.</p>
      <p>The exploration of utility is also worth looking at. The
data we have used is a generic CV dataset, but practical
applications may face domain-specific or firm-specific
needs. Future research could therefore explore domain
adaptive techniques to enable models to be adapted to
the needs of diferent domains or companies.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>tics</surname>
          </string-name>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>5307</fpage>
          -
          <lpage>5315</lpage>
          . URL:
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          https://aclanthology.org/D19-1534. doi:
          <volume>10</volume>
          .18653/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          v1/
          <fpage>D19</fpage>
          - 1534. [24]
          <string-name>
            <given-names>E.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Weiler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          , Skill extraction
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Multimodality</surname>
          </string-name>
          , and Interaction: 12th International
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>Conference of the CLEF Association</article-title>
          ,
          <year>CLEF 2021</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Virtual</given-names>
            <surname>Event</surname>
          </string-name>
          ,
          <source>September 21-24</source>
          ,
          <year>2021</year>
          , Proceedings
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          12, Springer,
          <year>2021</year>
          , pp.
          <fpage>116</fpage>
          -
          <lpage>128</lpage>
          . [25]
          <string-name>
            <given-names>I.</given-names>
            <surname>Kivimäki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dessy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Verdegem</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>of TextGraphs-8 graph-based methods for natural</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>language processing</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>