<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Recommender Systems, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Résumé Parsing as Hierarchical Sequence Labeling: An Empirical Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Federico Retyk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hermenegildo Fabregat</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Aizpuru</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariana Taglio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rabih Zbib</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>this system in production environments. In summary</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>22</lpage>
      <abstract>
        <p>Extracting information from résumés is typically formulated as a two-stage problem, where the document is first segmented into sections and then each section is processed individually to extract the target entities. Instead, we cast the whole problem as sequence labeling in two levels -lines and tokens- and study model architectures for solving both tasks simultaneously. We build high-quality résumé parsing corpora in English, French, Chinese, Spanish, German, Portuguese, and Swedish. Based on these corpora, we present experimental results that demonstrate the efectiveness of the proposed models for the information extraction task, outperforming approaches introduced in previous work. We conduct an ablation study of the proposed architectures. We also analyze both model performance and resource eficiency, and describe the trade-ofs for model deployment in the context of a production environment.</p>
      </abstract>
      <kwd-group>
        <kwd>Sequence labeling</kwd>
        <kwd>deep learning</kwd>
        <kwd>résumé parsing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Résumé parsing has become increasingly important in</title>
        <p>volves the extraction of relevant information about a
candidate from their résumé document into a structured
data model. The extracted information, when integrated
with downstream recommender systems, can in turn help
candidates and recruiters optimize their search.</p>
        <p>
          Résumés1 exhibit significant diversity in their format
and in their use of language due to diferences in the
background, industry, and location of the candidates [
          <xref ref-type="bibr" rid="ref15">1</xref>
          ]. But
despite their unstructured nature, these documents are
typically organized into sections. Each one of these text
blocks contains important details about the candidate,
such as their personal information, education history,
ther divided into groups. Within particular sections and
groups are diferent concepts or entities that are relevant
to the recruitment process. For example, the full name
of the candidate, their address, and their phone number
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Each group within the work experience section usually</title>
        <p>includes a period, a job title, and the employer’s name.</p>
        <p>Since résumé parsing occurs early on in the digitalized
icant efect on that of the downstream recommender
systems. However, the aforementioned diversity in the
use of language in résumés makes the parsing problem
cial intelligence (AI) approaches. Efective solutions for
this task, therefore, require the use of machine learning
techniques.</p>
      </sec>
      <sec id="sec-1-3">
        <title>Existing industry-scale solutions for résumé parsing do</title>
        <p>not make public detailed information about their systems.</p>
      </sec>
      <sec id="sec-1-4">
        <title>On the other hand, previous academic research in this</title>
        <p>domain focuses on constrained scenarios that are limited
in scope, in the complexity of the target label scheme, or
in terms of the size and quality of the annotated datasets.</p>
      </sec>
      <sec id="sec-1-5">
        <title>Moreover, these works address the problem in two or</title>
        <p>
          more stages. In the first stage, they segment the résumé
into sections and groups [
          <xref ref-type="bibr" rid="ref15 ref3">1, 2, 3</xref>
          ]. Since résumés are
sequence labeling model to extract the target entities
from the text of each section.
        </p>
      </sec>
      <sec id="sec-1-6">
        <title>In this work, we propose a joint model that labels the</title>
        <p>full document as a whole. This is an unusual setting in
long text sequences and the set of labels is relatively big.</p>
      </sec>
      <sec id="sec-1-7">
        <title>We show that the proposed system is not only eficient</title>
        <p>and convenient from an engineering point of view, but
compare it to previous approaches and we also study
several design-decisions of our system in terms of their efect
on accuracy as well as in time and memory eficiency.</p>
        <p>We share experimental observations on résumés in seven
tionally, some sections depict a chronological progression
previous work experience, and professional skills. Addi- long text documents, this is generally approached as text
classification of independent lines without
document(e.g. work experience, education), and are naturally fur- level context. The second stage uses a section-specific
are typically present in the contact information section. academic literature for sequence labeling, as résumés are
recruitment process pipeline, its accuracy has a signif- it is also competitive with the two-stage alternative. We</p>
        <sec id="sec-1-7-1">
          <title>RecSys in HR’23: The 3rd Workshop on Recommender Systems for</title>
        </sec>
        <sec id="sec-1-7-2">
          <title>Human Resources, in conjunction with the 17th ACM Conference on</title>
          <p>
            • Casting the task of résumé parsing as hierarchical Yadav and Bethard (2018) and Li et al. (2022) for a more
sequence labeling, with line-level and token-level comprehensive review of deep neural networks for
seobjectives, and presenting an eficient résumé quence labeling [8, 9].
parsing architecture for simultaneous labeling Prior work on parsing résumés usually divides the
at both levels. We propose two variants of this problem into two tasks, and tackles each separately [1,
model: one optimized for latency and the other 2, 3, 10, 11]. The résumé is first segmented into sections
optimized for performance. and groups, and then section-specific sequence labeling
models are applied to extract target entities. The early
• A comprehensive set of experiments on résumé work by Tosik et al. (2015) focuses on the second task
parsing corpora in English, French, Chinese, only, as they experiment with already-segmented
GerSpanish, German, Portuguese, and Swedish, each man résumés [
            <xref ref-type="bibr" rid="ref15">1</xref>
            ]. They train named entity recognition
covering diverse industries and locations. We models for the contact information and work experience
share our experience in the process of developing sections, each with a small set of labels. The architecture
such annotations. These experiments compare they apply uses word embeddings as direct features for
our proposed system to previous approaches and the CRF.
include an extensive ablation study, examining Zu et al. (2019) use a large set of English résumés
colvarious design choices of the architecture. lected from a single Chinese job board to experiment
• Insights into the process of deploying this model with several architectures for each of the two stages [2].
in a global-scale production environment, where For segmentation, they classify each line independently
candidates and recruiters from more than 150 (without document context). Then to extract entities,
countries use it to parse over 2 million résumés they train diferent models for each section type. The
per month in all these languages. We analyze the input to these sequence labeling models is the text of
trade-of between latency and performance for each independent line. While for the line classification
the two variants of the model we propose. task they use manually annotated samples, the sequence
labeling models are trained using automatic annotations
          </p>
          <p>
            Our empirical study suggests that the proposed hi- based on gazetteers and dictionaries.
erarchical sequence labeling model can parse résumés Barducci et al. (2022) work with Italian résumés. They
efectively and outperform previous work, even with a ifrst segment the résumé using a pattern-matching
aptask definition that involves labeling significantly large proach that relies on a language- and country-specific
text sequences and a relatively large number of entity dictionary of keywords [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ]. After this, they train
indelabels. pendent sequence labeling models for each section type.
The architecture they use for the sequence labeling
component is based on the approach described above that
2. Related Work uses BERT [7] with a classification layer on top.
Finally, Pinzon et al. (2020) work with a small corpus of
Our work builds upon prior research on deep learning English résumés [
            <xref ref-type="bibr" rid="ref6">12</xref>
            ]. They bypass the segmentation task
for sequence labeling, specifically those applying neu- (ignoring sections and groups) and propose a model that
ral networks in combination with Conditional Random directly extracts entities from the résumé text. They use a
Fields (CRFs) to various sequence labeling tasks. Huang BiRNN+CRF model for the token-level sequence labeling
et al. (2015) investigated an architecture based on Bidirec- task. Among the related work we examined, this is the
tional Recurrent Neural Networks (BiRNNs) and CRFs [4]. only one that made their dataset public. Nevertheless,
They use both word embeddings and handcrafted features a manual examination of the corpus led us to conclude
as initial representations. Lample et al. (2016) extended that the sample is far from representative of real-world
this architecture by introducing character-based repre- English résumés and that the labeling scheme they use
sentations of tokens as a third source of information for is limited and inadequate for our scope.
the initial features [5]. An alternative character-based We extend the previous work by exploring a joined
approach was proposed by Akbik et al. (2018), which uses architecture that predicts labels for both lines and tokens,
a BiRNN over the character sequence to extract contextu- treating each as a sequence labeling task. Furthermore,
alized representations that are then fed to a token-level as in Pinzon et al. (2020) [
            <xref ref-type="bibr" rid="ref6">12</xref>
            ], we unify the extraction of
BiRNN+CRF [6]. In addition, Devlin et al. (2019) intro- entities for any section. This setup is challenging, since
duce a simple Transformer-based approach that avoids résumés are unusually long compared to typical
Inforthe utilization of CRF. This consists of a pre-trained BERT mation Extraction tasks, and the set of labels for entities
encoder, which is fine-tuned, followed by a linear clas- is also bigger. But the advantage is the improvement of
sification layer applied to the representation of each to- eficiency in terms of execution time and memory usage,
ken [7]. We refer interested readers to the surveys by and the simplification of the engineering efort since only
Corpus
English
French
Chinese
Spanish
German
Portuguese
Swedish
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Task Description</title>
      <sec id="sec-2-1">
        <title>We cast résumé parsing as a hierarchical sequence labeling problem, with two levels: the line-level and the token-level. These two tasks can be tackled either sequentially or in parallel.</title>
        <p>For the first, we view the résumé as a sequence of
lines and infer the per-line labels that belong to diferent
section and group types. This is a generalization of the
task definition used in previous work, where the label
(class) for each line is inferred independent of information
about the text or the predicted labels of other lines. We
assume that section and group boundaries are always
placed at the end of a line, which is the case in all the
résumés we came across during this project. The label
set for this part of the task includes a total of 18 sections
and groups, which are listed in Appendix A.1.</p>
        <p>For the second level, we view the résumé as a long
sequence of tokens that includes all the tokens from every
line concatenated together. We infer the per-token labels
that correspond to the diferent entities. The label set
for this part of the task includes 17 entities, which are in
turn listed in Appendix A.2.</p>
        <p>The scope of this paper revolves around the extraction
task and therefore we do not focus on the conversion of
the original résumé (e.g. a docx or pdf file) into plain text
format. Rather, the systems studied in this work assume
textual input.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Since this efort is aimed at building a real-world ap</title>
        <p>plication, annotation quality is highly important. For
that purpose, we implemented a custom web-based
annotation tool that allows the user to annotate section and
group labels for each line of a résumé, and to annotate
entity labels for each arbitrary span of characters.</p>
        <p>We developed the annotation guidelines by starting
with a rough definition for each label and performing
exploratory annotations on a small set of English résumés
—a mini-corpus that we later used for onboarding the
annotators. The guidelines were then iteratively refined
for the whole duration of the project, achieving a stable
and rigorous version at the end. In Appendix A we define
the section, group, and entity objectives covered in our
corpora, and we provide a screenshot of the annotation
tool user interface for reference.</p>
        <p>Each language corpus was managed as an independent
annotation project. We recruited 2 or 3 annotators, who
are native speakers of the target language and without
specifically seeking domain expertise, through an online
freelance marketplace. The annotators did not
communicate with each other during the process, maintaining the
independence of the multiple annotations. Before
starting the annotations on the target corpus, we asked each
4. Corpora annotator to carefully read the guidelines, and annotate
the onboarding English mini-corpus. After reviewing
We built résumé parsing corpora in English, French, Chi- and providing feedback, the annotator was instructed to
nese, Spanish, German, Portuguese, and Swedish. Some annotate all the résumés in the target corpus.
statistics on the corpora are reported in Table 1. For each The estimated inter-annotator agreement (IAA) for
of these languages, résumés were randomly sampled from the corpus in each language, computed as suggested by
public job boards, covering diverse locations and indus- Brandsen et al. (2020) [13] in terms of  1, ranges from
tries. For all but Chinese, we controlled the sampling 84.23 to 94.35% and the median is 89.07%. Finally, we
adprocess in order to enforce diversity in locations. For ex- judicated the independent annotations in order to obtain
ample, although the English corpus is biased toward the the gold standard annotations. This process involved
USA, there is a fraction of résumés from other English- resolving any conflicting decisions made by individual
speaking countries including the UK, Ireland, Australia, annotators through the majority voting method. In cases
New Zealand, South Africa, and India. Although we did where a majority decision was not attainable, the
adjunot control for industry variability, we observe a high dicator was instructed to review the decisions of each
level of diversity in the selected collections. We then annotator and apply their own criteria to arrive at a final
used third-party software to convert into plain text the decision.
original files, which came in varied formats such as pdf,
doc, and docx.</p>
        <p>Text Sequence</p>
        <p>Input
Text Sequence</p>
        <p>Input
Text Sequence</p>
        <p>Input</p>
        <p>Token-level
Initial Features
Embeddings or 
Transformer
Token-level
Initial Features
Embeddings or 
Transformer
Token-level
Initial Features
Embeddings or 
Transformer</p>
        <p>Token-level
Representations</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Model Architecture and</title>
    </sec>
    <sec id="sec-4">
      <title>Training</title>
      <p>each line are concatenated to obtain the input sequence
for the BiRNN+CRF architecture. This is visually
described in Figure 2. Preliminary experiments, which are
The models we use in this work are based on the not presented here because of space constraints, showed
BiRNN+CRF architecture. Initial features are first ex- that avoiding the BiRNN component for this last
architracted for each token, then combined through bidirec- tecture, i.e. applying CRF directly on the output of the
tional recurrent layers, and finally passed through a CRF Transformer-based features, obtains markedly worse
relayer to predict the labels. Unless specified otherwise, sults. This is because the two layers capture
complementhe input to the model is the entire résumé text after ap- tary aspects of the context: the Transformer encodes
plying tokenization. We study two design-decisions: (1) tokens by exclusively considering the context of the
curthe choice for initial features, and (2) separate models for rent line, while the BiRNN layer on top contextualizes
predicting line and token labels vs. a multi-task model across every line. Because of the typical length of a
réthat predicts both jointly. sumé in terms of tokens, we did not explore encoding
Initial features. We explore two alternatives: the whole résumé at once with the Transformer encoders
used in this work.</p>
      <p>Single-task vs. Multi-task. We experiment with:
(a) A combination of FastText [14] word embeddings
and handcrafted features, which are detailed in</p>
      <p>Appendix B.
(b) Token representations obtained from the encoder
component of a pre-trained T5 [15] model (or an
mT5 [16], depending on the language) without
ifne-tuning.</p>
      <p>(a) Single-task models that perform either line-level
sequence labeling (sections and groups) or
tokenlevel sequence labeling (entities).
(b) Multi-task models that predict labels for both
linelevel and token-level tasks simultaneously.</p>
      <sec id="sec-4-1">
        <title>The T5 models are based on the Transformer [17] ar</title>
        <p>chitecture. For this second case, each line is encoded
individually2, and then the token representations for</p>
      </sec>
      <sec id="sec-4-2">
        <title>2Note that résumés are long text sequences, usually longer than 512</title>
        <p>tokens (see Table 1).</p>
        <p>Figure 1 illustrates the model variants. The
architecture shown in Figure 1a is a single-task model for
linelevel objectives (sections and groups). This architecture
takes as input the complete sequence of tokens in the
résumé and predicts one label for each line. We train
...
...
ta
cn
oC
coCn
Transformer</p>
        <p>Transformer</p>
        <p>Transf.
x1 x
2 x
3 x
4 x
5 x
6 x
7 x</p>
        <p>8 x9 .. xT-1 xT
Line 1</p>
        <p>Line 2</p>
        <p>Line 3</p>
        <p>Line L
token: r = ⃖h⃗


⊕⃖h⃖</p>
        <p>. The result is a sequence of line repre- commercial setting. From an operational perspective, the
sentations R = (r1, r2, … , r ), which is in turn processed
by another BiRNN layer. This aggregation mechanism is
training, testing, integration, and maintenance of a single
model is simpler and cheaper than for two models.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Experiments</title>
      <sec id="sec-5-1">
        <title>We next describe the results of our experiments using the</title>
        <p>corpora of Section 4. The main results are summarized in
for training and report the micro-average  1 scores (for
the positive labels only) on the held-out 10%3. The
results compare the two model architectures discussed in</p>
      </sec>
      <sec id="sec-5-2">
        <title>Section 5: Single-task and Multi-task, and for each archi</title>
        <p>tecture, the two alternatives for initial features: FastText
and Transformer-based T5.</p>
        <p>The  1 scores for the token sequence labeling task
(predicting entities) are reported in Table 2a. Those include
the results for the two single-task models that act only
on the token-level task, as well as the two multi-task
models. The  1 scores for the line sequence labeling task
(sections and groups) are shown in Table 2b, again for
the two single-task models that act only on the line-level
task, and the two multi-task models4.</p>
      </sec>
      <sec id="sec-5-3">
        <title>We make some observations. Comparing row 1 with</title>
        <p>row 3, and also row 2 with row 4, we see that using</p>
      </sec>
      <sec id="sec-5-4">
        <title>Transformer-based embeddings yields an improvement</title>
        <p>of 2.5% in the goals  1 on English, and a smaller
improvement on French, Spanish, Chinese, and Portuguese, but
is worse on German and Swedish5. FastText initial
features, on the other hand, perform as well or better than
Transformer-based features in the line-level task. It is
important to consider, though, that the improved error
rate of the Transformer-based model comes at a higher
computational cost during inference. This consideration
is especially important when the model is deployed in
a high-load commercial application where latency is a
crucial factor.</p>
      </sec>
      <sec id="sec-5-5">
        <title>A second important observation is that the multi-task models generally outperform their single-task counterparts for the token sequence labeling task. Additionally, the multi-task model has a significant advantage in a</title>
        <p>
          Section-specific Models
The simplification of model development and
maintenance is even more significant when we contrast the
unified multi-task model described above with the
typical two-stage approach for résumé parsing [
          <xref ref-type="bibr" rid="ref15 ref3">1, 2, 3</xref>
          ]. The
latter requires training and maintaining several models:
one for the initial line segmentation task, and then one
using a three-way split involving training, validation, and test sets.
        </p>
      </sec>
      <sec id="sec-5-6">
        <title>4Row 2 of both sub-tables evaluates the same underlying model (but</title>
        <p>for diferent tasks), and similarly for row 4</p>
      </sec>
      <sec id="sec-5-7">
        <title>5Swedish is an outlier, where the Transformer-based models are</title>
        <p>markedly less accurate. This might be due to the small size of
Swedish data used for pre-training mT5.
depicted in Figure 2.</p>
        <p>CRF receives as input the concatenation of: (i) the repre- 3Due to the relatively small size of the corpora, we opted against
for entity extraction within each specific section type stage, e.g. the error observed for the Single-task models
(e.g. one single-task model for the entities related to con- presented in Table 2b. The aim is to provide the
practact information, another single-task model for entities titioner with a quantifiable assessment of the trade-of
related to work experience, etc). By contrast, the unified between engineering simplicity and task accuracy.
multi-task model we proposed is used to label all the
entities across the whole résumé at once, regardless of Analysis and Details on Deployment
the section type. This simplification, however, comes at a
cost of increased error rate since a section-specific model The results already suggest that the Transformer-based
has to decide among a much smaller set of labels, and initial features perform generally better for the
tokenreceives a shorter text sequence as input. level sequence labeling task. Furthermore, they do not</p>
        <p>In this part, we attempt to quantify such degradation. need language-specific handcrafted features, so they can
We train section-specific models, i.e. individual models, be readily applied to new languages. On the other hand,
for the entities for three of the section types: contact infor- the alternative set of initial features (the combination
mation, work experience, and education. Each is trained of word embeddings and handcrafted features) performs
and evaluated only on the corresponding segment of better in the line sequence labeling task for detecting
the résumés. Segmentation is performed using the gold section and group labels.
standard annotations for sections, in order to focus our However, in terms of eficiency, our experiments
remeasurements on the token-level task. In Table 3, we re- veal that using word embedding initial features leads
port the micro-average  1 scores grouped by the relevant to a considerable improvement in time-eficiency
dursections, comparing the performance of each section- ing inference, when compared to the Transformer-based
specific model to the proposed unified, multi-task model. features. The inference time for the multi-task model
Results are reported for English, French, and Chinese. was measured under both feature sets. On a bare-metal</p>
        <p>
          We show a loss in  1 ranging from 1% to 5% depending server with a single GPU6, we observed a speedup of 7
on the section and language. Since the section-specific of the FastText models compared to Transformer-based
models benefit from the gold standard segmentation of features. Furthermore, when utilizing CPU-only
hardsections, the results should be considered as an upper ware7, the speedup increased substantially to 90. As an
bound of the degradation in error rate. A real-world
system implemented according to the two-stage approach
should expect a compound error carried from the first
6NVIDIA Tesla T4 and Intel® Xeon® Platinum 8259CL CPU @
2.50GHz.
7Intel® Xeon® CPU E5-2630 v2 @ 2.60GHz
example, we note that the multi-task model using Fast- Both result in a significant degradation of performance
Text initial features, deployed on CPU-only servers via with respect to the models including the BiRNN, again
TensorFlow Serving [
          <xref ref-type="bibr" rid="ref18">19</xref>
          ], yields a latency of 450 ms per showing the importance of the BiRNN for this task.
résumé without batch processing. The third group involves variants that also apply
Transformers to each individual line, but this time we allow
Ablations and Comparison with Previous Work for the Transformer encoder to be fine-tuned with the
task supervision. In this case, we do not employ a BiRNN
Table 4 presents an ablation study of the proposed archi- for contextualizing token representations across lines
tectures in order to empirically support our architectural because this would require a much more challenging
opdesign choices. Furthermore, some of the ablated vari- timization procedure8 and thus each line is processed
ants are re-implementations of systems proposed in pre- independently. Variant ⑧ involves a BERT encoder
(bevious work and thus act as baselines for the experiments ing fine-tuned) that computes representations for each
presented above in this section. token in the line, and uses a CRF layer to predict their
la
        </p>
        <p>
          The first group involves variants that use, as initial fea- bels. When compared to our proposed model (variant ④),
tures, the combination of FastText word embeddings and we observe a significant drop in performance, suggesting
handcrafted features. Variant ① is the multi-task model that the contextualization across diferent lines in the
presented in Table 2a. The first ablation, variant ②, in- résumé is the critical factor for the performance of the
volves replacing the top-wise CRF layer with a Softmax system. Interestingly, when variant ⑧ is compared to
layer. Both variants have comparable performance, with variant ⑦ —identical, except for fine-tuning— we do see
a small degradation when Softmax is used. The next abla- an improvement in performance, suggesting that without
tion, variant ③, removes the BiRNN layer and thus makes inter-line contextualization, fine-tuning is indeed helpful.
the CRF predict the token labels using the initial features Variant ⑨ is similar to the previous variant but
redirectly. This is a re-implementation of the system pro- places the CRF layer with Softmax. This model is a
posed by Tosik et al. (2015) [
          <xref ref-type="bibr" rid="ref15">1</xref>
          ], although they did not re-implementation of the NER system presented by
Deshare their handcrafted features (and therefore we use vlin et al. (2019) [7] and it is also equivalent to the system
those described in Appendix B). This other ablated vari- used for résumé parsing by Barducci et al. (2022) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
ant has a substantial degradation in performance with We can observe that the performance is similar to the
respect to our proposed model, suggesting that the role previous one (although CRF seems to achieve slightly
played by the BiRNN layer is critical. better results). Lastly, variant ⑩ is included as a
con
        </p>
        <p>The second group involves variants that apply frozen trol, in which we fine-tune the encoder component of T5
Transformers to each line individually, and then concate- and predict labels using a Softmax on top of the
Transnate every line to obtain the initial features (this is visu- former representations for each token. Again, T5
proally described in Figure 2). Variant ④ is the multi-task vides slightly better results with respect to the equivalent
model presented in Table 2a. The first ablation, vari- BERT variant.
ant ⑤, involves replacing the T5 (or mT5) encoder with In summary, the ablation experiments suggest that the
a BERT (or mBERT) encoder [7]. We observe an appre- BiRNN layer, which contextualizes the token
representaciable degradation in performance, suggesting that the tions across the entire résumé, has a significant impact
pre-trained T5 family of models produces representations
that are more useful for our task. Variant ⑥ and ⑦ use
T5 and BERT, respectively, but omit the recurrent layer.</p>
      </sec>
      <sec id="sec-5-8">
        <title>8A naïve implementation for this procedure would require keeping</title>
        <p>in memory as many copies of the Transformer as lines in the target
résumé.</p>
        <p>
          Table 4 trade-of analysis of the proposed variants and described
Ablation study. Variants are compared in terms of the micro- challenges for deployment in production environments.
average  1 obtained for the token sequence labeling task. Vari- The ablation experiments suggest that the BiRNN layer
ants ① and ④ represent the models discussed in the previous contextualizing across the résumé is critical for
perforpart of this section. Other model variants depart from either mance, and that the CRF component further provides a
one of these by changing one aspect at a time. In particular, smaller improvement.
vaanrdiavnatr③ianrte-⑨imispleeqmueivnatsletnhtetsoytshteemarocfhTitoescitkueret aplr.o(p20o1s5e)d[b1y], Potential directions for future research include the
folDevlin et al. (2019) [7] for other sequence labeling tasks. IF lowing: using character-based initial features [5, 6] for
denotes initial features. Each result is an average of three in- the FastText variants, as they can complement word
emdependent replications. beddings by incorporating information from the surface
form of the text and may even ofer the opportunity to
Model variant English French Chinese gradually replace handcrafted features; domain-adapting
F①asIFtT+eBxitRinNiNtia+lCfeRaFtures 89.03 86.90 92.66 the Transformer representations with unannotated
ré② IF+BiRNN+Softmax 88.86 86.53 92.67 sumés, considering the reported efectiveness of this
tech③ IF+CRF [
          <xref ref-type="bibr" rid="ref15">1</xref>
          ] 65.89 64.68 67.53 nique in enhancing downstream task performance [20];
and building multilingual models to improve sample
efiTransformer initial features ciency for low-resource languages. Furthermore,
alter(④froTz5e+nB)iRNN+CRF 90.94 88.65 92.61 native Transformer architectures designed specifically
⑤ BERT+BiRNN+CRF 88.91 86.34 91.79 for long input sequences [21, 22] could be used in order
⑥ T5+CRF 78.65 75.40 76.91 to encode the entire résumé in a single pass, while also
⑦ BERT+CRF 74.70 73.53 81.51 enabling the possibility to fine-tune the encoder.
Transformer initial features, linewise
(fine-tuned)
⑧ BERT+CRF 83.55
⑨ BERT+Softmax [
          <xref ref-type="bibr" rid="ref3">7, 3</xref>
          ] 83.13
⑩ T5+Softmax 84.18
on the performance. The CRF helps to further improve
the performance but in a smaller amount. The variants
that allow for fine-tuning the Transformer component
outperform their frozen-Transformer equivalents, but
they are in turn outperformed by our proposed solutions
(variants ① and ④).
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>7. Conclusion</title>
    </sec>
    <sec id="sec-7">
      <title>Limitations</title>
      <p>As discussed in Section 4, despite our best eforts to cover
as many locations, industries, and seniority levels, it is
not feasible for résumé parsing corpora with sizes of up
to 1200 résumés to actually contain samples from every
subgroup of the population under study. Therefore, we
would like to highlight that the findings presented in this
work apply specifically to résumés that are similar to
those included in the corpora, and may not generalize
with the same level of accuracy to other résumés
belonging to combinations of location, industry, and work
experience that were not seen by the model during
training.</p>
      <p>Résumé parsing is an important task for digitalized re- Ethics Statement
cruitment processes, and the accuracy of the parsing step
afects downstream recommender systems significantly. The system described in this work is intended for
pars</p>
      <p>In this work, we study résumé parsing extensively ing résumés of individuals from diferent backgrounds,
in seven languages. We formulated it as a sequence located around the globe. Considering the importance
labeling problem in two levels (lines and tokens), and of inclusivity in this context, we made a great efort to
studied several variants of a unified model that solves cover the diversity of the use of language in our
corboth tasks. We also described the process for develop- pora with the objective in mind. This helps us to provide
ing high-quality annotated corpora in seven languages. high-quality résumé parsing for individuals from various
We showed through experimental results that the pro- industries and locations.
posed models can perform this task efectively despite Furthermore, the data used for training and evaluating
the challenges of substantially long input text sequences our models consist of résumés that contain sensitive
inforand a large number of labels. We observed that the joint mation from real-world individuals. We have taken the
model is more convenient than the typical two-stage so- necessary privacy and security measures for protecting
lution in terms of resource eficiency and model life-cycle this information throughout every step of this project.
maintainability, and also found that in some cases the
joint model yields better performance. We provided a
A. Details on the Corpus Besides the categories listed above, lines that belong to
sections Work Experience, Education, and Internship
Annotations can belong to experience groups. Internally, the tool uses
the IOB (inside, outside, beginning) format for the groups
Annotators were provided access to a custom annotation within each section type. Note that, for example in the
tool that we developed for this task. In Figure 3 we show case of the Education section, the label I-edu denotes a
the user interface of this tool, with an example résumé line that is part of the Education section but it’s not part
annotated for sections, groups, and entities. of any particular group, whereas the label B-edu_group</p>
      <p>The annotators were asked to detect and highlight spe- denotes a line that lies at the beginning of a group in the
cific information in résumés written in their language. Education section.</p>
      <p>A.2. Labels for Tokens
The sequence labeling task at the token-level is intended
to extract entities. Entities are annotated on arbitrary
spans of characters in the résumé text, in order to
generalize the annotations for any possible tokenization. We
allow for a total of 17 labels. Although most entities are
usually found in specific sections, we allow for
annotating any entity in any part of the résumé.</p>
      <p>A.2.1. Contact Information entities</p>
      <sec id="sec-7-1">
        <title>The following entities are usually found in the contact</title>
        <p>information section.</p>
        <p>Name Candidate’s name.</p>
        <p>Phone number Candidate’s phone number.</p>
        <p>Email Candidate’s email address.</p>
        <p>St. address Unit information (apartment, floor) and
neigh</p>
        <p>borhood.</p>
        <p>ZipCode An alphanumeric postal code.</p>
        <p>City The city the candidate lives in.</p>
        <p>State First-level geopolitical subdivision where the candidate</p>
        <p>is located.</p>
        <p>A.2.2. Work Experience and Internship entities</p>
      </sec>
      <sec id="sec-7-2">
        <title>These entities are usually found in the groups of either</title>
        <p>the work experience or the internship sections.
Company Name of a candidate’s employing organization.
Job title Title that the employer gave to the candidate while
working for the company.</p>
        <p>Period Period in which the candidate held the position.
A.2.3. Education entities</p>
      </sec>
      <sec id="sec-7-3">
        <title>These entities are usually found in the groups of the</title>
        <p>education section.</p>
        <p>School name Name of an institution where the candidate
was formally educated.</p>
        <p>Degree title he name of the academic program in which a
student participates or was awarded a degree.</p>
        <p>Degree Period Period of time in which the candidate
attended the educational organization in fulfillment of
the degree.</p>
        <p>Major The core academic discipline the candidate had to
focus on while pursuing their degree.</p>
        <p>GPA Grade, or any other candidate score, presented as a
standardized measurement.</p>
        <p>A.2.4. Language entities</p>
      </sec>
      <sec id="sec-7-4">
        <title>The following entities are usually found in the Languages section.</title>
        <p>Language name The name of a language that the candidate
claims to be familiar with.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>B. Detail on the Handcrafted</title>
    </sec>
    <sec id="sec-9">
      <title>Features</title>
      <p>The non-Transformer models presented in Section 5
employ a combination of FastText word embeddings and
handcrafted features. The complete list and details of
each handcrafted feature is provided in GitHub because
of space considerations9. The final set of features was
determined based on the empirical results of preliminary
experiments, which are not included in this study due to
space constraints. Note that certain features necessitate
dictionaries of relevant terms to be computed, thereby
requiring separate dictionaries for each language.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>tems</surname>
          </string-name>
          , volume
          <volume>30</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2017</year>
          .
          <article-title>any span of characters in the text) include entities</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          URL: https://proceedings.neurips.cc/paper/2017/
          <article-title>The description for every type of annotation label is</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>file/</surname>
            3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. included below. [18]
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Abadi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Barham</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Davis,
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Devin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghemawat</surname>
          </string-name>
          , G. Irving,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Is- A.1</article-title>
          . Labels for Lines
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>lfow: A system for large-scale machine learning, The label set includes a total of 18 diferent labels</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>in: 12th USENIX Symposium on Operating Sys-</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>tems Design and Implementation (OSDI 16)</source>
          ,
          <year>2016</year>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>1</year>
          .1. Sections
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          pp.
          <fpage>265</fpage>
          -
          <lpage>283</lpage>
          . URL: https://www.usenix.org/system/
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          files/conference/osdi16/osdi16-abadi.pdf.
          <article-title>We allow the annotators to label lines into the following</article-title>
          [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Olston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Fiedel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gorovoy</surname>
          </string-name>
          , J. Harmsen, sections:
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>ml serving</source>
          ,
          <year>2017</year>
          . URL: https://arxiv.org/abs/1712.
          <article-title>Work Experience Information about the candidate's em-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          06139. doi:
          <volume>10</volume>
          .48550/ARXIV.1712.06139. ployment experience. [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gururangan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Marasović</surname>
          </string-name>
          , S. Swayamdipta,
          <article-title>EducationtionI.nformation about the candidate's formal educa-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>mains and tasks, in: Proceedings of the 58th An- Skills Information about the candidate's work-related abili-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>tics</surname>
          </string-name>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>8342</fpage>
          -
          <lpage>8360</lpage>
          . URL: https:// Summary Brief statement
          <article-title>intended to display a candidate's</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>740</volume>
          . doi:
          <volume>10</volume>
          .18653/ most compelling abilities and attributes.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          v1/
          <year>2020</year>
          .acl-main.
          <volume>740</volume>
          .
          <string-name>
            <surname>Objective</surname>
          </string-name>
          <article-title>Candidate's work-related goals</article-title>
          , usually written [21]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          , J. Carbonell, Q. Le, ifnorpfrroosme.aItjoebmoprhcaosmizepsanwyh.at the person is looking
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <article-title>Proceedings of the 57th Annual Meeting of the As- cations, Licenses and Certifications, etc</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <year>2019</year>
          , pp.
          <fpage>2978</fpage>
          -
          <lpage>2988</lpage>
          . URL: https://aclanthology.org/
          <article-title>Letter Letters embedded in the résumé (frequently cover</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <fpage>P19</fpage>
          -
          <lpage>1285</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          -1285. letters and reference letters). [22]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hutchins</surname>
          </string-name>
          , I. Schlag,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Dyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Neyshabur</surname>
          </string-name>
          ,
          <article-title>For sections, we use IO (inside, outside) format for the</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>Neural Information Processing Systems</source>
          , volume
          <volume>35</volume>
          ,
          <article-title>two consecutive lines of the same type of section belong</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Curran</given-names>
            <surname>Associates</surname>
          </string-name>
          , Inc.,
          <year>2022</year>
          , pp.
          <fpage>33248</fpage>
          -
          <lpage>33261</lpage>
          .
          <article-title>URL: to the same section element</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>