<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing Job-Skill Matching with LLM-Driven Data Augmentation and Fine-Tuned Bi-Encoders</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohab Ali</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer and Systems Engineering Department, Faculty of Engineering, Alexandria University</institution>
          ,
          <addr-line>Alexandria</addr-line>
          ,
          <country country="EG">Egypt</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>Recent advances in large language models (LLMs) have enabled a wide range of applications across diferent domains, including human resources (HR). A key challenge in this space is identifying relevant skills for a given job title. This paper details our approach developed for the TalentCLEF 2025 shared task on job-skill matching. We explore various AI-based strategies, with a significant focus on leveraging LLM-driven data augmentation to enhance model training. Our methodology evaluates cross-encoder models and sentence similarity architectures utilizing diverse pretrained models, analyzing their trade-ofs in mean Average Precision (mAP) and computational eficiency. Our final system achieved a mAP of 0.345 on the test set. We also discuss insights from contrasting experiments, including those less successful, to highlight modeling limitations and ofer recommendations for future improvements in this domain. This work contributes to the TalentCLEF 2025 shared task's aim to eficiently align occupational titles with relevant professional skills.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Job-Skill Matching</kwd>
        <kwd>Large Language Models</kwd>
        <kwd>Data Augmentation</kwd>
        <kwd>Sentence Embeddings</kwd>
        <kwd>Bi-encoder</kwd>
        <kwd>TalentCLEF</kwd>
        <kwd>Human Resources</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>1.1. Motivation</title>
        <p>Job-skill matching is a fundamental task in modern human resources (HR) and talent management.
With job postings often attracting thousands of applicants, there is a growing demand for automated
systems that can eficiently and accurately match candidates to job requirements. A key component
of this automation is identifying and ranking the most relevant skills for a given job title, a challenge
actively being addressed by research and shared tasks in the field. Recent advances in natural language
processing (NLP), particularly large language models (LLMs), have opened new avenues for building
intelligent HR systems. These technologies have the potential to power applications such as skill-based
candidate ranking, personalized up-skilling recommendations, and early detection of emerging skill
gaps. However, deploying such systems in real-world settings involves critical challenges related to
scalability, semantic ambiguity, and fairness, which this work aims to explore.</p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2. Task Definition</title>
        <p>
          The Job Title–Based Skill Prediction task, a part of the TalentCLEF 2025 lab within the Conference and
Labs of the Evaluation Forum (CLEF) [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], requires participants to rank a predefined set of candidate
skills for a given job title. The goal is to assign a relevance score (ranging from 0, not relevant, to 1,
highly relevant) to each skill, ensuring that more appropriate skills are placed higher in the resulting
ranked list.
        </p>
        <p>
          This task utilizes a dataset derived from the ESCO (European Skills, Competences, Qualifications,
and Occupations) database [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. The training data consists of three main components:
• Job–skill relations file: Maps job IDs to skill IDs, providing a label (essential/optional) for each
job-skill pair.
• Job titles file: Provides multiple names and aliases for each job ID.
        </p>
        <p>• Skill names file: Provides multiple names and aliases for each skill ID.</p>
        <p>Each job and skill ID corresponds to a unique URL in the ESCO taxonomy. For validation and testing,
the data is structured as follows:
• A skills corpus file: Lists all candidate skill entries (with their names and aliases).
• A job queries file: Provides a list of job titles for which skills need to be ranked. These job titles
were specifically selected by domain experts, which may result in a data distribution diferent
from the training set.
• A ground truth file: Contains the relevance labels, represented as pairs of job and skill indices
that reference their respective positions in the queries and corpus files.</p>
        <p>This formulation simulates a real-world retrieval scenario where a system must match a job title query
against a predefined skills corpus to identify the most relevant skills.</p>
      </sec>
      <sec id="sec-1-3">
        <title>1.3. Challenges</title>
        <p>This task presents several key challenges. First, scalability is a major concern: matching hundreds of
job titles with hundreds of candidate skills quickly becomes computationally expensive, particularly
when using cross-encoder models that require evaluating every possible pair, resulting in quadratic
complexity. Second, accessibility must be considered. High-performing solutions should ideally run on
afordable hardware to ensure practical adoption in real-world HR settings. Third, semantic ambiguity
and lack of context make it dificult for models to make accurate predictions based only on short job
titles or skill names. This issue afects both humans and machines. Finally, while LLM augmentation can
efectively provide additional context, its deployment is often constrained by significant computational
costs due to the need for high-end GPUs or reliance on expensive APIs. Designing a solution that
balances contextual reasoning and computational eficiency is therefore non-trivial.</p>
      </sec>
      <sec id="sec-1-4">
        <title>1.4. Contribution</title>
        <p>In this study, we investigate both cross-encoder and bi-encoder (sentence similarity) approaches for the
job–skill matching task, evaluating their trade-ofs in performance and eficiency. To address the lack
of input context, we experiment with augmenting job descriptions using lightweight LLMs that do not
require high-end hardware. Our experiments include successful and unsuccessful attempts, with all
results documented to inform future work in this area.1</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <sec id="sec-2-1">
        <title>2.1. Encoder models</title>
        <p>Encoder models are central to job–skill matching tasks, as they transform textual input, such as job
titles and skills, into dense vector representations that capture semantic similarity. In our experiments,
we evaluate several pretrained encoder models, each with diferent domain characteristics and training
strategies.</p>
        <p>
          Zhang et al. fine-tuned XLM-RoBERTa [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], a multilingual transformer model pretrained on 2.5 TB
of data, using the same ESCO database [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] provided for this shared task. Their goal was to adapt the
model to the HR domain by performing domain-adaptive pretraining on the ESCO taxonomy, which
spans 27 languages. The resulting model, ESCOXLM-R, was trained using masked language modeling
1The code for this study is available at https://github.com/mawhab/Job-Skill-Matching.
and additional multilingual objectives tailored to the ESCO structure. We selected this model for its
domain specificity and strong alignment with the dataset used in the task.
        </p>
        <p>
          Wang et al. introduced the E5 model family [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], a set of general-purpose sentence encoders with four
variants: small, base, large, and large-instruct. The models were pretrained using contrastive learning
on one billion multilingual text pairs, then fine-tuned on supervised tasks. The "instruct" variant
was further tuned using instruction-following objectives and achieved strong results across multiple
benchmarks. We include these models to explore trade-ofs between performance and eficiency.
        </p>
        <p>
          Reimers and Gurevych developed JobBERT [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], a domain-specific sentence encoder built on top of
all-mpnet-base-v2 [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It was fine-tuned on a large corpus of job titles and their associated skills and
requirements, with the goal of enabling semantic matching in HR and recruitment settings. The model
maps job-related text into 1024-dimensional embeddings and is optimized for job title similarity and
job–skill alignment tasks. We experimented with this model for its extensive pretraining on jobs and
job descriptions.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Large Language Models</title>
        <p>
          LLMs can complement encoder-based systems by generating additional context or augmenting inputs
with richer descriptions. While many LLMs are resource-intensive, recent open-weight models have
improved accessibility. In our experiments, we use LLaMA 3.1 8B [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], a compact but capable model
that fits on a single GPU. We use it to generate descriptions and contextual data that improve the
performance of embedding-based methods in similarity scoring.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Initial Approaches</title>
      <sec id="sec-3-1">
        <title>3.1. Architectural Decisions</title>
        <sec id="sec-3-1-1">
          <title>3.1.1. Cross-Encoder</title>
          <p>The first approach centered on a cross-encoder architecture. In this setup, a job title and a candidate
skill were concatenated, typically separated by a special [SEP] token, and fed as a single sequence into a
pretrained transformer encoder. For extracting a final representation for classification, we experimented
with two common strategies: using the final hidden state of the [CLS] token, and applying mean pooling
over the last hidden layer’s outputs (respecting attention masks).</p>
          <p>This pooled representation was then passed through a linear projection layer followed by a sigmoid
activation to predict a relevance score between 0 and 1. The core idea was that the cross-attention
mechanisms within the encoder would allow for deep contextual interaction between the job and skill,
enabling the model to learn diferent nuanced relevance signals. The intended advantage for this design
was better modeling of nuanced job-skill interactions. However, while conceptually powerful, this
approach sufered critically from computational ineficiency. Evaluating each job-skill pair
independently resulted in quadratic inference complexity (( · ) where J is the number of jobs and S is the
number of skills). With even a moderately sized test set, this led to prohibitive processing times. Despite
showing some promise in capturing relevance, achieving an initial mean Average Precision (mAP)
of approximately (0.23-0.26) after 1-5 epochs of fine-tuning with cross-entropy loss across diferent
encoders (ESCOXLM-R, E5-Large, and E5-Instruct), its lack of scalability made it impractical as a
standalone solution for the full task.</p>
          <p>Our key takeaway from this experiment was that while the cross-encoder approach efectively
captures relevance through deep job-skill interactions through attention mechanisms, its significant
computational expense at both training and inference makes it more suitable for re-ranking a smaller
set of candidates rather than for initial large-scale retrieval. Due to time and resource limitations, we
decided not to move forward with this approach.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.1.2. Bi-Encoder</title>
          <p>The second exploratory direction utilized a bi-encoder architecture, focusing on sentence similarity.
Here, job titles and skills were encoded into dense vector embeddings independently using a shared
pretrained encoder. The relevance score for a job-skill pair was then computed using the cosine similarity
between their respective embeddings.</p>
          <p>The intended advantage for this approach was its high computational eficiency, as job and skill
embeddings could be pre-computed and stored, allowing for fast nearest-neighbor searches or similarity
calculations. Although significantly more eficient, this approach often struggled with the lack of
explicit interaction context. When relying solely on the often concise raw job titles and skill names,
the encoders found it challenging to disambiguate meanings or infer relevance accurately without the
direct comparison ofered by cross-attention. This was reflected in its performance, resulting in a lower
mAP of approximately (0.13-0.22) across diferent encoders (ESCOXLM-R, E5-Large, E5-Instruct). It is
important to note that these bi-encoder results reflect zero-shot performance using the base pretrained
encoders without any task-specific fine-tuning at this exploratory stage.</p>
          <p>Our key takeaways from this experiment are that bi-encoders are eficient but require either highly
efective base encoders adept at capturing nuanced semantics from short texts or, more critically, richer
input representations that provide more context for the individual job and skill texts.</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>3.1.3. Summary of model architectures</title>
          <p>These initial explorations highlighted a key trade-of: the deeper contextual understanding of
crossencoders came at a steep computational price, while the eficiency of bi-encoders was often undermined
by a deficit in contextual awareness when using their base pretrained representations and basic textual
inputs. This understanding directly motivated our subsequent eforts to develop a more balanced
and efective methodology, focusing on enhancing the input to bi-encoder systems, as detailed in the
following section.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Data Formatting and Augmentation</title>
        <p>To address the context limitations of the bi-encoder design and improve its performance, we
experimented with several approaches to enrich job and skill representations, as detailed below.</p>
        <sec id="sec-3-2-1">
          <title>3.2.1. ESCO Descriptions</title>
          <p>ESCO’s taxonomy database contains not only information about jobs, skills, and their relations, but also
detailed descriptions for each job and skill entry. These descriptions represent an excellent opportunity
for contextual enrichment, being reliable, accurate, and directly relevant to the entities the encoder
processes.</p>
          <p>However, this approach quickly encounters a challenge with the validation and test sets. While the
skills in these sets are from ESCO (and thus have descriptions), the job titles are not directly from ESCO.
Instead, they are selected by domain experts and lack a direct mapping or URL to an equivalent ESCO
occupation (even if such an equivalent might conceptually exist).</p>
          <p>To overcome this, our strategy involved training models with the available ESCO descriptions and
then, for validation and testing, attempting to find the closest matching job title from the training set
(which has ESCO descriptions) and using its description as a proxy.</p>
          <p>We used two distinct methods to find these close matches. First, we tried Python’s get_close_matches
function from the Diflib library. This function performs sequence-based text matching, identifying the
closest n matches for a given string (a validation/test job title) from a list of strings (training set job
titles) based on a similarity cutof score. This approach, while efective, was relatively slow but could be
performed ofline and prepared in advance.</p>
          <p>The second method utilized JobBERT. This was a conceptually strong approach, as JobBERT is
specifically pretrained for job title similarity. We first generated embeddings for all job titles from
both the training set and the validation/test sets using JobBERT, then for each validation/test job title,
we identified the most similar training set job title via cosine similarity ranking. This method also
benefited from being significantly faster due to eficient batch processing and parallelization capabilities
for embedding generation.</p>
          <p>We evaluated this proxy-description augmentation technique with various encoders. A key
observation across both Diflib and JobBERT-based matching methods was that the model had the highest
score when using a very high similarity cutof (around 0.97). This strict cutof meant that approximately
95% of the validation set job titles did not receive a proxy description from the training set, as no
suficiently close match could be identified. However, using descriptions only for the matched 5%
still outperformed using no descriptions at all or using a lower cutof (which would introduce more,
potentially less accurate, descriptions). This finding strongly signaled that while accurate descriptions
were crucial, having no description was preferable to having an inaccurate one. This insight provided a
strong motivation for exploring LLM-based description generation, as discussed next.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2.2. LLM Generated Descriptions</title>
          <p>To generate informative descriptions, particularly for potentially ambiguous job titles, and ensure
system accessibility, we needed a high-performing LLM capable of running on a single consumer-grade
GPU. We opted to use LLaMA 3.1 8B instruct. We engineered the prompt with the following objectives:
conciseness (to fit encoder context windows), clarity (to avoid superfluous tokens or hallucinations),
inclusion of 2-3 relevant skills (to directly enhance job-skill similarity signals), direct output of only
the description (omitting conversational preambles), and avoidance of redundant information already
present in the job title. We worked with the following prompt:</p>
          <p>You are an expert HR assistant specialized in writing job descriptions. Generate a highly
concise, professional, and factual job description based ONLY on the provided job title.
Focus on typical key responsibilities. Crucially, explicitly list 2-3 common, essential
skills associated with the role within the description, often towards the end. Do
not add information not directly implied by the job title. The description must be brief,
ideally 2-3 sentences maximum, to fit within processing limits. Output only the description
text, without any preamble or introductory phrases.</p>
          <p>We generated descriptions for all job titles in the training, validation, and test sets. For the training
set, this was done despite the availability of oficial ESCO descriptions to ensure consistency in the style
and focus of descriptions across all data splits; LLM-generated descriptions might emphasize diferent
aspects than the ESCO originals, and training on a consistent distribution was deemed important.
Furthermore, we generated a unique description for each job title alias, even those mapping to the same
underlying ESCO job ID. This was to prevent the encoder from potentially learning to disregard the
specific alias text and overly rely on identical descriptions if they were shared.</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>3.2.3. Summary of Input Formatting</title>
          <p>Our exploration into enriching context, initially with oficial ESCO descriptions and then more efectively
with LLM-generated content, aimed to provide richer inputs for the encoders. We used LLMs to generate
descriptions for jobs for training, validation, and test sets. For skills, we used the available ESCO
descriptions as they were available across training, validation, and test sets. The final input text to the
encoders would be:
• For jobs: "Job: (job title) [SEP] Description: (LLM generated description)"
• For skills: "Skill: (skill name) [SEP] Description: (oficial ESCO description)"</p>
          <p>For E5-Instruct, we include an instruction prompt to make use of the model’s instruction tuning. The
input was adjusted to the following:
• For jobs: "Instruct: Given a job title and its description, retrieve relevant skills based on their
descriptions. Query: Job: (job title) [SEP] Description: (LLM generated description)"
• For skills: "Skill: (skill name) [SEP] Description: (oficial ESCO description)"</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <sec id="sec-4-1">
        <title>4.1. Overall System Architecture</title>
        <p>Our final system uses a bi-encoder architecture. Pretrained transformer models serve as the backbone
to independently encode job and skill textual representations. These encoders are fine-tuned using
a contrastive learning approach. At inference, the relevance of a skill to a job is determined by the
cosine similarity between their respective vector embeddings, allowing for eficient ranking. This
architecture was selected to balance predictive performance with the computational eficiency required
for potentially large-scale matching, informed by our initial explorations in Section 3.1.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Input Representation</title>
        <p>The textual inputs for jobs and skills were combined with their descriptions to provide rich contextual
information to the encoders, as detailed in Section 3.2.3.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Encoder Models and Embedding Extraction</title>
        <p>We evaluated three distinct pretrained transformer models as the backbone for our bi-encoder:
ESCOXLM-R, E5-Large, and E5-Instruct. These models were fully fine-tuned. For each input text,
the dense vector embedding was derived by mean pooling over the last hidden layer’s token
embeddings, respecting attention masks.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Fine-tuning setup</title>
        <p>
          The bi-encoder models were fine-tuned using Multiple Negatives Ranking Loss (MNRL) following
Henderson et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. MNRL is a contrastive loss function designed to train retrieval models by
ensuring that positive (relevant) pairs have higher similarity scores than negative (irrelevant) pairs.
For constructing training batches, we processed our list of training jobs. For each job, one of its aliases
was randomly sampled, along with one skill associated with it and a randomly sampled alias for that
skill. This formed a positive (anchor, positive) pair. This sampling strategy ensured that each unique
job ID appeared at most once per batch, which is crucial for the correctness of using in-batch negatives
and aids training eficiency on the large dataset. All other skill texts within the same batch served as
in-batch negatives for each anchor job text.
        </p>
        <p>
          A comprehensive list of hyperparameters can be found in Table 1. We used the AdamW optimizer
[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] with a linear learning rate decay schedule and a 6% warmup period. The final model weights were
selected from the epoch that achieved the highest mAP on the validation set. All training was conducted
on a P100 GPU via the Kaggle platform.
        </p>
      </sec>
      <sec id="sec-4-5">
        <title>4.5. Inference</title>
        <p>During inference, all job texts and skill texts along with their generated descriptions are first encoded
using the fine-tuned bi-encoder to generate their respective dense vector embeddings. For a given query
job, its embedding is compared against all skill embeddings in the target corpus using cosine similarity.
Skills are then ranked for that job in descending order of their similarity scores.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Discussion</title>
      <p>Our best model achieved a mAP score of 0.345 on the oficial TalentCLEF test set. In this section, we
present diferent encoder models’ results on the validation set.</p>
      <sec id="sec-5-1">
        <title>5.1. Results</title>
        <p>Our experiments reveal several key insights into building efective job-skill matching systems. The
ifnal fine-tuned bi-encoder architecture, particularly when using E5-Instruct with LLM-generated
job descriptions, demonstrated strong performance, achieving 0.313 mAP on our validation set and
generalizing well to the oficial TalentCLEF test set with a mAP of 0.345.</p>
        <sec id="sec-5-1-1">
          <title>5.2.1. Interpreting Encoder Performance</title>
          <p>The E5 model family, especially E5-Instruct, consistently outperformed ESCOXLM-R in our
finetuned setups (Table 2). The E5 models’ robust performance, even with minimal context (as seen in
the "No Descriptions" fine-tuned column), suggests their large-scale diverse pretraining provides a
strong foundation. The instruction tuning of E5-Instruct likely further aided its ability to focus on the
retrieval task when presented with augmented inputs. The comparatively lower scores for ESCOXLM-R,
despite its ESCO-specific pretraining, were notable. The largest performance gap between
ESCOproxied descriptions and LLM-generated descriptions was observed for ESCOXLM-R (an increase
from 0.2689 to 0.298 mAP). This could imply that ESCOXLM-R is highly sensitive to the style and
ESCOXLM-R</p>
          <p>E5-Large</p>
          <p>E5-Instruct
5.2. Discussion
completeness of contextual information; when most validation job titles lacked a high-quality ESCO
proxy description, its performance sufered more acutely than the E5 models. Conversely, it benefited
more when comprehensive LLM descriptions were provided for all titles. This suggests that while
domain pretraining is valuable, adaptability to varied descriptive styles, or a consistent supply of rich
context, is equally crucial.</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>5.2.2. Impact of Augmentation and Fine-tuning</title>
          <p>An important finding is the significant improvement from zero-shot performance (e.g., 0.2197 for
E5-Instruct) to fine-tuned performance with LLM descriptions (0.313 for E5-Instruct). This suggests that
while pretrained encoders have general semantic understanding, task-specific fine-tuning is essential
for optimal retrieval. The LLM-generated descriptions, which explicitly included 2-3 relevant skills
as per our prompt, likely created stronger semantic anchors for the bi-encoder, directly facilitating
better job-skill similarity assessments. The guarantee of having a description for every job title via
LLM generation, regardless of its presence in ESCO, was a key advantage over the proxy-based ESCO
description approach. Our final fine-tuned bi-encoder system also surpassed the performance of our
initial cross-encoder explorations (0.23-0.26 mAP), demonstrating a more scalable and ultimately more
efective solution.</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>5.2.3. Limitations</title>
          <p>Despite these results, limitations exist. Our choice of LLaMA 3.1 8B for generating job descriptions
balanced performance with accessibility; larger LLMs might ofer further improvements at a higher
computational cost. The prompt guiding this LLaMA 3.1 8B generation process, though iterated upon,
could be further refined. Similarly, the instruction prompt prepended to the input for E5-Instruct, while
simple and efective, was not exhaustively optimized and could be a subject for further experimentation.
Additionally, our hyperparameter search for fine-tuning, while guided by common practices, was not
exhaustive due to resource constraints.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>In this work, we explored various approaches for the TalentCLEF job-skill matching task, systematically
evaluating diferent system architectures, encoder models, input augmentation strategies, and
finetuning techniques. Our findings highlight the significant impact of rich contextual input (particularly
through LLM-generated job descriptions) and task-specific fine-tuning (using MNRL) on the performance
of bi-encoder models. Our most efective system utilized a fine-tuned E5-Instruct bi-encoder, augmented
with LLM-generated job descriptions and oficial ESCO skill descriptions. This approach achieved a
mAP of 0.345 on the oficial TalentCLEF test set, demonstrating a robust and scalable solution.</p>
      <p>For future research, several avenues warrant exploration. Investigating the use of larger or more
specialized LLMs for job description generation, as well as experimenting with LLM-generated descriptions
for skills, could further enhance input representations. Additionally, more extensive hyperparameter
tuning for the fine-tuning process and dedicated experimentation with more sophisticated instruction
prompts for E5-Instruct may unlock additional performance. Finally, exploring a two-stage hybrid
architecture, where our eficient fine-tuned bi-encoder retrieves an initial set of candidates subsequently
re-ranked by a more computationally intensive cross-encoder, could ofer a compelling balance of
eficiency and accuracy. These directions could build upon the insights gained in this study to further
advance the state-of-the-art in job-skill matching.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author used Gemini 2.5 Pro in order to: Paraphrase and reword,
Grammar and spelling check. After using this tool, the author reviewed and edited the content as
needed and takes full responsibility for the publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Gasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fabregat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>García-Sardiña</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Estrella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Deniz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rodrigo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zbib</surname>
          </string-name>
          ,
          <article-title>Overview of the TalentCLEF 2025: Skill and Job Title Intelligence for Human Capital Management, in: International Conference of the Cross-Language Evaluation Forum for European Languages</article-title>
          , Springer,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>le Vrang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Papantoniou</surname>
          </string-name>
          , E. Pauwels,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fannes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vandensteen</surname>
          </string-name>
          , J. De Smedt,
          <article-title>Esco: Boosting job matching in europe with semantic interoperability</article-title>
          ,
          <source>Computer</source>
          <volume>47</volume>
          (
          <year>2014</year>
          )
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Van Der</given-names>
            <surname>Goot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Plank</surname>
          </string-name>
          , Escoxlm-r:
          <article-title>Multilingual taxonomy-driven pre-training for the job market domain</article-title>
          ,
          <source>arXiv preprint arXiv:2305.12092</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <article-title>Multilingual e5 text embeddings: A technical report</article-title>
          , arXiv preprint arXiv:
          <volume>2402</volume>
          .05672 (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.-J.</given-names>
            <surname>Decorte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Van</given-names>
            <surname>Hautte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Demeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Develder</surname>
          </string-name>
          ,
          <article-title>Jobbert: Understanding job titles through skills</article-title>
          ,
          <source>arXiv preprint arXiv:2109.09605</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Nils</given-names>
            <surname>Reimers</surname>
          </string-name>
          and
          <article-title>Sentence Transformers contributors, all-mpnet-base-v2, https://huggingface.co/ sentence-transformers/all-mpnet-base-</article-title>
          <string-name>
            <surname>v2</surname>
          </string-name>
          ,
          <year>2022</year>
          . Accessed:
          <article-title>April 2025</article-title>
          .
          <article-title>Hugging Face Model Card</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Grattafiori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dubey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jauhri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pandey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kadian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Al-Dahle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Letman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mathur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schelten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaughan</surname>
          </string-name>
          , et al.,
          <source>The llama 3 herd of models, arXiv preprint arXiv:2407.21783</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Henderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Al-Rfou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Strope</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Sung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lukács</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Miklos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kurzweil</surname>
          </string-name>
          ,
          <article-title>Eficient natural language response suggestion for smart reply</article-title>
          ,
          <source>arXiv preprint arXiv:1705.00652</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>I.</given-names>
            <surname>Loshchilov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hutter</surname>
          </string-name>
          ,
          <article-title>Decoupled weight decay regularization</article-title>
          ,
          <source>arXiv preprint arXiv:1711.05101</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>