<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Dynamic Jobs-Skills Knowledge Graph</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alejandro Seif</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sarah Toh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hwee Kuan Lee</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Artificial Intelligence Practice, GovTech</institution>
          ,
          <addr-line>S117438</addr-line>
          ,
          <country country="SG">Singapore</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Bioinformatics Institute, Agency for Science, Technology and Research (A</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>STAR)</institution>
          ,
          <addr-line>S138671</addr-line>
          ,
          <country country="SG">Singapore</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Skills Development Group, SkillsFuture Singapore</institution>
          ,
          <addr-line>S408533</addr-line>
          ,
          <country country="SG">Singapore</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Adult learners aiming to upskill themselves are often bombarded with overwhelming and contradicting information in terms of what skills are relevant to their occupations at a given time. Companies and governments also face the same concerns as they navigate ways to optimally manage their workforce and citizens' employability. The goal of this paper is to explore the use of knowledge graphs to better understand the complex relationship between occupations and skills, aiming to provide a cohesive answer to the question of how a skill relates to an occupation. We focus on harnessing taxonomies represented as relationships within a temporal knowledge graph. The contribution of this paper is a methodology to construct a comprehensive Jobs and Skills knowledge base that evolves dynamically over time, starting with the Singapore SkillsFuture Skills Framework as a foundational resource. Through the integration of expert knowledge and extracted entity information from labour market data, we employ data analytics techniques to refine and update the knowledge base to accurately reflect the evolving needs of the Singaporean economy. Through this, technologists in organizations working with jobs and skills can seek to build knowledge bases that incorporate expert knowledge and labour market signals so as to answer jobs and skills questions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;adult learners</kwd>
        <kwd>data analytics</kwd>
        <kwd>education</kwd>
        <kwd>graph database</kwd>
        <kwd>job posting</kwd>
        <kwd>knowledge graphs</kwd>
        <kwd>labour market</kwd>
        <kwd>occupations</kwd>
        <kwd>online job ads</kwd>
        <kwd>property graph</kwd>
        <kwd>skill mismatch</kwd>
        <kwd>taxonomies</kwd>
        <kwd>temporal knowledge graphs</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The cacophony of signals coming from workforce
development specialists, human resource consultants and job
market advocates, including adult education oferings, can
leave individuals seeking to upskill (or re-skill) themselves
confused about what skills are critical for a certain
occupation and which are not.</p>
      <p>Companies express their needs by hiring individuals with
specific skills. Hence, incorporating labour market data
from job postings is crucial to get a direct reading of the
skills demand for given occupations. Additionally, expert
knowledge coming from specialists in various sectors of
the economy presents contextualised views of the overall
role that specific occupations and skills play. Together, the
labour market insights and domain knowledge give a holistic
picture of how skills and occupations are related at certain
point in time.</p>
      <p>In this paper, we will focus on unlocking the power of
taxonomies represented as relationships between entities
in a temporal knowledge graph to address the question
of what skills are required by occupations and how those
relationships evolve over time.</p>
      <p>
        With the transition from traditional newspapers to
online job portals, researchers initially [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] turned towards
natural language processing (NLP) and topic modelling
techniques to automate the identification of skills from job
advertisements and their categorization within specific fields
or occupations. However, concerns [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] persisted regarding
the generalization of skill identification methods. Some
approaches to mitigate these concerns involved the creation of
comprehensive skill bases that encompass diverse skill
terminologies to enhance skill identification in job ads. Human
resources experts have published standardized occupations
and skill classifications such as ISCO [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], ESCO [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and
O*NET [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Some authors have further tailored published
knowledge bases for the particular economic needs of their
own country [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        SkillsFuture Singapore (SSG), the national skills authority
of the Singapore public service, together with employers,
industry associations, institutes of higher learning and unions,
has created its very own occupation and skill knowledge
base, the SSG Skills Framework [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The SSG Skills
Framework provides information on key sectors, occupations, job
roles, and the required existing and emerging skills. The
Singapore Department of Statistics also maintains an
occupational classification known as Singapore Standard
Ocupational Classification (SSOC) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which is used by the Skills
Framework to classify Job Roles (Skills Framework) into
their respective Occupations (SSOC).
      </p>
      <p>
        Relying on rich occupational and skill bases, labour
market data analytics and knowledge graph applications have
emerged as pivotal tools for understanding workforce
dynamics and optimizing skills matching in the job market.
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] introduces the concept of the Occupation Space, a
network representation of French workers’ job mobility
patterns, revealing intense skill-relatedness between
occupations. Leveraging O*NET, [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] highlights quantification of
skill polarization in the labour market, impacting career
mobility and wage distribution. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] review explores
knowledge graph-based frameworks for education and
employability, enhancing skill-job matching and skill-course alignment.
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] survey of the state of the art underscores the significance
of online job ads in providing large-scale insights into job
market needs, fueled by advancements in artificial
intelligence. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] presents a graph-theoretic approach to skills
valuation, relying on O*NET data, emphasizing skills’
facilitation of occupational transitions as a measure of value.
These approaches however, take existing knowledge bases
as-is, even though that knowledge represents a snapshot in
time when they were created and not a constantly updated
source of information.
      </p>
      <p>
        Meanwhile, [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] discusses the importance of knowledge
graph refinement and the challenges in measuring and
validating such updates. In this context, [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] proposes a Skills
and Occupation Knowledge Graph that leverages ISCO and
ESCO, refined with job posting data, to facilitate skills-based
job matching and career pathfinding. As opposed to static
knowledge graphs where facts do not change with time,
temporal knowledge graphs [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] incorporate new
information over time to refine the relationships. It is in this area
that we would like to frame our study.
      </p>
      <p>Our contributions are as follows: (i) presenting a
comprehensive Jobs and Skills knowledge base, and (ii) creating
an approach for that knowledge base to evolve over time.
Using the SSG Skills Framework as a starting point, the Jobs
and Skills knowledge base combines expert knowledge with
extracted entity information from large amounts of
unstructured data such as job postings, course listings and other
available datasets. Data analytics in the form of machine
learning classifiers and sliding window weighted averages
are used to refine the information extracted from
unstructured data and integrate it with expert knowledge, enabling
the Jobs and Skills knowledge base to evolve to reflect the
needs of the ever-changing Singapore economy.</p>
    </sec>
    <sec id="sec-2">
      <title>2. The Jobs and Skills Knowledge</title>
    </sec>
    <sec id="sec-3">
      <title>Graph ( JSKG)</title>
      <p>
        A knowledge graph is a structured model of information,
comprising entities, relationships, and semantic
descriptions [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The growing availability of diverse data sets
has prompted researchers to explore semantic and
conceptual representations, leading to the popularity of knowledge
graphs (KGs). KGs serve as fundamental components for
various information systems that rely on structured
knowledge [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>Within knowledge graphs, there is a sub type called
property graphs (example provided in Fig. 1). These graphs
represent data using nodes (entities) and relationships (edges)
with properties associated with both nodes and
relationships. These graphs are flexible and intuitive, making
them suitable for representing complex relationships and
attributes within data sets. By following the direction of the
relationship, it is possible to compose sentences in English
that provide the semantic information of the relationship
(e.g. Occupation (Instrumentation Engineer) - REQUIRES
SKILL→− Skill (electrical, electronic and control
engineering)")</p>
      <p>We present the Jobs Skills Knowledge Graph (JSKG), a
knowledge graph that can be queried based on the proximity
of entities within it. By capturing the information in the
form of a knowledge graph, relationships between entities
such as Occupations and Skills can be easily investigated to
determine commonalities and relatedness between entities.
Figure 2 provides an example of various types of nodes and
their relationships. The following section provides details
on how the JSKG was constructed, with a focus on the
Occupation-Skill relationships.</p>
      <p>The JSKG itself is structured as a property graph and
stored in a graph database, which can be queried through
an API. A website and a large language model-enabled
chatbot make use of this API to consume the results of queries
between entities (e.g. two occupations) and present the
information in a compelling way to non-expert users.</p>
      <p>Details on the implementation and software tools utilised
can be found in the Appendix A.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Data Sources</title>
      <p>We leverage a comprehensive dataset encompassing three
types of entities to examine the evolving landscape of skills
demand and occupational structures in Singapore.</p>
      <p>Firstly, the primary data source comprises of job postings
collected from the internet spanning the years 2018 to 2023,
ofering insights into the dynamic nature of skill demand
within the job market over this period. We shall refer to this
as Labour Market data.</p>
      <p>
        Secondly, complementing the labour market data, we
incorporate expert knowledge from SkillsFuture’s
Singapore Skills Framework (SFw) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], a knowledge base for job
roles and skills, which also includes a mapping between job
roles and the Singapore Standard Occupation Classification
(SSOC) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>Thirdly, we also include data on courses and training
providers validated by SkillsFuture Singapore to enrich the
JSKG with information on options for increasing skills
supply to meet market demand.</p>
      <p>The integration of these diverse entities allows for a
nuanced exploration of the intersections between job market
dynamics, skill evolution, and occupational structures,
contributing to a comprehensive understanding of the
Singaporean workforce and continuous education landscape.</p>
      <sec id="sec-4-1">
        <title>3.1. Singapore Standard Occupation</title>
      </sec>
      <sec id="sec-4-2">
        <title>Classification</title>
        <p>The JSKG uses the Singapore Standard Occupation
Classification (SSOC) from the year 2020, which ofers a
standardized framework for classifying and organizing occupations
within the Singaporean context. SSOC uses a 5 digit
hierarchical classification, so that every 5 digit code is contained
within a 4 digit code and so on for for fewer digits. Table 1
contains an example SSOC 3-digit, 4-digit and 5-digit:</p>
        <p>At the highest level of abstraction, major groups of
occupations are categorised at the SSOC 1D level given by Table
2.</p>
        <p>Throughout the rest of this manuscript, we’ll refer to
SSOCs and Occupations interchangeably.
Job portals such as "LinkedIn" or "Indeed.com" are widely
used by employers to advertise Job Postings for candidates
to apply to. Given that the job market evolves continuously,
periodically analysing Job Postings provides a way to
measure the relative changes in demand for workers, as well as
the evolving descriptions of the Job Roles they perform.</p>
        <p>By relying on data purchased from a job aggregator
company, the authors analysed over 20M job postings between
years 2018 and 2023. Each job posting is defined by a title
and plain text description. Given the unstructured nature of
the data, additional tools were needed to classify a job
posting by the occupation it represents and to identify the skills
it requires. More details on the classification and extraction
tools are presented below in section 3.5.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. SSG Skills Data</title>
        <p>SkillsFuture Singapore (SSG) has three variants of skills data.
In this section we’ll present them in greater detail.</p>
        <sec id="sec-4-3-1">
          <title>3.3.1. SSG’s Skills Framework (SFw)</title>
          <p>SSG’s Skills Framework (SFw) provides an expert view on
the relationships between Occupations and Skills. It is
important to note that the SFw defined a taxonomy of skills
which were classified into two types: Technical Skills and
Competencies (TSCs) and Critical Core Skills (CCS).</p>
          <p>TSCs refer to over 11,000 skills that are typically
nontransferrable across occupations, such as "Digital
Marketing" or "Airport Service Quality Management". CCSs refer
to 16 skills typically referred as soft skills, such as "Problem
Solving" or "Communication". Every Job Role in the SFw has
TSCs and CCSs associated with it. In turn, every Job Role
is linked to an occupation found in the SSOC (Occupations
are a superset of Job Roles).</p>
          <p>We note that the SFw accounts for similar skills
belonging to diferent sectors as two diferent skills. For example,
"Digital Marketing" in the "Tourism" sector is considered
a diferent skill than "Digital Marketing" in the "Financial
Services" sector. This is because the SFw is developed and
validated by domain experts representing diferent sectors,
such as other statutory boards or industry and trade
associations.</p>
          <p>Additionally, each skill in the SFw is further broken down
into proficiency levels. Each proficiency level, represents a
set of knowledge and abilities. An example of this is the TSC
skill "Financial Modelling" which has a level 3 proficiency
associated to the knowledge "Basic accounting theories of
recording and reporting financial information" and a level 6
proficiency associated to the knowledge "Specialised models
including real options".</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>3.3.2. Sector- and proficiency-agnostic skills</title>
          <p>As mentioned in the subsection 3.3.1, the SFw labels skills
using sector allowing for some of the skill titles to be duplicates
or synonyms of each other across diferent sectors. This is
also the case for similar skills classified into diferent
proficiency levels. In order to obtain a more general view, SSG
has created sector- and proficiency-agnostic skills. After
de-duplicating and mapping, there are ∼ 2, 000 sector- and
proficiency-agnostic skills (compared to ∼ 11, 000
sectorbased skills). We note that this de-duplication was
performed based on expert knowledge.
An example of the mapping is presented in the Table 3.</p>
          <p>During the extraction of skills from a job posting, there
might not be any information on what sector the skill
belongs to. The sector-agnostic skills are then a suitable choice
for automatic extraction of skills from raw text. More details
on the extraction will be provided in section 3.5.</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>3.3.3. Applications and Tools</title>
          <p>The skills considered up until now are generic. For example,
the skill programming and coding does not specify which
programming language is involved. When looking for courses
and job postings, specifics can be very important. For this
reason, SSG has created a list of ∼ 1000 Application and
Tools.</p>
          <p>Applications and Tools are often required to deliver
specific tasks of many occupations. These refer to brand specific
software applications or tools such as the programming
language Python or the 3D creation platform Unreal Engine.
As opposed to SFw Skills, Applications and Tools are not
related to a specific sector of the economy.</p>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>3.4. Courses</title>
        <p>The third data source is courses present on the SkillsFuture
Singapore portal. Each of these courses has a description of
its learning goals (in plain text), alongside other information
such as the fees and duration. A subset of the courses also
have specific SFw skills and proficiency levels associated to
them with the validation of SkillsFuture Singapore.</p>
        <p>In order to link Labour Market and Course data with the
SFw, we leveraged two diferent tools to identify the Skills
associated with every job post and course listing, and the
SSOC associated with every job post analysed. The tools
were not developed by the authors, but we will provide an
overview of how they work in the following section.</p>
      </sec>
      <sec id="sec-4-5">
        <title>3.5. Tools to link Labour Market and Course data with SFw</title>
        <p>Two key entities are required to analyse job postings and
course data - Occupation associated with any job posting,
and the skills required by a job or taught by a course. By
extracting these entities from any job post or course listing,
it is possible to aggregate and create a knowledge graph
representation in which the entities are Occupations, Skills
and Courses.</p>
        <p>
          Two proprietary tools, developed by other teams within
the Singapore Public Service, were utilised to achieve this
linkage. These two tools rely on Natural Language
Processing (NLP) and Named Entity Recognition (NER) ([
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]
[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]) to perform extractions from unstructured text. In the
context of this paper we will only provide a high level
explanation of how they work, but will otherwise take them
as black boxes as our contribution is centered around how
the knowledge base is constructed.
        </p>
        <p>
          First we will present the tool to extract SSOC: The SSOC
Autocoder. The SSOC Autocoder is a tool that leverages
transformers [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] and neural networks to determine the
most likely 5-digit SSOC code for a body of unstructured
text of &lt; 300 words. Its loss function is optimized so that
the loss is maximized when the first digits are incorrect,
and minimized when the last few digits are incorrect. In
this way, if the results are not correct to the 5 digit level, it
is possible to make use of the 4-digit code to retain some
accuracy.
        </p>
        <p>Secondly, we will present the proprietary tool to extract
Skills: The Skill Extraction Algorithm (SEA). Given the
possibility of highly similar skills existing in diferent SFw sectors,
the SEA developers decided to utilise the ∼ 2000
sectorand proficiency-agnostic de-duplicated skills (see section 3),
removing the sectoral dependence. Including the curated
list of approximately 1000 Applications and Tools, the SEA
is able to extract 3000 unique skills, where each skill has a
type given by one of the following: ’TSC’, ’CCS’ and
’App/Tools’. The algorithm uses a NER transformer model and
cosine similarity to identify skills that most closely match
spans extracted from unstructured text found in Labour
Market job postings and course descriptions.</p>
        <p>We highlight that if several job postings for the same
occupation have skills extracted using SEA, it is then possible
to count how often each skill is needed for a given
occupation. This ability to quantify the presence of a skill for a
given occupation will be utilised in the following section to
rank the importance of a skill.</p>
      </sec>
      <sec id="sec-4-6">
        <title>3.6. Entities in the Jobs and Skills</title>
      </sec>
      <sec id="sec-4-7">
        <title>Knowledge Graph ( JSKG)</title>
        <p>The knowledge graph is composed of entities intrinsic to
the topic of jobs and skills. An example of such entities and
their relationships is depicted in Fig 2. The entity types in
the JSKG are comprehensively listed here:
1. SFw Skills: Defined through expert knowledge in the</p>
        <p>
          Skills Framework.
2. Skills: Sector- and proficiency-agnostic skills &amp;
Application and Tools, extracted using the Skills
Extraction Algorithm.
3. Occupations: As defined by the SSOC 2020 [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
4. Courses: Available on the MySkillsFuture web portal
[
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]
5. Training Providers: Providers of courses on the
        </p>
        <p>MySkillsFuture web portal
6. Proficiencies : Specific capabilities within a SFw Skill
7. Job Roles: Specific roles within an SSOC occupation
At this point we have presented the data sources and the
processing done on them to extract the entities (nodes) and
relevant relationships between them. In the next section we
will present how the entities are linked in the JSKG.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Methodology</title>
      <sec id="sec-5-1">
        <title>4.1. Using Job Postings to Proxy Evolving</title>
      </sec>
      <sec id="sec-5-2">
        <title>Demand for Occupations and Skills</title>
        <p>Job Postings, found on Job Portals such as "LinkedIn",
"JobStreet" and "MyCareersFuture", provide a window into the
instantaneous demands of the labour market. However, we
must highlight that these portals represent a proxy for the
total number of jobs in the entire economy. Specific
companies/sectors of the economy might rely on alternative
ways to attract job seekers and it is important we highlight
that a pure reliance on Job Posting portals does not give a
complete picture1.</p>
        <p>Additionally, the retrieval at scale of Job Posting from
such portals can be a source of inaccuracy in itself. The
portal’s changes in popularity by employers posting jobs,
server downtime and errors in data processing can lead to
lfuctuating sampling of Job Postings. The outcome of these
variations can be large diferences in the volume of job
postings for two diferent time periods that are not solely
due to an observed economic shift in the workforce, but
rather potentially attributable to duplicated postings,
repostings and intermittent accessibility to portals, just to
name some causes. For example, the year of 2018 could have
a total of ∼ 1.5 million job postings, whereas a diferent
time period (the year of 2019) can have up to ∼ 2.3 million
job postings 2</p>
        <p>This immediately leads us to realize that comparisons on
extensive quantities (e.g. job post counts across years) are
unreliable. A specific occupation could garner double its job
post count from one period to the next, but that could be
a data sampling problem rather than a reflection of a stark
change in the economy.</p>
        <p>To mitigate this problem, we opt to represent and compare
quantities derived from Job Postings in terms of intensive
quantities (e.g. ranking), which are less sensitive to sampling
lfuctuations.</p>
        <p>Particularly for the case of quantifying the relationship
between an occupation and a skill, we can compare the top
 most popular skills for a given occupation across the
years analysed without having to be as concerned about
the absolute number of job postings. The popularity of the
skill, for a given occupation, is given by counts of skill
presence when extracting job postings associated to a specific
occupation. By ranking these counts for a given occupation
, we can then compute a ranking (, , ) of skills  in
each time period . Below, we will present the notation to
represent the relationships between occupations and skills.</p>
        <sec id="sec-5-2-1">
          <title>4.1.1. Determining skill rankings using counts of skill presence</title>
          <p>For each skill  and occupation  relationship ( − ),
we compute a ranking (, , ) in terms of the weight
between  −  for a given time period . The weights are
based on the confidence of the SEA extraction, the frequency
with which a job posting mentions a skill, and the frequency
with which a skill is required for an occupation. Additionally,
1We’d like to note that this study solely focuses on Job Postings listed
on Singapore.
2There is no observable macroeconomic indicator that signaled that
Singapore experienced a ∼ 53% growth in demand of workers, more
likely, this increase in job postings is due to a variation in sampling.
we can also rank the −  relationships derived from expert
knowledge in the SFw, which is represented as  (, ).</p>
          <p>In section 4.2 we’ll use the 1/(, ) as a score for every
 −  relationship. The same is applied to  (, ) for
−  relationship derived from the SFw. In this way, we can
aggregate scores over time periods and incorporate expert
knowledge and labour market data. More importantly, it
also enables us to deal with the issue of missing  − 
relationships during certain periods. By working uniquely
with rank, we would be forced to assign a rank value to a
missing skill. Operating with inverse values allows us to
assign 1/(, ) = 0 for missing skills. Table 4 depicts an
example for the Occupation Data Scientist on selected skills
and year.</p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>4.2. Refinement of the Occupation-Skill relationships</title>
        <p>The temporal aspect of the JSKG requires relationships to
be updated as new information comes in. In particular, this
afects the relationships between occupations and skills. Our
goal is to compute a score that can be updated over time to
better reflect the reality of the changes in the job market.</p>
        <p>Every time period, a new dataset of newly posted job
postings is analysed. Aggregated relationships between
skills and occupations are thus computed for that period.</p>
        <p>Although we keep relationships labeled by year, having an
overall consolidated relationship between occupations
and skills is best for analysts that need holistic
information, while retaining information of the evolution of the
relationship.</p>
        <p>Using the SFw as a starting point, we can use each new
set of job postings to update or reinforce the original expert
knowledge (if any).</p>
        <p>There are three considerations at this point:
1. Skills of type App/Tools, which are not present in
the SkillsFuture’s Singapore Skills Framework (SFw).</p>
        <p>Their starting weight derived from expert
knowledge is zero.
2. The SFw had by design 6 skills associated to a given
occupation. The SEA is unbounded and able to
retrieve any number of skills from a single Job Posting.</p>
        <p>This implies that skills in SFw will have a boosted
weight, as compared to those that are not in the SFw.
3. Expert knowledge (SFw) on TSCs and CCSs is
considered ever-green. The weight contribution from
the SFw should not decay over time.</p>
        <p>To address these considerations, we make a judgment call
to take the SFw as a starting point, from which variations are
made based on signals from labour market data. Secondly,
we compare App/Tools separately from TSCs and CCSs, to
prevent these from consistently having weaker relationships
(by virtue of having zero representation in the SFw).</p>
        <p>To control the influence of the ever-green SFw, we can
introduce parameter 0 &lt;  SFw ≤ 1 to regulate how strong
the influence of expert knowledge is on the computed score
from labour market signals (job postings). The more time
passes before the SFw is updated with expert knowledge,
the more  SFw should approach 0.</p>
        <p>We propose the mathematical formula given by Eq 2 to
compute an evolving score for the weight  of the
relationship between an occupation  and a skill . LM represents
the ranking of the values of a sliding window of length
∆  in which weights over the years are computed using
We present the ranking and scores for Occupation Data Scientist and selected skills over the years. We can observe the values
of (, , ) and 1/(, , ) for  = 2018 and the expert knowledge from the Singapore SkillsFuture Skills Framework
(SFw), handling the missing skill as 0 contribution.</p>
        <p>Skill
Data and Statistical Analysis</p>
        <p>Programming and Coding
Mathematical Concepts Application
missing
(SFw)</p>
        <p>(2018)
1
2
2
5
38
( )
1</p>
      </sec>
      <sec id="sec-5-4">
        <title>4.3. Computing diferentiator skills for an occupation</title>
        <p>
          While the frequency with which a skill is required by an
occupation is an indicator of its importance for the tasks
the occupation entails, it is also useful to be able to identify
skills that diferentiate one occupation from another,
especially within a cluster of similar occupations. To achieve
this, similar to [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], we adopted Term Frequency - Inverse
Document Frequency (TF-IDF) given by Eq. 3 to identify
diferentiator skills for each occupation relative to the full
set of SSOC 2020 occupations, as well as clusters defined by
the SSOC 2020 hierarchical taxonomy.
        </p>
        <p>Using the consolidated skills requirement as the
relationship between occupations and skills, diferentiator skill
scores are calculated using this formula:
diferentiator score</p>
        <p>=
The terms in formula are:</p>
        <p>1
LM
×
︂(
log
︂( 1 + N
1 + n
+ 1
︂)
(3)
• LM is the rank of the Consolidated Skill Requirement
detailed in the section above.</p>
        <p>the specified cluster
•  refers to the total number of occupations within
•  refers to the number of occupations within the
cluster requiring the skill in question</p>
        <p>This formula represents an adaptation of TF-IDF, in which
we substitute the inverse rank of consolidated skills
requirement relationships for TF. The inverse rank gives a measure
of how important a skill is to a specific occupation. This
substitution is necessary because consolidated skills requirement
relationships do not have a direct measure of the frequency
with which a skill is needed for a given occupation, which
is a more direct parallel for TF.
diferentiator skills ("DIFFERENTIATOR").
relationships through Large Language</p>
      </sec>
      <sec id="sec-5-5">
        <title>Models</title>
        <p>
          The recent advances in Large Language Models (LLMs) [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ],
particularly applied to tasks involving Retrieval Augmented
Generation (RAG) [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], provide us with an opportunity to
perform an otherwise time-consuming human-supervised
evaluation of the skills associated to occupations (SSOC),
while minimising the efect of hallucinations by providing
contextual information. In this case, the team leveraged
the AI tool "Open AI GPT-4o" [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] function calling, together
with a team of Jobs Skills Analysts from SkillsFuture SG to
evaluate the labeling of skills-to-occupations relationships
produced by GPT-4o (in the form of confidence labels). The
way the RAG was utilised to minimise hallucination, was in
terms of providing the SSOC Occupation definition (a
paragraph of text) given by [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and the top 50 skills associated
to an SSOC 5D through consolidated skills requirement. The
output of GPT-4o was two labels for every such
relationship between skill and occupation - True / False for whether
the skill was relevant to the occupation, and a confidence
level for this evaluation - high, medium or low, where high
implies high confidence in the evaluation.
        </p>
        <p>
          Below is the methodology followed:
1. Take top  = 50 TSC and CCS skills associated to
an SSOC given by consolidated skills requirement
2. Take the title and paragraph definition of SSOC 5D
definition given by [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
3. Prompt OpenAI’s GPT-4o to label every selected TSC
and CCS skill linked to an SSOC 5D given its
occupation definition. Each skill-occupation relationship
is to be labeled True or False with confidence labels
high, medium, low.
4. Leverage on a team of human Job Skills Analysts in
SkillsFuture Singapore to evaluate a sample of the
relationship labels produced by GPT-4o
        </p>
        <p>The Jobs Skills Analysts evaluated a sample of ∼
SSOC 5D across the total 1002 SSOC 5D in the
classifica400
tion. In that sample, they found that 67% of relationships
labeled False with high confidence should be dropped, as
they deemed that the skills are unrelated to the occupations.
For skill-to- SSOC relationships labeled False by GPT-4o
with medium and low confidence, the human evaluation
indicated that most relationships are not unrelated. For
those with medium confidence,
&lt; 5% were deemed to be
unrelated, while for those with high confidence, 0% was
deemed to be unrelated. For relationships labeled True, for
those with low confidence,</p>
        <p>&lt; 1% were deemed to be
unrelated and 0% of those labeled mid, high were found to be
unrelated. In favour of leaning towards less but true
values, all relationships labelled False with high confidence by
GPT-4o were dropped across all SSOC 5D. The downstream
implication is that now the top 50 ranked skills might be
less than 50 distinct skills on several SSOC 5D. The
following section presents the results obtained in applying these
methodologies.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Results: Bringing it all together</title>
      <sec id="sec-6-1">
        <title>5.1. The JSKG schema</title>
        <p>By leveraging the entities presented in section 3.6 we can get
a full picture of how occupations and skills come together.
In Figure 4 we present the example of the Instrumentation
Engineer mentioned in Table 5.</p>
        <p>Given that the focus of our discussion has been
regarding the refinement/updating of the SSOC-Skill relationship,
we’d like to present two illustrative use cases of how the
JSKG can be used to derive value for two distinct types of
users.</p>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Evaluation of the Labour Market</title>
      </sec>
      <sec id="sec-6-3">
        <title>Inferred Skills</title>
        <p>In order to provide an evaluation of the methods to extract
skills across all SSOC, we consider the SFw Skill-Occupation
linkage as ground truth and run a comparison with the skills
associated purely through the labour market data as
inference between the years 2020 to 2023. Borrowing from
machine learning approaches, we take the top  skills
extracted by those two approaches and compare how these
skills overlap using precision and recall metrics in terms
of true values (Top  SFw TSC+CCS Skills) and inferred
values (Top  Labour Market TSC+CCS Skills). The
evaluation can be made in terms of the confusion matrix presented
in Table 6. App/Tools must be excluded from the comparison,
as these are not present in the Skills Framework.</p>
        <p>Precision and recall are given by the following formulas:
   
precision  =   +  and recall  =   +  . We
can contextualise our findings in terms of the domains of
SSOC 1D and by the median amount of job postings analysed
per SSOC 5D, which is captured as Job Density also on Table
7.</p>
        <p>Table 7 also includes the amount of skills-occupation
relationships dropped using the approach presented in 4.4,
aggregated at the SSOC 1D level. We can observe that high
precision is correlated with lower drop rates of
occupationskill relationships. Definitions for each of the major groups
in SSOC 1D are given by Table 2.</p>
        <p>Across all SSOCs found in both the SFw and labour market
data, 71% of the top 10 skills are the same. This decreases
to 59% for the top 20 skills and 37% for the top 50 skills.
Some divergence in the skills signals for the labour market
and the SFw is expected due to changes over time, so these
ifgures suggest that the skill-occupation relationships in the
knowledge graph are most useful when considering the most
important skills for each occupation. This is particularly so
for jobs that are advertised more commonly in Singapore
(high job post count). Possible ways to improve the rest
of the skill-occupation relationships further are covered in
section 7.1.</p>
      </sec>
      <sec id="sec-6-4">
        <title>5.3. Illustrative Case Studies</title>
        <p>Next, we illustrate the use of the JSKG by providing two
diferent customer use cases.</p>
        <p>Mid-career Switcher. The user is a mid-career
marketing manager who is considering their options to exit their
current occupation in favour of one that is more rewarding
monetarily and that brings some new challenges. The user
needs to be able to take stock of their existing Skills and
expected pay in their occupation, and use that information
to gauge adjacent occupations in terms of Skill similarity
and salary increment.</p>
        <p>The user is able to tap on a web portal linked to the JSKG
to perform a comparison in terms of skill similarity,
obtaining the occupations that have the most skills in common,
while providing an increase in pay. The user can then
retrieve the skill gap through an occupation comparison and</p>
        <p>Jobs Density</p>
        <p>Job Posts</p>
        <p>Drop Rate
ifnd optimal courses that will equip them with the required
skills for the best cost and least amount of time. In this
example, the marketing manager can identify the occupation
of Business Development Manager as an occupation with
9 skills in common and a skills gap of 9 skills. The skill
gap can then be fulfilled by identifying the optimal course
set required to achieve such skills for the minimum cost,
duration, and wait time to earliest opportunity for training.</p>
        <p>At this point, the user can get a sense of the dificulty
involved in making the career switch, but can also quantify
the monetary and time commitment required to achieve
basic literacy to bridge the skill gap and achieve their goals
of switching careers.</p>
        <p>The Learning and Development HR analyst. Some
of the key tasks for a L&amp;D analyst involve three aspects: (i)
analyzing the gaps between the current skills of employees
and the skills required for their roles or for future
organizational needs; (ii) Designing and developing training
programs and learning materials tailored to address
identiifed skill gaps;(iii) Implementation of Training Programs to
re-skill or up-skill the workforce to achieve organizational
needs organically.</p>
        <p>The L&amp;D analyst, might be tasked to identify roles for
which demand is declining to train individuals to move into
adjacent occupations with higher demand, such as training
secretaries to become compliance oficers. By comparing
occupations, the analyst identifies the usual skill gaps
between the roles. In this way, it is possible to determine the
optimal course set to obtain the missing diferentiator skills
to achieve basic literacy in the new roles. By considering
the costs in terms of fees for training, the analyst can also
determine the optimal training strategy and the programs
to carry out such transformation.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Lessons Learned and Challenges</title>
      <p>One of the primary challenges encountered was with the
Skills Extractor Algorithm (SEA), which at times provided
incorrect skill extractions. This issue arose from the need for
the tool to operate quickly enough to ensure computational
feasibility. Consequently, skills were sometimes omitted or
incorrectly extracted based on the phrasing of job
descriptions. To curb the incorrectly extracted skills, we found
that ranking extracted skills by their popularity typically
relegated incorrect extractions to lower tiers. Furthermore,
relying on ground truth SkillsFuture (SFw) skills and
updating them with labour market data helped to purge incorrect
extractions. Additionally, using the GPT-4 Large Language
Model (LLM) to prune skills, based on its general knowledge,
further reduced the number of incorrect extractions, albeit
at the cost of occasionally removing some correct ones.</p>
      <p>We also learned that rankings, rather than absolute
values, provide more intuitive insights to users. Given the
wide variation in the popularity of occupations, absolute
numbers often do not convey meaningful information. For
instance, in our job posting dataset for 2018, the occupation
"Marketing Manager" had 2039 extractions, while
"Management Consultant" had only 13. In that same year, the skill
"Business Opportunities Development" appeared 314 times
for "Marketing Manager," making it the 13th most popular
skill for that occupation, whereas for "Management
Consultant," the same skill appeared 10 times, ranking as the most
popular skill for that occupation.</p>
      <p>We faced the dilemma of whether to feature jobs that are
not typically ofered in the job market, such as "Legislator,"
through synthetic methods such as large language models
(LLMs). The absence of such jobs in the data is likely due to
the nature of hiring for those occupations rather than a lack
of data. However, given the definitions of these occupations
and their associated tasks, it is feasible to include them in
the knowledge graph through synthetic means.</p>
      <p>Another challenge was dealing with changes in skill or
occupation taxonomy. Should the SEA extraction tool adopt
a new list of skills (e.g., new Technical Skills and
Competencies (TSCs) included) or if a new version of the SSOC is
released, a re-extraction of SSOCs or skills across all
historical job postings would be necessary. This re-extraction is
required for consistency and to ensure fair comparison.</p>
      <p>Finally, Section 4.2 discusses the use of 1/(, ) as a
score to compute the aggregated efect within a sliding time
window and diferentiator scores. We found that using a
succession of 1 provides an over-estimation of the diference
between the highest-ranking skills compared to directly
comparing rankings (, ). However, the variation in the
number of elements across rankings makes the direct use of
(, ) challenging, particularly when a skills-occupation
pair is absent in certain time periods. This 1 approach helps
address the problem of missing rankings, ofering a practical
solution.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Conclusion</title>
      <p>This research contributes to a deeper understanding of
workforce dynamics and supports timely and informed
decisionmaking in skills matching and career planning initiatives.
The creation of the Jobs Skills Knowledge Graph delineates a
dynamic repository of occupational roles, skills, and
associated entities within the realm of continuous adult education.
Leveraging occupation and skill classifiers serves as
foundational mechanisms for constructing a knowledge base of
occupations and competencies. Central to this endeavour
is the ongoing refinement and updating of relationships
between occupations and skills, which serves as
cornerstone for deriving ever-green graph analytics insights. The
collaboration of LLMs and human experts further refines
the end-result. An evaluation of the results of this
methodology is presented in terms of precision and recall, after
the human-LLM layered drops the more unrelated skill
associations. The evaluation results throw light on which
occupations major groups (SSOC 1D) are more accurately
represented by the methodology proposed.</p>
      <sec id="sec-8-1">
        <title>7.1. Future Work</title>
        <p>Through the linkages derived between occupations and
skills, it is possible to derive a network purely of skills (or
occupations). Such a network allows us to investigate the
inter-relationship of skills. This is particularly relevant for
the linkage between Technical Skill Competencies (TSCs)
and the ever growing list of Applications/Tools that allow
the skill to be applied.</p>
        <p>Investigating the network using centrality metrics or
clustering methods could enable us to identify key
clusters of skills that are highly transferable or that are often
required together. These can serve as strong
recommendations for training providers to adapt their oferings to the
ever-changing landscape of the needs of the economy.</p>
        <p>Robust evaluation of the results remains only partially
addressed. Although the precision and recall results in Table
7 give a quantified figure of how well the occupation
categories (SSOC 1D) match expert knowledge, this does not tell
us if the diference is due to an evolution in the job
definition in Singapore, or due to limitations in the approach.
Potential verifications could involve leveraging datasets from
other countries (which would need to be contextualised in
the taxonomies of occupation and skills used in this paper)
or leveraging on another round of costly Labour Market
experts to validate the findings with updated validated
relationships. Both approaches are highly taxing in terms
of efort and cost, however, rounds of human validation of
small samples can give indications of the benefits of the
method for smaller costs.</p>
        <p>
          Lastly, recent growing interest in applications between
knowledge graphs and large language models, underline the
importance of reliable knowledge bases to power such
models through graph retrieval-augmented generation
(GraphRAG) [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. Areas of further refinement lie in the creation
of triplets for the knowledge graph from unstructured text
[
          <xref ref-type="bibr" rid="ref28">28</xref>
          ] [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ], to the alleviation of hallucinations in Large
Language Models (LLMs) [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] through the facts captured in
knowledge graphs.
        </p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The authors would like to gratefully acknowledge the
support of SkillsFuture Singapore in facilitating the research
and application of the Jobs and Skills Knowledge Graphs.
The authors also extend their sincere thanks to Ong Meng
Hwee Victor for his valuable comments and help with the
review of the manuscript.</p>
      <p>Disclaimer: The content of this paper do not imply that
the research outcomes are currently deployed or planned
for deployment by SkillsFuture Singapore at any given time.
The views and opinions expressed in this paper are solely
those of the authors and do not necessarily represent those
of SkillsFuture Singapore, Government Technology Agency
of Singapore or A-STAR.</p>
    </sec>
    <sec id="sec-10">
      <title>A. Software tools utilised to build</title>
      <p>and deploy a proof of concept</p>
    </sec>
    <sec id="sec-11">
      <title>JSKG</title>
      <p>The graph database utilized was Neo4j version 5.3, selected
due to its eficiency and ease in handling relationships
within complex datasets. The application programming
interface (API) implemented for seamless communication
and data exchange was FastAPI.</p>
      <p>The web portal presented in Figure 5 was crafted using
Streamlit, a user-friendly and interactive framework for
data visualization and exploration. To ensure the
accessibility and availability of the developed API and web portal,
both were deployed on the Heroku platform, ofering a
scalable and reliable cloud infrastructure. For the hosting of
the Neo4j graph database, the AuraDB service provided by
Neo4j was utilized, facilitating a cloud-native, fully managed
database-as-a-service solution.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Debortoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. vom Brocke</surname>
          </string-name>
          ,
          <article-title>Comparing business intelligence and big data skills: A text mining study using job advertisements</article-title>
          ,
          <source>Business Information Systems Engineering</source>
          <volume>6</volume>
          (
          <year>2014</year>
          )
          <fpage>289</fpage>
          -
          <lpage>300</lpage>
          . URL: http:// dx.doi.org/10.1007/s12599-014-0344-2. doi:
          <volume>10</volume>
          .1007/ s12599-014-0344-2.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>A. De Mauro</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Greco</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Grimaldi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Ritala</surname>
          </string-name>
          ,
          <article-title>Human resources for big data professions: A systematic classification of job roles and required skill sets</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>54</volume>
          (
          <year>2018</year>
          )
          <fpage>807</fpage>
          -
          <lpage>817</lpage>
          . URL: http://dx.doi.org/10.1016/j.ipm.
          <year>2017</year>
          .
          <volume>05</volume>
          . 004. doi:
          <volume>10</volume>
          .1016/j.ipm.
          <year>2017</year>
          .
          <volume>05</volume>
          .004.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khaouja</surname>
          </string-name>
          , I. Kassou,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ghogho</surname>
          </string-name>
          ,
          <article-title>A survey on skill identification from online job ads</article-title>
          ,
          <source>IEEE Access 9</source>
          (
          <year>2021</year>
          )
          <fpage>118134</fpage>
          -
          <lpage>118153</lpage>
          . URL: http://dx.doi.org/10.1109/ ACCESS.
          <year>2021</year>
          .
          <volume>3106120</volume>
          . doi:
          <volume>10</volume>
          .1109/access.
          <year>2021</year>
          .
          <volume>3106120</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>International</given-names>
            <surname>Labour</surname>
          </string-name>
          <string-name>
            <surname>Ofice</surname>
          </string-name>
          ,
          <article-title>The International Standard Classification of Occupations (ISCO-08) Companion Guide</article-title>
          , Geneva,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Smedt</surname>
          </string-name>
          , M. le
          <string-name>
            <surname>Vrang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Papantoniou</surname>
          </string-name>
          , Esco:
          <article-title>Towards a semantic web for the european labor market</article-title>
          ,
          <source>in: LDOW@WWW</source>
          ,
          <year>2015</year>
          . URL: https://api. semanticscholar.org/CorpusID:14184714.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>National</given-names>
            <surname>Academies</surname>
          </string-name>
          Press,
          <year>2010</year>
          . URL: http://dx.doi. org/10.17226/12814. doi:
          <volume>10</volume>
          .17226/12814.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.-I.</given-names>
            <surname>Dascalu</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Marin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. V.</given-names>
            <surname>Nemoianu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.-F.</given-names>
            <surname>Puskás</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hang</surname>
          </string-name>
          ,
          <article-title>An ontology for educational and career profiling based on the romanian occupation classification framework: Description and scenarios of utilisation</article-title>
          ,
          <source>in: ICERI Proceedings, ICERI2022, IATED</source>
          ,
          <year>2022</year>
          . URL: http://dx.doi.org/10.21125/iceri.
          <year>2022</year>
          .
          <year>1881</year>
          . doi:
          <volume>10</volume>
          .21125/iceri.
          <year>2022</year>
          .
          <year>1881</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Skills</surname>
            <given-names>frameworks</given-names>
          </string-name>
          ,
          <year>2023</year>
          . URL: https://www. skillsfuture.gov.sg/skills-framework.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>[9] Singapore standard occupational classification ssoc</article-title>
          <year>2024</year>
          ,
          <year>2024</year>
          . URL: https://www.singstat.gov.sg/ standards/standards-and
          <article-title>-classifications/ssoc.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Joyez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lafineur</surname>
          </string-name>
          ,
          <article-title>The occupation space: network structure, centrality and the potential of labor mobility in the french labor market</article-title>
          ,
          <source>Applied Network Science</source>
          <volume>7</volume>
          (
          <year>2022</year>
          ). URL: http://dx. doi.org/10.1007/s41109-022-00453-3. doi:
          <volume>10</volume>
          .1007/ s41109-022-00453-3.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Alabdulkareem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>AlShebli</surname>
          </string-name>
          , C. Hidalgo,
          <string-name>
            <surname>I. Rahwan</surname>
          </string-name>
          ,
          <article-title>Unpacking the polarization of workplace skills</article-title>
          ,
          <source>Science Advances</source>
          <volume>4</volume>
          (
          <year>2018</year>
          ). URL: http://dx.doi.org/10.1126/sciadv.aao6030. doi:
          <volume>10</volume>
          .1126/sciadv.aao6030.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fettach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ghogho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Benatallah</surname>
          </string-name>
          ,
          <article-title>Knowledge graphs in education and employability: A survey on applications and techniques</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>80174</fpage>
          -
          <lpage>80183</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2022</year>
          .
          <volume>3194063</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vista</surname>
          </string-name>
          ,
          <article-title>Data-driven identification of skills for the future: 21st-century skills for the 21st-century workforce</article-title>
          ,
          <source>SAGE Open 10</source>
          (
          <year>2020</year>
          )
          <article-title>215824402091590</article-title>
          . URL: http://dx.doi.org/10.1177/2158244020915904. doi:
          <volume>10</volume>
          . 1177/2158244020915904.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Knowledge graph refinement: A survey of approaches and evaluation methods</article-title>
          ,
          <source>Semantic Web</source>
          <volume>8</volume>
          (
          <year>2016</year>
          )
          <fpage>489</fpage>
          -
          <lpage>508</lpage>
          . URL: http://dx.doi.org/10.3233/ SW-160218. doi:
          <volume>10</volume>
          .3233/sw-160218.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>M. de Groot</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Schutte</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Graus</surname>
          </string-name>
          ,
          <article-title>Job postingenriched knowledge graph for skills-based matching</article-title>
          ,
          <source>ArXiv abs/2109</source>
          .02554 (
          <year>2021</year>
          ). URL: https://api. semanticscholar.org/CorpusID:237420688.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          , E. Cambria,
          <string-name>
            <given-names>P.</given-names>
            <surname>Marttinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>A survey on knowledge graphs: Representation, acquisition, and applications</article-title>
          ,
          <source>IEEE Transactions on Neural Networks and Learning Systems</source>
          <volume>33</volume>
          (
          <year>2022</year>
          )
          <fpage>494</fpage>
          -
          <lpage>514</lpage>
          . URL: http://dx.doi.org/10.1109/TNNLS.
          <year>2021</year>
          .
          <volume>3070843</volume>
          . doi:
          <volume>10</volume>
          .1109/tnnls.
          <year>2021</year>
          .
          <volume>3070843</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          , E. Cambria,
          <string-name>
            <given-names>P.</given-names>
            <surname>Marttinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>A survey on knowledge graphs: Representation, acquisition, and applications</article-title>
          ,
          <source>IEEE Transactions on Neural Networks and Learning Systems</source>
          <volume>33</volume>
          (
          <year>2022</year>
          )
          <fpage>494</fpage>
          -
          <lpage>514</lpage>
          . doi:
          <volume>10</volume>
          .1109/TNNLS.
          <year>2021</year>
          .
          <volume>3070843</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <article-title>A survey on application of knowledge graph</article-title>
          ,
          <source>Journal of Physics: Conference Series</source>
          <volume>1487</volume>
          (
          <year>2020</year>
          )
          <article-title>012016</article-title>
          . URL: http://dx.doi.org/10.1088/
          <fpage>1742</fpage>
          -6596/1487/1/012016. doi:
          <volume>10</volume>
          .1088/
          <fpage>1742</fpage>
          -6596/ 1487/1/012016.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jensen</surname>
          </string-name>
          , R. van der Goot,
          <string-name>
            <given-names>B.</given-names>
            <surname>Plank</surname>
          </string-name>
          ,
          <article-title>Skill extraction from job postings using weak supervision (</article-title>
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2209.08071.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fareri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Melluso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chiarello</surname>
          </string-name>
          , G. Fantoni,
          <article-title>Skillner: Mining and mapping soft skills from any text</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>184</volume>
          (
          <year>2021</year>
          )
          <article-title>115544</article-title>
          . URL: http://dx.doi.org/10.1016/j.eswa.
          <year>2021</year>
          .
          <volume>115544</volume>
          . doi:
          <volume>10</volume>
          . 1016/j.eswa.
          <year>2021</year>
          .
          <volume>115544</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>N. H. N.</given-names>
            <surname>Minh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. K.</given-names>
            <surname>Doan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. Q.</given-names>
            <surname>Huy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. X.</given-names>
            <surname>Loc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. N.</given-names>
            <surname>Vu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. C.</given-names>
            <surname>Phap</surname>
          </string-name>
          ,
          <article-title>Information Technology Skills Extractor for Job Descriptions in vku-ITSkills Dataset Using Natural Language Processing</article-title>
          , Springer Nature Switzerland,
          <year>2023</year>
          , p.
          <fpage>250</fpage>
          -
          <lpage>261</lpage>
          . URL: http://dx. doi.org/10.1007/978-3-
          <fpage>031</fpage>
          -36886-8_
          <fpage>21</fpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>031</fpage>
          -36886-8_
          <fpage>21</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Proceedings of the 31st International Conference on Neural Information Processing Systems</source>
          , NIPS'17, Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <year>2017</year>
          , p.
          <fpage>6000</fpage>
          -
          <lpage>6010</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>My</surname>
            <given-names>skillsfuture portal</given-names>
          </string-name>
          ,
          <year>2024</year>
          . URL: https://www. myskillsfuture.gov.sg/.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Minaee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nikzad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chenaghlu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Amatriain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <source>Large language models: A survey</source>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2402.06196. doi:
          <volume>10</volume>
          .48550/ARXIV.2402.06196.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Searching for best practices in retrieval-augmented generation</article-title>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2407.01219. doi:
          <volume>10</volume>
          .48550/ ARXIV.2407.01219.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26] Gpt-4
          <source>turbo (gpt-4o)</source>
          ,
          <year>2024</year>
          . URL: https://platform. openai.com/docs/models/gpt-4o.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Grag: Graph retrieval-augmented generation</article-title>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2405</volume>
          .
          <fpage>16506</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Carta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Giuliani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Piano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Podda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pompianu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. G.</given-names>
            <surname>Tiddia</surname>
          </string-name>
          ,
          <article-title>Iterative zero-shot llm prompting for knowledge graph construction</article-title>
          ,
          <source>arXiv preprint arXiv:2307.01128</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chepurova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bulatov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kuratov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Burtsev</surname>
          </string-name>
          ,
          <article-title>Better together: Enhancing generative knowledge graph completion with language models and neighborhood information, in: Findings of the Association for Computational Linguistics: EMNLP 2023, Association for Computational Linguistics</article-title>
          ,
          <year>2023</year>
          . URL: http://dx. doi.org/10.18653/v1/
          <year>2023</year>
          .findings-emnlp.
          <volume>352</volume>
          . doi:
          <volume>10</volume>
          . 18653/v1/
          <year>2023</year>
          .findings-emnlp.
          <volume>352</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>W.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Q. Li,</surname>
          </string-name>
          <article-title>Graph machine learning in the era of large language models (llms</article-title>
          ),
          <year>2024</year>
          . URL: https://arxiv.org/abs/2404.14928. doi:
          <volume>10</volume>
          .48550/ARXIV.2404.14928.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>