<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>European Workshop on Algorithmic Fairness, July</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Building Job Seekers' Profiles: Can LLMs Level the Playing Field?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Susana Lavado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leid Zejnilovic</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nova School of Business and Economics</institution>
          ,
          <addr-line>Lisbon</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>0</volume>
      <fpage>1</fpage>
      <lpage>03</lpage>
      <abstract>
        <p>This study investigates the impact of language complexity on the performance of an NLP-based recommender system that assists job seekers in adding relevant occupation labels and skills to their profiles. The system, deployed by Job Market Finland (JMF), was evaluated to determine whether it biases its recommendations towards more complex language inputs, potentially disadvantaging users who employ simpler language. Additionally, the study explores the efectiveness of using large language models (LLMs) to enhance simpler descriptions and mitigate potential biases. By utilizing a stratified sample of occupations and crafting varied descriptions (original, simple, complex, and LLM-improved), we analyzed the system's recommendations against a ground truth. Results indicate that the system favored more complex language, improving occupation label suggestions (but not skill recommendations). This bias is not mitigated by the use of an LLM, suggesting potential unintended consequences for users who employ simpler language and highlighting the opacity in optimizing such systems.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Job matching</kwd>
        <kwd>Large language models</kwd>
        <kwd>Natural language processing</kwd>
        <kwd>Algorithmic bias</kwd>
        <kwd>Human-machine interaction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The usage of systems relying on natural language processing (NLP) techniques, especially those
systems using large language models (LLMs), is swiftly increasing across diverse tasks [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
However, due to their black-box nature [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], understanding how to maximize these models’
utility poses considerable challenges to users [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The complexity of optimizing the performance
of NLP systems has sparked discussions regarding the importance of developing domain-specific
prompt engineering skills [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ]. When opaque NLP systems are available to casual users
without clear information on how to optimally use them, biased outputs can emerge [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>As NLP-based recommender systems (including LLMs) are more often an external face of
organizations interacting with citizens, decision-makers are confronted with a trade-of between
scalability and eficiency of services on one side and potential biases on the other. The question
of biases is complex, as they emerge from the interaction between the technological artifact (a
model) and a human, and are context dependent. In this study, we explore a type of bias that
may emerge due to the form (rather than the content) of the interaction between a model and a
human. More precisely, we examined the potential bias present in an NLP-based recommender
system designed to assist job seekers in adding skills and occupation labels to their profiles,
which are subsequently used to match them to relevant job ofers.</p>
      <p>
        Humans may have diferent skills to craft the prompts with which they interact with the
system and the system may hypothetically produce more relevant output to better crafted
prompts. As the quality of the prompts is related to individual skills, or education, the system
may introduce unintended bias that favor more skilled people. A way to address such bias
may lie within the NLP tools themselves. Recent evidence suggests that LLMs might be more
efective when used by individuals with lower skills, hinting at a potential for these systems
to "level the playing field" [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We investigated whether feeding a simpler description to an
LLM before submitting it to the recommender system would reduce unintended bias. Hence,
we hypothesized that:
1. An NLP system that recommends skills based on candidates’ job descriptions will more
accurate when more complex vs. simpler language is used.
2. Improving of the candidates’ profile using an LLM, before inputting it to recommender
system would eliminate the complex language bias.
      </p>
      <sec id="sec-1-1">
        <title>1.1. Context</title>
        <p>This study investigated potential language bias of an NLP tool deployed at Job Market Finland
(JMF), the front-page of the Finnish PES e-services1. JMF ofers job seekers the possibility of
creating a profile to be shown to potential employers, which can then contact the job seekers.
Besides adding a description of their job applicant profile in natural language, job seekers may
select the relevant job occupation label and skills. These occupations labels and skills are defined
the multilingual classification of European Skills, Competences, Qualifications and Occupations
(ESCO)2. However, this is arguably a daunting task, considering that ESCO contains more than
3 thousand occupations and 13 thousand skills. To assist job seekers, JMF ofers a NLP tool,
based on Word2Vec, that recommend job occupation labels and skills to be added to their profile.
These job occupation labels and skills will then be used to match the job seeker to the available
job opportunities, which are also tagged with occupation(s) label(s) and relevant skills.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Method</title>
      <p>To create employees’ profiles, we randomly selected 63 occupations (corresponding to 2% of the
3008 ESCO occupations), stratified by the highest hierarchy level of ESCO (the 10 job groups).
Because ESCO descriptions are associated to occupations’ labels and a skill set, using them
as employees’ profiles gave us a ground truth to which we could compare the performance
of the system. From the ESCO occupations’ descriptions, we removed the occupation tag and
shifted the language from the third person singular to the first person singular (e.g., we replaced
"Dietitians create dietary plans" by "I create dietary plans"). Then, we crafted two alternative
descriptions by keeping the original content but either using simpler or more complex language
1https://tyomarkkinatori.fi/en
2https://esco.ec.europa.eu/en/about-esco/what-esco
(i.e, using vocabulary and sentence structure that is easily understandable to a wide audience
vs. using advanced vocabulary and nuanced sentence constructions). This process involved
leveraging a language model (ChatGPT 3.5). To simplify the descriptions, we opened a new
chat for each occupation, where we entered the following prompt "Can you please help me
simplify some sentences?" and then proceeded to introduce each sentence of the occupation’s
description, one by one. We revised each of the simplified sentence to guarantee they kept the
same content, and asked the LLM for further simplification if needed. For the more complex
descriptions, we followed the same procedure, but replaced the prompt for "Can you please
help me rewrite some sentences using more polished/complex language?" To assert the validity
of our created materials, we computed embedding vectors of the descriptions using
sentenceBERT, a Transformers model, and computed the Cosine similarity between the vectors. To
verify whether LLM-improved descriptions would achieve better results than the original, we
crafted an additional occupation description by inputting the simplified description into another
LLM (Gemini), preceded by the following prompt: "I am creating my candidate profile in a
website where employers can find my profile and contact me to ofer me a job interview. I need
help to polish my description. Can you help me improve my paragraph? Here is my original
description:." We then verbatim copied the paragraph improved by Gemini, but removed any
suggested occupation by Gemini by either deleting it or replacing it with the word "professional",
like we had done in the remaining descriptions. The job descriptions used can be obtained from
the authors upon request.</p>
      <p>We then inserted the original, simple, complex, and LLM-improved descriptions into the JMF
profile tool. The profile tool recommends seven occupation labels and 20 skills to be added
to any description inputted by the user. We compared the tool’s output with the occupation
labels and skills associated with the ESCO occupation descriptions to compute the following
dependent variables for each occupation:
1. Occupation position: The position in which the tool recommended the correct
occupation label (8 if not recommended);
2. Skills hit: The number of matches between the set of skills suggested by the tool
and by ESCO.</p>
      <p>To assess the impact of the diferent descriptions on the performance of the JMF tool across
our two dependent variables, we conducted a Repeated-measures ANOVA with pairwise
comparisons with a Bonferroni correction. We considered a significance level of  = 0.05. Figure 1
presents an overview of the research methodology.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <sec id="sec-3-1">
        <title>3.1. Similarity of the occupation descriptions</title>
        <p>Table 1 presents a comparison of the average number of sentences, words and characters in the
original, simple, complex and LLM-improved descriptions.</p>
        <p>We computed the cosine similarity between embedding vectors of the diferent descriptions
to validate that, despite using diferent words, the original, simple, and complex descriptions
maintained the same content. There were no diferences between the similarity of the original
and the simple descriptions (M = 0.84, SD = 0.06) and the similarity of the original and the
complex descriptions (M = 0.85, SD = 0.06). The simple and complex descriptions had, on
average, a distance of 0.75 (SD = 0.08). There were similar distances between the LLM-improved
description and the original (M = 0.71, SD = 0.10), the complex (M = 0.71, SD = 0.09) and the
simple (M = 0.68, SD = 0.09) descriptions.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Occupation position</title>
        <p>The analysis revealed significant diferences among the descriptions levels for the position the
correct occupation label was suggested , F(3, 186) = 25.26, p &lt; 0.001). Pairwise comparisons with
a Bonferroni correction revealed that the description improved by the LLM performed the worst
(M = 4.52, SD = 2.82), suggesting the correct occupation significantly later than the original
description (M = 1.79, SD = 1.76, corrected-p &lt; 0.001) and the complex description (M = 2.70, SD
= 2.28, corrected-p &lt; 0.001). No diferences were found between the occupation position for
the simple (M = 3.49, SD = 2.51) and the LLM-improved description (corrected-p = 0.100). The
complex description performed better than the simple description (corrected-p = 0.038), but
worse than the original description (p = 0.001).</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Skills hit</title>
        <p>The pattern of the means for each of the levels generally mimic the results obtained for the
occupation position variable: LLM-improved description: M = 4.92, SD = 4.41; Simple description:
M = 5.02, SD = 4.12; Original description: M = 5.84, SD = 4.25; Complex description: M = 5.67,
SD = 4.24. However, pairwise comparisons with a Bonferroni correction revealed no significant
diferences in the number of skills correctly recommended by JMF using the diferent description,
F(3, 186) = 2.30, p = 0.079.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
      <p>
        Results partially supported hypothesis 1. The NLP system exhibited unintended biases that
favored more complex language over simpler language in recommending occupation labels,
but this bias did not extend to skills recommendations. However, we found no support for
hypothesis 2. Results showed that improving the inputted description using an LLM did not
improve the performance of the system. While better results may have been found if more
time was invested in crafting the prompt asking the LLM to improve the description, we were
interested in the behavior of an average user, who may stop interacting with the LLM after it
provides a (seemingly) successful response [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        This study has implications in the context of ensuring fair access to employment opportunities
through digital platforms. Since the occupations and skills associated with a profile are crucial
for matching job seekers to relevant job opportunities, biases in NLP-based recommender
systems, specifically those favoring complex language, may inadvertently disadvantage job
seekers who use simpler language. This unequal performance could lead to less relevant job
matches for these individuals, potentially limiting their employment prospects. Additionally,
while LLMs have been presented as tools to potentially level the playing field [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], our findings
suggest that their use alone does not eliminate such biases. In fact, LLM-enhanced inputs did not
improve the recommender system’s performance, underscoring the need for more transparent
and carefully designed interventions. These findings highlight the opacity surrounding the
optimization of NLP-based systems like this one, which may pose challenges for users seeking
to maximize their utility.
      </p>
      <p>We acknowledge that our operationalization of language complexity in this study remains
underdeveloped. This was an exploratory study, and as such, we did not adopt a formal
framework for characterizing language complexity. Future work should address this limitation
by providing a more standardized measure of language complexity.</p>
      <p>The NLP-based recommender system analysed in this work was based on word2vec algorithm.
Future research should investigate whether systems leveraging advanced LLMs exhibit similar
biases, and examine the boundary conditions of these findings. Future research should also
explore the real-world impact of using LLMs in hiring processes, specifically by examining
whether LLM-enhanced descriptions result in increased engagement from potential employers.</p>
      <p>In conclusion, this study highlights important considerations for policymakers and developers
in designing AI-powered employment services that ensure all job seekers are equally empowered.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          , T. Han,
          <string-name>
            <surname>S</surname>
          </string-name>
          . Ma, J. Zhang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Qiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <article-title>Summary of ChatGPT-Related research and perspective towards the future of large language models</article-title>
          ,
          <source>Meta-Radiology</source>
          <volume>1</volume>
          (
          <year>2023</year>
          )
          <article-title>100017</article-title>
          . doi:https://doi.org/10.1016/j.metrad.
          <year>2023</year>
          .
          <volume>100017</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Deciphering the Enigma: A Deep Dive into Understanding and Interpreting LLM Outputs, TechRxiv (</article-title>
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .36227/techrxiv.24085833.
          <year>v1</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Haag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Kruse</surname>
          </string-name>
          ,
          <article-title>Negotiating with llms: Prompt hacks, skill gaps, and reasoning deficits</article-title>
          ,
          <source>arXiv preprint arXiv:2312.03720</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Heston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Khun</surname>
          </string-name>
          , Prompt Engineering in Medical Education,
          <source>International Medical Education</source>
          <volume>2</volume>
          (
          <year>2023</year>
          )
          <fpage>198</fpage>
          -
          <lpage>205</lpage>
          . doi:
          <volume>10</volume>
          .3390/ime2030019.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Giray</surname>
          </string-name>
          ,
          <article-title>Prompt Engineering with ChatGPT: A Guide for Academic Writers</article-title>
          ,
          <source>Annals of Biomedical Engineering</source>
          <volume>51</volume>
          (
          <year>2023</year>
          )
          <fpage>2629</fpage>
          -
          <lpage>2633</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10439-023-03272-4.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <article-title>Unleashing chatgpt's power: A case study on optimizing information retrieval in flipped classrooms via prompt engineering</article-title>
          ,
          <source>IEEE Transactions on Learning Technologies</source>
          <volume>17</volume>
          (
          <year>2024</year>
          )
          <fpage>629</fpage>
          -
          <lpage>641</lpage>
          . doi:
          <volume>10</volume>
          .1109/TLT.
          <year>2023</year>
          .
          <volume>3324714</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Noy</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. Zhang,</surname>
          </string-name>
          <article-title>Experimental evidence on the productivity efects of generative artificial intelligence</article-title>
          ,
          <source>Science</source>
          <volume>381</volume>
          (
          <year>2023</year>
          )
          <fpage>187</fpage>
          -
          <lpage>192</lpage>
          . doi:
          <volume>10</volume>
          .1126/science.adh2586.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zamfirescu-Pereira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Y.</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hartmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Why johnny can't prompt: How non-ai experts try (and fail) to design llm prompts</article-title>
          ,
          <source>in: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI '23</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .1145/3544548.3581388.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>