<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Annotation Imputation to Individualize Predictions: Initial Studies on Distribution Dynamics and Model Predictions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>London Lowmanstone</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ruyuan Wan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Risako Owan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaehyung Kim</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dongyeop Kang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Minnesota</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Notre Dame</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Annotating data via crowdsourcing is time-consuming and expensive. Due to these costs, dataset creators often have each annotator label only a small subset of the data. This leads to sparse datasets with examples that are marked by few annotators. The downside of this process is that if an annotator doesn't get to label a particular example, their perspective on it is missed. This is especially concerning for subjective NLP datasets where there is no single correct label: people may have diferent valid opinions. Thus, we propose using imputation methods to generate the opinions of all annotators for all examples, creating a dataset that does not leave out any annotator's view. We then train and prompt models, using data from the imputed dataset, to make predictions about the distribution of responses and individual annotations. In our analysis of the results, we found that the choice of imputation method significantly impacts soft label changes and distribution. While the imputation introduces noise in the prediction of the original dataset, it has shown potential in enhancing shots for prompts, particularly for low-response-rate annotators. We have made all of our code and data publicly available.1</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;natural language processing</kwd>
        <kwd>imputation</kwd>
        <kwd>matrix factorization</kwd>
        <kwd>content filtering</kwd>
        <kwd>large language models</kwd>
        <kwd>annotation</kwd>
        <kwd>NLPerspectives</kwd>
        <kwd>LeWiDi</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>tions for individual annotators as an imputation problem:
given a spreadsheet with rows corresponding to text and
Natural language processing (NLP) models rely on large columns corresponding to annotators, how would one
amounts of data that is expensive and time-consuming accurately fill in the spreadsheet in order to correctly
to label [1]. Crowdsourcing has emerged as a popular predict how each annotator will label each piece of text?
solution to this problem, but it comes with its own chal- Figure 1 visualizes this approach, which, ideally, enables
lenges, principal among them being annotator disagree- dataset creators to generate additional annotations
withment [2, 3]. Although there are many possible causes of out extensive crowdsourcing.
disagreement, the common causes are annotator subjec- We postulate that annotators who have historically
tive judgment and language ambiguity [4]. Not taking assigned the same labels to identical text segments may,
into account the inherent subjectiveness and ambiguity given similar contexts in unseen data, continue to
demonof some instances can lead to inaccurate predictions [5]. strate congruent labeling behavior. Thus, imputation
Thus, in recent years, researchers have begun to recog- methods, which take in data containing all of the dataset
nize the importance of disagreement, advancing models annotations, should be able to discover patterns to relate
and datasets that accurately reflect disagreement, rather annotators and annotations in order to make accurate
than ignoring it or working around it [6]. predictions as to how a particular annotator might
la</p>
      <p>In order for models to accurately reflect disagreement, bel a particular example, based on how other annotators
they must accurately model true human populations. labeled the same or similar examples.
Here, we frame the problem of making accurate predic- Matrix factorization techniques used in
recommendation systems and annotator-level models of disagreement
2nd Workshop on Perspectivist Approaches to NLP both make predictions about individual annotations made
* Corresponding author. by individual annotators. Thus, our analyses can be
ap† These authors contributed equally. plied to both types of models in order to reveal diferences
($R. lWowanm)a;0o1w6a@n0u0m2n@.eudmun(.Le.dLuo(wR.mOawnsatno)n;eja);erhwyuanng@knimd.@edkuaist.ac.kr between the original data and imputed data created by
(J. Kim); dongyeop@umn.edu (D. Kang) these models. In our work, we impute datasets by
uti0000-0002-5457-6553 (L. Lowmanstone) lizing two matrix factorization methods, kernel matrix
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License factorization and neural collaborative filtering, and a
su1 hCPWrEooUrctkReshtdoinpgpssIhStpN:/c1e:6u1r3-w/-0s.o7r3g/giACtthtEriubUubtRi.ocnoW4m.0oIn/rtmekrnsianhtinoonpeals(PoCCrtoaBnYce4l.p0e)/.dainnngost(aCtiEoUnR-i-mWpSu.otartgio)n pervised learning model (Multitask) proposed by [7], that</p>
      <p>Individual Imputed Annotation
b) Soft Label Analysis, which focuses on shifts
in the soft label after imputation compared
to the original data. We provide a
visualization technique for viewing how the soft
labels change after imputation.
c) Usage Analysis, which focuses on how
models perform after training on or
being prompted with imputed data. We show
that kernel matrix factorization, neural
collaborative filtering, and Multitask
imputation tend to harm the capabilities of
Multitask and GPT models to make
individual, soft-label, and aggregate predictions,
except in the case of using imputation to
increase the number of shots to prompts
for making individualized predictions for
low-response-rate annotators.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>models disagreement at the annotations level [8, 9, 7]. Disagreement in NLP Disagreement has been found
Through our analyses, we find that imputation greatly within NLP datasets for many years [11, 12, 13, 6].
Howtransforms the distribution of annotations (including low- ever, recently, there has been much work done on
deering the variance of the data) and creates noticeable veloping and evaluating models that model
disagreechanges in examples’ soft labels. ment within datasets, rather than ignore disagreement</p>
      <p>After imputing and analyzing the data, we use the im- [7, 14, 5, 15, 16].
puted datasets to train and prompt models that make In particular, the SemEval-2023 Learning with
Disindividualized predictions. For training, we use the Mul- agreements (LeWiDi) task invites competitors to create
titask model from [7] in order to make aggregate and models that predict soft labels of human disagreement
individualized predictions and find that training on im- for diferent text inputs [ 6]. While hard labels provide a
puted data harms prediction performance. For prompting, definitive categorization for data points, soft labels ofer
we use GPT-3 (text-davinci-003) and ChatGPT (3.5-turbo) a probabilistic interpretation, capturing the uncertainties
and provide the models with prompts containing either or nuances in classification. Multiple submissions for
imputed or non-imputed data to determine their impact this task used models proposed by [7] in order to make
on the models’ ability to make individualized and distri- predictions at the individual level. Success at the task
butional label predictions. We find that adding prompt was determined by micro F1 score on gold labels and
shots via imputation improves ChatGPT’s performance cross-entropy on soft-labels. Within the task, all dataset
for predicting annotations of low-response-rate annota- labels were binary, and no metric was used to measure
tors, but does not consistently improve other areas of success at the level of individual annotators.
prediction such as distributional label prediction, indi- The authors of [17] propose multiple diferent methods
vidualized prediction for high-response annotators, or for evaluating models that make individualized
predicmerely replacing human annotations with imputed data tions. Among these are Jensen-Shannon divergence, a
[10]. symmetric variation of Kullback-Leibler (KL) divergence
In summary, our primary contributions are: and cross-entropy. F1 score is also a proposed metric, but
only for aggregate labels, not individual labels.</p>
      <sec id="sec-2-1">
        <title>Another model for approaching disagreement is Jury</title>
        <p>Learning, where individuals’ annotations are modeled in
order to form “juries" of diferent demographics [ 14]. In
their paper, the authors analyze how using data generated
by “juries" afects the aggregate label, particularly in the
case of contentious texts [14].
1. Framing individualized prediction as an
imputation problem
2. Analysis techniques to compare imputed data to
real data:
a) Distribution Analysis, which focuses on
transformations of the underlying
distribution of annotations after imputation. We
show that diferent imputation methods
significantly change the underlying
annotation distributions.</p>
        <sec id="sec-2-1-1">
          <title>Collaborative Filtering in Recommendation Sys</title>
          <p>tems Similar to modeling disagreement, collaborative
Individual Imputed Annotation
Majority Vote for Original Data</p>
          <p>Majority Vote for Imputed Data</p>
          <p>Imputed vs Original Data
Individual Classifier</p>
          <p>Aggregated Prediction
Imputed Training</p>
          <p>Sentences_i
Shared Pre-trained</p>
          <p>Layers</p>
          <p>Imputed Prompting
Prompt:
The particular annotator's response to
other examples: 1
Other annotators labeled the target
example: 75% 0 and 25% 1 0 0 0 1
How would the particular annotator
annotatethe target example?</p>
          <p>Sentences_i
Generated Results from GPT
ifltering systems also create individualized predictions Kernel matrix factorization relies on kernels to project
of human behavior in order to make relevant recommen- data to a higher dimensional space where more complex
dations. Contrary to disagreement models in natural patterns can be found in order to generate a matrix
factorlanguage processing, these systems are entirely depen- ization which is used for imputation. NCF matrix
factordent on user-provided annotations and lack the ability ization relies on neural networks rather than kernels to
to predict the reactions of new users to unseen text. compute a matrix factorization of the data, and Multitask</p>
          <p>When evaluating performance of collaborative filter- relies purely on neural networks to make individualized
ing systems, metrics are generally focused on accuracy predictions. All methods employ a core process:
identifyof predictions, rather than quantifying and visualizing ing patterns between annotators and annotations across
changes in the distribution of data [18, 19]. These metrics the dataset.
provide good signals for the success of a model, but do For our experiments, kernel matrix factorization is
not help with understanding how models modify data implemented primarily using of-the-shelf code [ 20]. In
when they do not match the original data. addition, we add a grid search component which
determines the best model hyperparameters by holding out 5%
of the given training data as validation data, and choosing
3. Annotation Imputation For the hyperparameters that resulted in the lowest RMSE
Individualized Predictions score on the validation data. See Appendix C for details.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Neural collaborative filtering was implemented based</title>
        <p>First, we compared how various imputation methods han- on the work of [9]. The details of our implementation
dle and fill in the missing annotations. We then trained can be found via our code. For this model, we also use
supervised models and used GPT-based prompting to an additional grid search component which determines
evaluate imputation’s impact on aggregate and individu- the best model hyperparameters. However, we choose
alized prediction. the hyperparameters for this model based on the lowest</p>
      </sec>
      <sec id="sec-2-3">
        <title>RMSE score when evaluated on all training examples,</title>
        <p>3.1. Annotation Imputation rather than a held-out validation set. See Appendix C for
details.</p>
      </sec>
      <sec id="sec-2-4">
        <title>In order to understand how individualized prediction</title>
        <p>afects data, we use three diferent methods: kernel
matrix factorization, neural collaborative filtering, and a
Multitask supervised neural network model from [7].1</p>
      </sec>
      <sec id="sec-2-5">
        <title>1The hyperparameters used for each of the models can be found in</title>
        <p>Appendix C.
3.2. Imputed Training
In this stage, we use the Multitask model from [7] on
both original and imputed data and compare the
evaluation results in order to understand how imputed data
impacts model training. We follow a similar setup to
[7] by using 5-fold validation and averaging the results
across the folds [7]. However, in order to account for
dataset imbalance in our datasets, we report weighted F1
scores, rather than macro F1 scores. Note that the data
from each validation fold is hidden from the imputer, so
as not to cause data leakage. Details of the model’s
architecture can be found in Appendix A, and hyperparameter
details can be found in Appendix C. The same model is
used both for imputation and training (see Section 4).
3.3. Imputed Prompting</p>
        <p>In each of these conditions, outputs are considered
correct if, after removing whitespace, they only contain
the correct label. We conducted initial studies to discard
particularly low-performing skeletons and infills. The
remaining skeletons and infills are used for all conditions.
(Details are provided in our code.) We then measure
success of a condition based on the highest weighted F1 score
achieved by a prompt skeleton within that condition.</p>
        <p>We also conducted three key experiments using GPT-3 4. Experiments
(text-davinci-003) and ChatGPT (3.5-turbo) to better
understand the impact of imputation on predictions made Our experiments involve: (1) comparing imputed and
by GPT-based models [10]: The first experiment tests the original data, (2) conducting training using imputed data,
impact of using imputed data when making individual- and (3) prompting generation based on imputed data, all
ized predictions for low-response-rate annotators. The illustrated in Figure 2.
second experiment makes individualized predictions for
all annotators (not just low-response-rate annotators), Datasets In order to ensure a diversity of data, we
but also adds original distribution information, imputed utilize six diferent datasets in our analysis: Social
Chemdistribution information, or the original majority-voted istry (SChem) [21], Social Bias Inference Corpus (SBIC)
label near the end of the prompt in addition to the in- [22], Gab Hatespeech Corpus (GHC) [23, 24], Sentiment
cluded individual examples to quantify the impact of the dataset [25], and Politeness dataset [26]. Additionally,
extra information on predictions. The third tests indi- we isolate examples from the SChem dataset that were
lavidualized predictions when either original or imputed belled by 5 annotators in order to form the SChem5Labels
data from three distinct annotators is provided in the dataset. Our datasets are summarized in Table 1, and
prompt. Of these, imputation only had a positive impact more details can be found in Appendix B.
on making individualized predictions for
low-responserate annotators; the other two experiments are included
in Appendix D. 4.1. Imputed vs Original Data</p>
      </sec>
      <sec id="sec-2-6">
        <title>For all experiments, we create prompt skeletons, which are then filled in with data and/or text, depending on the experiment run (see Appendix F). This enables us to understand the influence of diferent prompts and data.</title>
        <sec id="sec-2-6-1">
          <title>Individualized Predictions for Low-Response-Rate</title>
          <p>Annotators In this experiment, we first isolated from
each dataset the 30 annotators with the lowest number of
annotations in the dataset. We then generated a prompt
for each of those 30 annotators. Each prompt consists of
at most 30 sentences and annotations from that annotator
(if there were more, we discarded the extras and chose
one to hold out, and if there were less, we included all but
one to hold out). Following the real examples, we also
included an additional 30 examples whose sentences are
from the dataset (and difer from the previous 30
examples and the held-out example), but whose annotations
are imputed via NCF. The final section of the prompt
then asks ChatGPT to predict the annotator’s annotation
on the held-out example.</p>
        </sec>
      </sec>
      <sec id="sec-2-7">
        <title>In the experiment, we test for diferences between three diferent conditions:</title>
      </sec>
      <sec id="sec-2-8">
        <title>1. Including both the original and imputed data</title>
        <p>Imputation We impute each of the datasets with each
of the imputation methods. However, in order to judge
which methods have the best performance, we also test
imputing the data while withholding 5% of the
annotations for evaluation. Withheld data is chosen in a manner
that reduces duplicate examples and annotators within
the withheld data in order to provide a more diverse test
set (details can be found in our code).</p>
        <p>Table 2 summarizes the RMSE score for each of the
methods on each of the datasets when evaluated on the
withheld data. Note that the Politeness dataset collects
labels ranging from 1 to 25, implying a broader variance
compared to other datasets. Consequently, RMSE values
are expected to be higher for the Politeness dataset. We
also find that while Multitask and NCF perform best on
diferent datasets, kernel matrix factorization is never
the best method, and is in fact always dominated by the</p>
      </sec>
      <sec id="sec-2-9">
        <title>NCF method.</title>
      </sec>
      <sec id="sec-2-10">
        <title>After the data is imputed, we use two analyses in order to better understand how imputed data difers from original data.</title>
        <p># instances # annotators # annotation</p>
        <p>SChem
SChem5Labels</p>
        <p>SBIC</p>
        <p>GHC
Sentiment
Politeness
400
8007
45223
27538
14070
4338
100
102
304
18
1481
219
50
5
3
3-4
4-5
5</p>
        <p>No one believes (0), occasionally believed (1),
controversial (2), common belief (3), universally true (4)</p>
        <p>Not ofensive (0), maybe (0.5), ofensive (1)</p>
        <p>Not hate speech (0), hate speech (1)
Very negative (-2), somewhat negative (-1), neutral (0),
somewhat positive (1), very positive (2)
A scale from polite (1) to impolite (25).</p>
        <p>Distributional Analysis The first analysis (distribu- cause significant changes to the distribution of the data as
tion analysis) applies principal component analysis (PCA) shown in Figure 3.2 Each imputation method generates
to both imputed and original data to visualize shifts in the an extremely diferent underlying distribution for the
distribution of example ratings. In order to apply PCA, annotations.
we represent each text as a vector of its annotations, In addition, we compute how the variance and
diswhere missing annotations are filled in with a value of 10, agreement rate change after imputation with NCF matrix
which is far outside the range of valid annotation labels factorization. Our results are compiled in Table 3, and we
for these datasets [27]. We also calculate the change in also provide Figure 4 to display the results on the SChem
variance between imputed and original data, and graph dataset. Results from other methods can be found in
Apthis variance against the disagreement rate across exam- pendix H. Across all datasets, we find that imputation
ples. The disagreement rate is computed as the number of decreases variance, indicating that NCF matrix
factorizaannotations that disagree with the majority-voted label tion does not accurately model the diversity of human
for that example, divided by the total number of anno- annotations. We can observe this lowered variance in
tations for that example. The majority-voted label for both Figure 3 and Figure 4 by comparing the scale of
imputed data is computed on the imputed data. the plots in the PCA visualization and by comparing the</p>
        <p>When we project the annotations to two dimensions heights of the points in the variance plot. We also find
using PCA, we find that diferent imputation methods</p>
      </sec>
      <sec id="sec-2-11">
        <title>2Other datasets’ results can be found in Appendix G.</title>
        <p>Original SBIC</p>
        <p>Kernel Imputed SBIC</p>
        <p>Multitask Imputed SBIC
NCF Imputed SBIC
300
200
100
0
-100
-200
50
40
30
20
10
0
-10
300
150
100
50
0
-50
-100
-0.096
-0.061
-0.044
0.044
-0.004
-0.006
merically quantify the diference between distributions.</p>
        <p>Through our soft label analysis, we find that diferent
imputation methods lead to varying changes of the soft
label of examples after imputation. Figure 5 demonstrates
how imputation changes the distribution of the data for
a given example and allows one to directly compare
different imputation methods to see how they modify the
data. In the case of Figure 5, we see an example from the</p>
      </sec>
      <sec id="sec-2-12">
        <title>SChem dataset which shows that kernel matrix factoriza</title>
        <p>tion predicted a much smaller proportion of annotators
to give the highest rating (in pink) than was in the
original dataset, while NCF matrix factorization predicted a
moderately larger proportion of users to give the
secondhighest rating (in blue) than the original dataset.</p>
      </sec>
      <sec id="sec-2-13">
        <title>Since we are interested in understanding how these</title>
        <p>soft labels difer from the original data, we also compute
Figure 4: A graph displaying how the variance has decreased the KL divergence between the imputed data and the
origafter using NCF matrix factorization. Each point represents an inal data. Note that if one would like a symmetric metric,
example. Variation is across annotations for that example, and the Jensen-Shannon divergence could be computed here
wdiistahgtrheeemmeanjotrriatyte-viosttehdeapnenrocteanttiaogne. of people who disagree as well [17]. In the particular example in Figure 5, we see
that the KL divergence score on Example 97 for Kernel
is 0.105, compared to the 0.123 divergence score by NCF.</p>
      </sec>
      <sec id="sec-2-14">
        <title>We provide a selection of multiple examples in Appendix</title>
        <p>that NCF matrix factorization tends to, but does not al- I. We also provide the average and standard deviation
ways, lead to more agreement with the majority-voted of KL divergence from the original data for each dataset
annotations. and each imputation method in Table 4. Overall, we see</p>
        <p>Overall, the chosen method for individualized predic- that NCF matrix factorization tends to best preserve the
tion has a large impact on the structure underlying the soft label of the original dataset when compared to
kerpredictions, even within the same dataset. We also find nel matrix factorization and the multitask model, as it
imputers can lower variance and raise agreement within is always either best or second-best. However,
perforthe dataset, demonstrating that imputation models may mance is dataset-dependent, and kernel and multitask
not always capture the diversity and disagreement of real achieve the best fidelity to the original soft labels for the
human annotators. Sentiment and SBIC datasets, respectively.</p>
      </sec>
      <sec id="sec-2-15">
        <title>Overall, soft labels do not remain consistent through</title>
        <p>Soft Label Analysis The second analysis visualizes imputation, and some methods of individualized
predicdiferences between soft labels of examples between the tion may tend to better preserve soft labels than others.
original and imputed data. To create the visualization, In our case, NCF matrix factorization best preserved the
we assign each label to a color and then generate hori- soft labels.
zontal bars of equal size where the proportion of the bar
containing that color corresponds to the proportion of 4.2. Imputed Training
annotations with that label. This enables us to directly
compare how diferent imputation methods alter the soft For imputed training, we train the Multitask model on
label distribution. Similar to [17], we also calculate the the original data, data imputed by NCF, and data imputed
Kullback-Liebler (KL) divergence between the original by a separate Multitask model. (Since RMSE scores from
distribution of data and the imputed data in order to nu- kernel matrix factorization are worse than NCF on each
dataset, we omit kernel matrix factorization from this of performance, and diferent datasets observed diferent
experiment.) After training the Multitask model on the results. Generally, using the original data resulted in the
original and imputed data, we report the average and best outcomes, followed by using the Multitask model to
standard deviation of the weighted F1 score for individ- impute the training data. Using NCF matrix factorization
ual and aggregate predictions over 5 folds using 5-fold to impute the data resulted in the lowest performance.
validation in Table 5 and Table 6. We observe that train- When we break out the model’s performance to
exing on imputed data from NCF and Multitask results in amine success on examples with difering levels of
disperformance worse than if we had used just the original agreement (Table 7), we see that the model tends to
perdata. This indicates that the predictions made by each form much better on examples with higher agreement
of the methods biases the data in a way that does not among annotators. We also see that the drop in
permatch the true predictions that the annotators would formance from imputing data is fairly consistent across
have made. disagreement levels, except for the GHC dataset anomaly</p>
        <p>However, not all prediction models had the same level on low disagreement examples, where imputation helped
or less), using only imputed data increases performance.</p>
        <p>We conjecture that since low-response-annotators in
these datasets generally have far less than 30 original
annotations, imputation enables us to provide more shots
to ChatGPT than the original dataset could provide, thus
enabling more accurate predictions than can be made
without imputation. While more data is needed to
determine why combining both imputed and original data
performs poorly, we provide supporting experiments in</p>
      </sec>
      <sec id="sec-2-16">
        <title>Appendix E to demonstrate that the performance improvement from using imputed data is particular to lowresponse-rate annotators and is caused by the imputed data, not the prompt text.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Discussion</title>
      <p>NCF 0.22± 0.007 0.912± 0.010 0.611± 0.91 Our analyses shed light on the impact of various
impuMultitask 0.34± 0.016 0.915± 0.009 0.611± 0.91 tation methods on the structure, soft label, and
training/prompting viability of imputed data in the context
of NLP annotation tasks in comparison to purely
humanTable 6 labeled data. We demonstrate that diferent imputation
Average weighted F1 score of aggregate (majority-voted) pre- methods can lead to significantly diferent underlying
edricattieodnsbymeaidtheebryNthCeFMmualttriitxasfkacctloarsisziafiteiorntraoirnaedseopnardaatteaMgeunl-- distributions of the data, which can, in turn, afect the
titask model. All the values and error bars are mean and performance of models trained on this data. Furthermore,
standard deviation across five folds. The best and the sec- while imputation can introduce noise, diminishing the
ond best results on each dataset are indicated in bold and accuracy of predictions for the original dataset, it is
essenunderline, respectively. tial to consider that the original dataset may not wholly
capture the full spectrum of reality due to the absence
of some annotator opinions. This has important
implislightly. How disagreement levels are computed is dis- cations for the design and evaluation of individualized
cussed in detail in Appendix J. prediction models in various applications, as well as for</p>
      <p>Overall, this indicates that diferent methods of indi- understanding and quantifying the biases that may be
vidualized predictions can introduce diferent biases into introduced by such models.
the data that cause methods trained on these predictions Each one of our analyses focuses on a particular area
to perform worse than if they had trained on just the of interest, which, together, help researchers and
pracoriginal data. Since we expect performance to increase titioners to better understand the predictions made by
with the amount of data provided, we conclude that these individualized prediction models. The distribution
analyparticular methods of individualized prediction likely in- sis provides information to those who are interested in
troduce strong biases that do not reflect reality [28]. ensuring that their model’s predictions match the
distribution of the original data and tools for analyzing
changes in disagreement and variation. For those who
4.3. Imputed Prompting are interested in soft labels, such as competitors in future
Here, we highlight the results of using imputed data LeWiDi tasks, our visualization helps with understanding
to improve individualized predictions on low-response- how models estimate the soft label and computational
rate annotators, as shown in Table 8. (As mentioned tools for determining which models mimic the original
above, other experiments are detailed in Appendix D) soft label best. As we see a rise in human-level predictions
From the data, we observe that using solely imputed data from systems, it is important to understand if models can
outperforms using original data or adding original data be trained or prompted with data created by
individualto the imputed data for all datasets except for Politeness. ized prediction models. We provide analyses from base</p>
      <p>Politeness is likely an outlier due to the high range of systems indicating how the chosen imputation method
potential labels in the Politeness dataset, leading the NCF may afect performance. Regardless of the scenario, our
method to impute labels that are unlikely to occur in the provided analyses enable researchers and practitioners
real dataset, causing ChatGPT to also predict unlikely who use models that make individualized predictions to
labels. However, when the amount of labels is smaller (5 better understand the diferences between their model’s
Politeness</p>
      <p>GHC
SChem</p>
      <sec id="sec-3-1">
        <title>Not Imputed</title>
      </sec>
      <sec id="sec-3-2">
        <title>Disagreement N</title>
        <p>Low 1306
Medium 2267
High 762
Low 20344
Medium 814
High 6392
Low 133
Medium 137
High 130</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>7. Limitations</title>
    </sec>
    <sec id="sec-5">
      <title>6. Future Work</title>
      <sec id="sec-5-1">
        <title>While we include two diferent matrix factorization meth</title>
        <p>ods from collaborative filtering, content-based
recommendation systems also provide individualized
predictions, so future work includes applying our methods to a
content-based recommendation system.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Also note that each of the methods we use is not state</title>
        <p>of-the-art in their respective field. We have chosen
baseline models for ease of implementation. Future work
includes running our methods on more advanced sys- 8. Conclusions
tems that may make more accurate predictions.</p>
        <p>We also have not conducted a user study to verify and We have proposed and utilized four diferent methods
quantify that our analysis methods help with understand- of understanding how the predictions made by
individing how predicted data difers from original data. Our ualized prediction models difer from the original data.
analysis here is based on the fact that previous methods We found that for kernel matrix factorization, and NCF
rely on aggregate metrics and do not provide fine-grained matrix factorization, the original soft label for the data
and comparative data between original and predicted shifts in diferent ways based on the method used, the
data. Conducting a user study would allow us to provide variance in labels is overall lowered, and training on data
explicit evidence of the exact amount of improvement created by these methods results in generally worse
preour methods provide in general for understanding how diction performance, while imputed data can be used to
individualized prediction impacts data. increase the number of shots in prompts.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Overall, we hope that our analysis methods for models</title>
        <p>that make individualized predictions are applied to future
models in order to help researchers and practitioners to
While our methods are extendable to any model that
makes individualized predictions, we only test our
methods on baseline models for both disagreement modeling
and collaborative filtering. Thus, when used on
state-ofthe-art methods, our methods may give very diferent
results. However, we still expect these methods to be
useful for understanding how imputation modifies the
underlying data, even if those modifications do not match
our results.
better understand how their models’ predictions difer
from real human annotators.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Ethics Statement</title>
      <sec id="sec-6-1">
        <title>Any methods which attempt to make individualized pre</title>
        <p>dictions carry the risk of learning how to replicate aspects
of individuals’ identities in order to make better
predictions. This may be viewed as data misuse, a violation of
privacy, or a violation of the right to be forgotten.</p>
        <p>Furthermore, there’s an inherent ethical challenge
in the goal of generating synthetic perspectives and
opinions. The ability to synthetically generate opinions
might inadvertently discourage practitioners from
seeking real human input. This poses two primary risks: 1)
it may lead to erroneous assumptions based on the
synthetic data rather than actual human sentiments, and</p>
      </sec>
      <sec id="sec-6-2">
        <title>2) it might marginalize authentic human participation, thereby weakening the quality and inclusivity of dataset and model development.</title>
      </sec>
      <sec id="sec-6-3">
        <title>While our methods are designed to help detect when</title>
        <p>models may be incorrectly predicting human behaviors,
they are most efective when applied to models
performing imputation. Thus, advocating for the success of this
work may inadvertently promote the creation and usage
of models with the ethical concerns described above.</p>
      </sec>
      <sec id="sec-6-4">
        <title>We urge creators of individualized prediction systems</title>
        <p>to always obtain consent from their users before applying
models to their data and to maintain open and
consistent communication about how their data may be used.</p>
      </sec>
      <sec id="sec-6-5">
        <title>We also advocate for a balanced approach, ensuring that while we progress in model development, real human perspectives remain at the core of our datasets and models.</title>
        <p>Language and Computation 6 (2008) 333–353. URL: Y. Choi, Social bias frames: Reasoning about
sohttps://doi.org/10.1007/s11168-008-9059-1. doi:10. cial and power implications of language, ArXiv
1007/s11168-008-9059-1. abs/1911.03891 (2020).
[13] R. Wan, K. Badillo-Urquiola, Dragonfly_captain [23] B. Kennedy, M. Atari, A. Davani, L. Yeh, A.
Omat SemEval-2023 task 11: Unpacking disagree- rani, Y. Kim, K. Coombs, S. Havaldar, G.
Portilloment with investigation of annotator demograph- Wightman, E. Gonzalez, et al., Introducing the
ics and task dificulty, in: Proceedings of the 17th gab hate corpus: Defining and applying hate-based
International Workshop on Semantic Evaluation rhetoric to social media posts at scale (2018).
(SemEval-2023), Association for Computational Lin- [24] B. Kennedy, M. Atari, A. M. Davani, L. Yeh, A.
Omguistics, Toronto, Canada, 2023, pp. 1978–1982. rani, Y. Kim, K. Coombs, S. Havaldar, G.
PortilloURL: https://aclanthology.org/2023.semeval-1.272. Wightman, E. Gonzalez, et al., Introducing the
doi:10.18653/v1/2023.semeval-1.272. gab hate corpus: defining and applying hate-based
[14] M. L. Gordon, M. S. Lam, J. S. Park, K. Patel, rhetoric to social media posts at scale, Language
J. T. Hancock, T. Hashimoto, M. S. Bernstein, Resources and Evaluation 56 (2022) 79–108.
Jury Learning: Integrating Dissenting Voices into [25] M. Diaz, I. L. Johnson, A. Lazar, A. M. Piper, D.
GerMachine Learning Models, 2022. URL: http:// gle, Addressing age-related bias in sentiment
analarxiv.org/abs/2202.02950. doi:10.1145/3491102. ysis, Proceedings of the 2018 CHI Conference on
3502004, arXiv:2202.02950 [cs]. Human Factors in Computing Systems (2018).
[15] M. L. Gordon, K. Zhou, K. Patel, T. Hashimoto, [26] C. Danescu-Niculescu-Mizil, M. Sudhof, D. Jurafsky,
M. S. Bernstein, The Disagreement Deconvolu- J. Leskovec, C. Potts, A computational approach
tion: Bringing Machine Learning Performance Met- to politeness with application to social factors, in:
rics In Line With Reality, in: Proceedings of the ACL, 2013.
2021 CHI Conference on Human Factors in Com- [27] B. Roy, All About Missing Data Handling.
puting Systems, ACM, Yokohama Japan, 2021, pp. Missing data is a every day problem. . . |
1–14. URL: https://dl.acm.org/doi/10.1145/3411764. by Baijayanta Roy | Towards Data Science,
3445423. doi:10.1145/3411764.3445423. 2019. URL: https://towardsdatascience.com/
[16] R. Wan, J. Kim, D. Kang, Everyone’s voice mat- all-about-missing-data-handling-b94b8b5d2184.
ters: Quantifying annotation disagreement us- [28] J. Kaplan, S. McCandlish, T. Henighan, T. B.
ing demographic information, arXiv preprint Brown, B. Chess, R. Child, S. Gray, A. Radford,
arXiv:2301.05036 (2023). J. Wu, D. Amodei, Scaling Laws for Neural
Lan[17] A. N. Uma, T. Fornaciari, D. Hovy, S. Paun, guage Models, 2020. URL: http://arxiv.org/abs/
B. Plank, M. Poesio, Learning from Disagree- 2001.08361. doi:10.48550/arXiv.2001.08361,
ment: A Survey, Journal of Artificial Intelligence arXiv:2001.08361 [cs, stat].</p>
        <p>Research 72 (2021) 1385–1470. URL: https://www. [29] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT:
jair.org/index.php/jair/article/view/12752. doi:10. Pre-training of Deep Bidirectional Transformers
1613/jair.1.12752. for Language Understanding, in: Proceedings
[18] J. L. Herlocker, J. A. Konstan, L. G. Terveen, J. T. of the 2019 Conference of the North American
Riedl, Evaluating collaborative filtering recom- Chapter of the Association for Computational
Linmender systems, ACM Transactions on Information guistics: Human Language Technologies, Volume
Systems 22 (2004) 5–53. URL: https://dl.acm.org/ 1 (Long and Short Papers), Association for
Comdoi/10.1145/963770.963772. doi:10.1145/963770. putational Linguistics, Minneapolis, Minnesota,
963772. 2019, pp. 4171–4186. URL: https://aclanthology.org/
[19] F. O. Isinkaye, Y. O. Folajimi, B. A. Ojokoh, Rec- N19-1423. doi:10.18653/v1/N19-1423.
ommendation systems: Principles, methods and [30] H. F. , bert-base-uncased · Hugging Face, 2023. URL:
evaluation, Egyptian Informatics Journal 16 https://huggingface.co/bert-base-uncased.
(2015) 261–273. URL: https://www.sciencedirect.
com/science/article/pii/S1110866515000341. doi:10.</p>
        <p>1016/j.eij.2015.06.005.
[20] Q.-V. Do, Matrix Factorization, 2022. URL: https:
//github.com/Quang-Vinh/matrix-factorization,
original-date: 2020-06-04T00:10:11Z.
[21] M. Forbes, J. D. Hwang, V. Shwartz, M. Sap, Y. Choi,</p>
      </sec>
      <sec id="sec-6-6">
        <title>Social chemistry 101: Learning to reason about social and moral norms, ArXiv abs/2011.00620 (2020). [22] M. Sap, S. Gabriel, L. Qin, D. Jurafsky, N. A. Smith,</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>A. Multitask Model Details</title>
      <sec id="sec-7-1">
        <title>The multitask model follows the specifications by [ 7].</title>
      </sec>
      <sec id="sec-7-2">
        <title>Specifically, let  be the Hugging Face “bert-base</title>
        <p>uncased" model, which takes in a text, , and outputs
the embedding of the [CLS] token for that text [29, 30].
Then, let  represent a linear layer which takes in
the embedding output by  and outputs  values,
where  is the number of valid annotation classes. We
have  of these linear layers, one for each annotator
. Finally, let  be a single-dimensional array whose
th entry is 1 if the corresponding annotation , is
not missing (is valid), and 0 if it is missing (is not valid).</p>
      </sec>
      <sec id="sec-7-3">
        <title>Finally, let  represent the cross entropy function of</title>
        <p>two vectors.</p>
        <p>Then, the output of the model  for a given text  is
computed as a single-dimensional array whose th value
is</p>
        <p>, =  ( ()).</p>
      </sec>
      <sec id="sec-7-4">
        <title>And the loss for the model is computed as</title>
        <p>( ⊙ , ).</p>
      </sec>
      <sec id="sec-7-5">
        <title>Exact implementation details can be found in our code.</title>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>B. Dataset Details</title>
      <p>C.2. Kernel Matrix Factorization
The hyperparameters for kernel matrix factorization are:
• Factors: [1, 2, 4, 8, 16, 32]
• Epochs: [1, 2, 4, 8, 16, 32, 64, 128, 256]
• Kernels: [linear, rbf, sigmoid]
• Gammas: (always set to auto)
• Regularization: [0.1, 0.01, 0.001]
• Learning Rate: [0.01, 0.001, 0.0001]
• Initial Mean: (always set to 0)
• Initial Standard Deviation: (always set to 0.1)
• Random Seed: [42 85]</p>
      <sec id="sec-8-1">
        <title>The hyperparameters used for each imputation task are picked automatically based on a randomly-chosen heldout validation set consisting of 5% of the training data.</title>
        <p>C.3. Multitask Model</p>
      </sec>
      <sec id="sec-8-2">
        <title>The hyperparameters for the Multitask model are: • Epochs: (always set to 10) • Learning rate: (always set to 5e-5)</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>D. Additional Imputed Prompting</title>
    </sec>
    <sec id="sec-10">
      <title>Experiments</title>
      <sec id="sec-10-1">
        <title>Each dataset consists of two files: a text and annotation</title>
        <p>ifle. The text file consists of  texts, such that  refers to
the th text, where 1 ≤  ≤  . The annotation file con- Method / Dataset GHC SBIC SChem
sists of annotations of text, and is a   matrix, where
, refers to the annotation given by the th annotator ONrCigF. DDiisstt.. 8833..3333%% 7766..6677%% 5500..0000%%
for the th text, where 1 ≤  ≤  . Maj. Voted 88.89% 83.33% 45.00%</p>
        <p>For all datasets, , is an integer rating of the text.</p>
        <p>While diferent datasets have upper bounds of potential
ratings, ratings which are numerically close to one an- Table 9
other signify annotations which are semantically close Accuracy of GPT-3 at making individualized predictions for
to one another. In other words, for the datasets we use, a given text when provided with 1. The original distribution
a rating of 1 is similar to a rating of 2 and less similar 2a.nTnhoteaNtioCnF-fiomr pthuattedtedxitstribution and 3. The majority-voted
to a rating of 5. This is in contrast to standard
classification tasks, where class labels may difer significantly in
semantics despite being close numerically.</p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>C. Hyperparameters</title>
      <p>C.1. NCF Matrix Factorization</p>
      <sec id="sec-11-1">
        <title>The hyperparameters for NCF matrix factorization are</title>
        <p>• Factors: [4, 8, 16, 32, 64, 128]
• Learning Rates: [0.001, 0.0005, 0.0001, 0.00005]</p>
      </sec>
      <sec id="sec-11-2">
        <title>Hyperparameters are picked automatically through a grid search of all possible values during each run based on whichever hyperparameters achieve the lowest RMSE score on all of the training data.</title>
        <p>Method / Dataset</p>
        <p>GHC</p>
        <p>SBIC
Not Imputed</p>
        <p>Imputed</p>
        <p>Not Imputed</p>
        <p>Imputed
original soft label 2. The imputed soft label or 3. The
majority-voted label. Note that we expect a lower
accuracy for SChem in comparison to SBIC or GHC since</p>
      </sec>
      <sec id="sec-11-3">
        <title>SChem has 5 labels, while SBIC and GHC have 3 and 2 labels respectively.</title>
      </sec>
      <sec id="sec-11-4">
        <title>Interestingly, there was no impact to accuracy based</title>
        <p>on whether or not imputed versus original data was used.</p>
      </sec>
      <sec id="sec-11-5">
        <title>While Section 4.1 clearly indicates diferences between imputed soft labels and original soft labels, GPT-3 appears to be robust to these diferences when making individualized predictions.</title>
        <p>We do see that providing the majority-voted
annotation rather than the soft label improves performance by
roughly 5% on GHC and 7% on SBIC. However, it also
drops performance on SChem by 5%. This appears to
indicate that for datasets with less labels, providing the
majority-voted label enables GPT-3 to make better
predictions than if one were to provide a soft label. However,
as the number of labels increases, soft labels may provide
more informative information for accurate predictions.</p>
      </sec>
      <sec id="sec-11-6">
        <title>In Table 10 we display the impact of imputed data</title>
        <p>on making individualized predictions for one of three
annotators whose data was provided in the prompt. The
data clearly shows that imputation has a negative impact
on ChatGPT’s ability to make accurate individualized
predictions.</p>
        <p>Similar results are shown in Table 11 where we display
the impact of imputed data on making soft label
predictions. A high KL divergence score indicates a worse
prediction; for GHC and SChem, imputation seems to harm
the predictions, whereas for SBIC, imputation seems to
help significantly. However, if we analyze the standard
deviation, we see that it is often near if not greater than
the mean, indicating a distribution that is skewed highly
to the right, and suggesting that any changes in
performance are not particularly significant.</p>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>E. Experiments to Support</title>
    </sec>
    <sec id="sec-13">
      <title>Low-Response Imputation</title>
      <p>Overall, based on the data we have compiled into Table
12, there is no clear pattern for annotators with high
response rate as to whether using imputed data rather
than real data is more beneficial for making
individualized predictions. We cannot test if this is the case on the
annotators with a low-response-rate, as they do not have
enough annotations to replace the imputed annotations.</p>
      <p>Table 13 indicates that swapping the prompts may
increase results in some cases, but, again, there’s is no
clear trend similar to the trend we saw for using imputed
data, which can be verified again in this data by noticing
that the imputed column consistently outperforms other
columns for all datasets but Politeness.</p>
      <sec id="sec-13-1">
        <title>Together, these two experiments show that the in</title>
        <p>crease in F1 score is not due to the text before the prompt,
and that it is the moderate increase in examples that
imputation can provide, rather than the imputed data itself,
that is likely the cause of the increased performance.</p>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>F. Imputed Prompting Prompt</title>
    </sec>
    <sec id="sec-15">
      <title>Details</title>
      <p>F.1. Description of Prompts
For the highlighted ChatGPT experiment and ablation
studies, each of the text portions was chosen from a list of
possible options, and each possible combination of these
options, along with multiple prompt versions, was used
for an initial run on SBIC and politeness. After this initial
run, the worst-performing prompts and prompt options
were removed, and all datasets were run again. The
results reported are the best results among all prompts used.</p>
      <sec id="sec-15-1">
        <title>Exact details, including all of the full prompts, prompt</title>
        <p>options, examples, and outputs can be found in our code.</p>
        <p>For GPT-3, we provide either the true
(original/nonimputed) soft label, the imputed soft label, or the true
majority-voted (aggregate) label for the target text. For
the distributional label, we ignore the annotator’s actual
label when computing the distribution, so as not to cause
data leakage. However, when computing the original
majority-voted annotation, we leave in the annotator’s
label for the target example. For the non-highlighted</p>
      </sec>
      <sec id="sec-15-2">
        <title>ChatGPT experiments, when making soft label predic</title>
        <p>tions we use the soft label from the imputed data, rather
than real data. When making individualized annotation
predictions, the example shots are chosen to difer from
the original such that the imputed annotation can be
used.</p>
        <p>F.2. Prompt Skeletons</p>
      </sec>
      <sec id="sec-15-3">
        <title>This section displays the skeletons of each of the prompts used. In practice, the portions of the skeleton surrounded by curly braces are replaced with data, which can be seen in Section F.3.</title>
        <sec id="sec-15-3-1">
          <title>F.2.1. Highlighted ChatGPT Original Data Skeleton Prompt 1</title>
          <p>{dataset_description}
{orig_examples_header}
{orig_examples}
{target_example_header}
{target_example}
{final_words}</p>
        </sec>
        <sec id="sec-15-3-2">
          <title>F.2.2. Highlighted ChatGPT Original Data</title>
        </sec>
        <sec id="sec-15-3-3">
          <title>Skeleton Prompt 2</title>
          <p>{dataset_description}
{target_example_header}
{target_example}
{final_words}</p>
        </sec>
        <sec id="sec-15-3-4">
          <title>F.2.3. Highlighted ChatGPT Original Data</title>
        </sec>
        <sec id="sec-15-3-5">
          <title>Skeleton Prompt 3</title>
          <p>{dataset_description}
{instructions}
{target_example_header}
{target_example}
{final_words}</p>
        </sec>
        <sec id="sec-15-3-6">
          <title>F.2.4. Highlighted ChatGPT Combined Data Skeleton Prompt</title>
          <p>{target_example_header}
{target_example}</p>
        </sec>
        <sec id="sec-15-3-7">
          <title>F.2.5. Highlighted ChatGPT Imputed Data</title>
        </sec>
        <sec id="sec-15-3-8">
          <title>Skeleton Prompt 1</title>
          <p>{imputed_examples}
{target_example}</p>
        </sec>
        <sec id="sec-15-3-9">
          <title>F.2.6. Highlighted ChatGPT Imputed Data</title>
        </sec>
        <sec id="sec-15-3-10">
          <title>Skeleton Prompt 2</title>
        </sec>
        <sec id="sec-15-3-11">
          <title>F.2.7. GPT-3 Distributional Skeleton Prompt</title>
          <p>Here’s a description of a dataset:
{dataset_description}
will be given {n_shots} samples of how that
particular annotator has responded to other
examples and {k_shots} sample of how others
have annotated the target example, and will
then complete the prediction for the target
example as that annotator would.</p>
          <p>Here’s the samples of how the particular
annotator has responded to other examples:
{shots}
Here’s the samples of how others have
annotated the target example:
{other_shots}
How would the particular annotator annotate
the target example?
{target_example_line}
ANSWER:
Given the previous dataset description, F.2.9. GPT-3 Majority-Voted Skeleton Prompt
your goal is to predict how one of the
annotators of the previous dataset would Here’s a description of a dataset:
annotate an example from that dataset. You {dataset_description}
will be given {n_shots} samples of how that
particular annotator has responded to other Given the previous dataset description,
examples and be shown the distributional your goal is to predict how one of the
label of how all annotators have annotated annotators of the previous dataset would
the target example, and will then complete annotate an example from that dataset. You
the prediction for the target example as will be given {n_shots} samples of how that
that annotator would. particular annotator has responded to other
examples and be shown what the plurality
Here’s the samples of how the particular of annotators gave as a label, and will
annotator has responded to other examples: then complete the prediction for the target
{shots} example as that annotator would.
Here’s how the distributional label of how
all annotators have annotated the target
example:
{other_shots}
Here’s the samples of how the particular
annotator has responded to other examples:
{shots}
Here’s how the plurality of annotators
How would the particular annotator annotate labeled the target example:
the target example? {other_shots}
{target_example_line}
ANSWER:</p>
        </sec>
        <sec id="sec-15-3-12">
          <title>F.2.8. GPT-3 Individual Skeleton Prompt</title>
          <p>Here’s a description of a dataset:
{dataset_description}
Given the previous dataset description, {soft_label_examples}
your goal is to predict how one of the {prediction_text}
annotators of the previous dataset would
annotate an example from that dataset. You
How would the particular annotator annotate
the target example?
{target_example_line}
ANSWER:</p>
        </sec>
        <sec id="sec-15-3-13">
          <title>F.2.10. ChatGPT Soft Label Skeleton Prompt F.2.11. (Unused) ChatGPT Contextual Soft Label Skeleton</title>
          <p>Fillers for final_words:
Here is a description of a dataset: 1. "Your output should be a single integer
{dataset_description} corresponding to the label."
2. "Your output should be a single integer
Your goal is to predict the soft label given and nothing else."
by the raters on a particular text. 3. "The only valid output is a single
integer."
Here are a few examples of texts and their 4. "If you output anything other than
soft label: a single integer, your output will be
{soft_label_examples} considered invalid."
5. "If you output anything other than a
Now, you will make your prediction (if you single integer, your output will harm the
are unsure, just give your best estimate): integrity of our dataset."
{prediction_text} 6. "If you output anything other than
a single integer (and absolutely nothing
F.2.12. ChatGPT One of Three Individualized else, including explanatory text), your
Skeleton Prompt output will invalidate the dataset."
7. "If you output anything other than
Description of the dataset: a single integer (and absolutely nothing
{dataset_description} else, including explanatory text), your
output will invalidate the dataset. So,
How annotator A has labeled some examples: please only output a single integer."
{annotator_A_examples} Explicit details of all filler options for all experiments
can be found in the code.</p>
        </sec>
      </sec>
      <sec id="sec-15-4">
        <title>As an example, the options for the filler may look like:</title>
        <p>Example 3:
Text: It’s good to find proof of cheating
you perceive to be happening.</p>
        <p>Annotation from annotator: 3
F.3. Full Prompts
1. orig_examples_header
2. imputed_examples_header
3. target_example_header
4. instructions
5. final_words
This section demonstrates how the skeletons are com- Examples from the dataset (there may not be
pleted with full data for the prompts that are actually any):Examples from the dataset (there may
sent to the OpenAI GPT models. Version numbers such not be any)
as “v4.-1.0.-1.1" indicate which fillers were used for the Example 1:
prompt. The value “v4.-1.0.-1.1" indicates that the 5th Text: It’s good to trust those you care
option was used for the first filler, the second filler was about.
not included in the prompt, the first option was used for Annotation from annotator: 3
the third filler, the second to last filler was not included,
and the 2nd option was used for the last filler.</p>
      </sec>
      <sec id="sec-15-5">
        <title>The fillers, in order, are:</title>
        <p>Example 2:
Text: It’s wrong to grow apart from your
partner.</p>
        <p>Annotation from annotator: 1
How annotator B has labeled some examples:
{annotator_B_examples}</p>
        <sec id="sec-15-5-1">
          <title>F.3.1. Highlighted ChatGPT Original Data Full Prompt 1 (v4.-1.0.-1.1)</title>
          <p>How annotator C has labeled some examples:
{annotator_C_examples}
In this dataset, people are given rules
of thumb, such as "You should be accepting
How might annotator {target_annotator} of other people’s bodies." People then rate
label this example? Output the annotation the rule of thumb as to how much of the world
that seems most likely. population they think would agree with that
{target_example} rule of thumb. 4 means "universal". 3 means
"common". 2 means "controversial". 1 means
"occasional". 0 means "almost no one".</p>
          <p>Example 4:
Text: It’s good to get a friend’s help.</p>
          <p>Annotation from annotator: 3
Example 24:
Text: It is distracting to pin everyone in
an admin team on a discord server.</p>
          <p>Annotation from annotator: 3
Example 25:
Text: It is okay to not share something
that you did a lonf time ago.</p>
          <p>Annotation from annotator: 3
Example 12:
Text: It is okay to show emotions.</p>
          <p>Annotation from annotator: 2
Example 13:
Text: It’s bad to hate babies.</p>
          <p>Annotation from annotator: 3
Example 14:
Text: You shouldn’t scare your coworkers.</p>
          <p>Annotation from annotator: 3
Example 10:
Text: It’s good to give advice to your Example 22:
friends. Text: It’s ok to make your own reproductive
Annotation from annotator: 3 choices.</p>
          <p>Annotation from annotator: 2
Example 11:
Text: You shouldn’t judge people by how Example 23:
they look Text: It’s good for a parent to want
Annotation from annotator: 3 their children to make the best decisions
possible.</p>
          <p>Annotation from annotator: 3
Example 15: Example 26:
Text: It’s not okay to rely on technology Text: It is good to be yourself.
Annotation from annotator: 3
Example 27:
Text: It’s expected that you’ll stay home
from work when you are sick.</p>
          <p>Annotation from annotator: 3
rule of thumb. 4 means "universal". 3 means
"common". 2 means "controversial". 1 means
"occasional". 0 means "almost no one".</p>
          <p>Example 29:
Text: It’s good to want to be able to back
up your words.</p>
          <p>Annotation from annotator: 4
Example 30:
Text: You should treat employees well
Annotation from annotator: 3
Target Example:
Text: It’s OK not to take life too
seriously. Example 2:
Annotation from annotator: Text: It’s not normal to dislike everybody
Your output should be a single integer and Annotation from annotator: 3
nothing else.</p>
        </sec>
        <sec id="sec-15-5-2">
          <title>F.3.4. Highlighted ChatGPT Combined Full Prompt (v1.0.0.-1.-1</title>
        </sec>
        <sec id="sec-15-5-3">
          <title>F.3.3. Highlighted ChatGPT Original Data Full Prompt 3 (v-1.-1.0.1.0)</title>
          <p>In this dataset, people are given rules
of thumb, such as "You should be accepting
of other people’s bodies." People then rate
the rule of thumb as to how much of the world
population they think would agree with that
Example 7:
Text: You shouldn’t be so scared about
getting married.</p>
          <p>Annotation from annotator: 2
Example 8:
Text: It’s wrong to want to hide who you</p>
        </sec>
        <sec id="sec-15-5-4">
          <title>F.3.2. Highlighted ChatGPT Original Data Full Prompt 2 (v-1.-1.0.-1.0)</title>
          <p>In this dataset, people are given rules
of thumb, such as "You should be accepting Example 4:
of other people’s bodies." People then rate Text: It’s okay to be upset if someone you
the rule of thumb as to how much of the world report to at work lets you take the blame
population they think would agree with that for something you didn’t do.
rule of thumb. 4 means "universal". 3 means Annotation from annotator: 3
"common". 2 means "controversial". 1 means
"occasional". 0 means "almost no one".
Example 22:
Text: You should not get angry when your
Example 11: chronically ill relative needs help.
Text: You should thank someone when they Annotation from annotator: 3
invite tou to thier wedding.</p>
          <p>Annotation from annotator: 3
Example 25:
Text: Relationships are expected to have
Example 14: clear communication.</p>
          <p>Text: If you are uncomfortable with your Annotation from annotator: 3
partner feeding your kids, you should not
let your partner watch your children.</p>
          <p>Annotation from annotator: 3
Example 15:
Text: It’s good to want to keep your
children safe in life.</p>
          <p>Annotation from annotator: 3
Example 16:
Text: It’s okay to have sex on camera.</p>
          <p>Annotation from annotator: 2
Example 26:
Text: You shouldn’t try to get sympathy from
people just because you’re in the hospital.</p>
          <p>Annotation from annotator: 2
Example 27:
Text: You shouldn’t be friends with someone
that talks badly about you behind your back.</p>
          <p>Annotation from annotator: 3
Example 29:
Text: It’s not okay to be in a romantic
Example 18: relationship with a teacher.
Text: It is reasonable to want to hold onto Annotation from annotator: 3
loving memories after losing someone.</p>
          <p>Annotation from annotator: 3
Example 3: Example 1:
Text: It’s inappropriate to make passes at Text: It’s wise to never let yourself get
an underling. embarrassed by your significant other.
Annotation from annotator: 3 Annotation from annotator: 2
Example 4: Example 2:
Text: Couple should work together to solve Text: It’s not normal to dislike everybody
problems Annotation from annotator: 3
Annotation from annotator: 3
Example 7:
Text: You can’t expect someone to pay off
their loans immediately.</p>
          <p>Annotation from annotator: 3
Example 12:
Text: It’s not okay to have to rely on
Example 9:
Text: You shouldn’t have sex with people if
you have herpes.</p>
          <p>Example 8:
Text: It’s not necessary to be strong
Annotation from annotator: 3
Example 6:
Text: It’s wrong to spend other people’s
Example 9: money.</p>
          <p>Text: You shouldn’t put up an antenna if Annotation from annotator: 3
it’s snowing outside.</p>
          <p>Annotation from annotator: 1
Example 10:
Text: It’s good to have a best friend.</p>
          <p>Annotation from annotator: 3
Example 8:
Example 11: Text: It’s wrong to want to hide who you
Text: It’s bad to lose track of your really are.
belongings. Annotation from annotator: 2
Annotation from annotator: 3
Example 22:
Text: You should not get angry when your
Example 11: chronically ill relative needs help.
Text: You should thank someone when they Annotation from annotator: 3
invite tou to thier wedding.</p>
          <p>Annotation from annotator: 3
Example 23:
Text: You shouldn’t be bitter and salty
Example 12: about petty things.</p>
          <p>Text: You can’t expect your family members Annotation from annotator: 3
to support you if you’re lazy.</p>
          <p>Annotation from annotator: 3
Example 24:
Text: It’s good to play games of chance
Example 13: where you can win money.</p>
          <p>Text: It’s expected that revealing secrets Annotation from annotator: 2
will lead to trouble.</p>
          <p>Annotation from annotator: 3
Example 25:
Text: Relationships are expected to have
Example 14: clear communication.</p>
          <p>Text: If you are uncomfortable with your Annotation from annotator: 3
partner feeding your kids, you should not
let your partner watch your children.</p>
          <p>Annotation from annotator: 3
Example 15:
Text: It’s good to want to keep your
children safe in life.</p>
          <p>Annotation from annotator: 3
Example 16:
Text: It’s okay to have sex on camera.</p>
          <p>Annotation from annotator: 2
Example 26:
Text: You shouldn’t try to get sympathy from
people just because you’re in the hospital.</p>
          <p>Annotation from annotator: 2
Example 27:
Text: You shouldn’t be friends with someone
that talks badly about you behind your back.</p>
          <p>Annotation from annotator: 3
Example 20:
Text: It’s okay to have an abortion.</p>
          <p>Annotation from annotator: 2
Text: People shouldn’t betray the trust of
those they’re close to.</p>
          <p>Annotation from annotator:
Example 28:
Text: It’s okay to carry your pet in public
Example 17: as a status symbol
Text: It’s wrong for partners to lie to Annotation from annotator: 1
each other.</p>
          <p>Annotation from annotator: 3
Example 29:
Text: It’s not okay to be in a romantic
Example 18: relationship with a teacher.
Text: It is reasonable to want to hold onto Annotation from annotator: 3
loving memories after losing someone.</p>
          <p>Annotation from annotator: 3</p>
        </sec>
        <sec id="sec-15-5-5">
          <title>F.3.6. Highlighted ChatGPT Imputed Data Full Prompt 2 (v-1.0.0.-1.-1)</title>
          <p>Example 2:
Text: It’s not normal to dislike everybody
Annotation from annotator: 3
Example 4:
Text: It’s okay to be upset if someone you
report to at work lets you take the blame
for something you didn’t do.</p>
          <p>Annotation from annotator: 3
Example 5:
Text: You should always respect your
partner’s choices.</p>
          <p>Annotation from annotator: 3
Example 6:
Text: It’s wrong to spend other people’s
money.</p>
          <p>Annotation from annotator: 3
Example 7:
Text: You shouldn’t be so scared about
getting married.</p>
          <p>Annotation from annotator: 2
Example 8:
Text: It’s wrong to want to hide who you
really are.</p>
          <p>Annotation from annotator: 2
Example 9:
Text: You shouldn’t have sex with people if
you have herpes.</p>
          <p>Annotation from annotator: 3
Example 10:
Text: You shouldn’t video tape someone
without their permission.</p>
          <p>Annotation from annotator: 3
Example 11:
Text: You should thank someone when they
invite tou to thier wedding.</p>
          <p>Annotation from annotator: 3
Example 12:
Text: You can’t expect your family members
to support you if you’re lazy.</p>
          <p>Annotation from annotator: 3
Example 13:
Text: It’s expected that revealing secrets
will lead to trouble.</p>
          <p>Annotation from annotator: 3
Example 14:
Text: If you are uncomfortable with your
partner feeding your kids, you should not
let your partner watch your children.</p>
          <p>Annotation from annotator: 3
Example 15:
Text: It’s good to want to keep your
children safe in life.</p>
          <p>Annotation from annotator: 3
Example 16:
Text: It’s okay to have sex on camera.</p>
          <p>Annotation from annotator: 2
Example 17:
Text: It’s wrong for partners to lie to
each other.</p>
          <p>Annotation from annotator: 3
Example 18:
Text: It is reasonable to want to hold onto
loving memories after losing someone.</p>
          <p>Annotation from annotator: 3
Example 19:
Text: You shouldn’t play pranks on people.</p>
          <p>Annotation from annotator: 2
Example 20:
Text: It’s okay to have an abortion.</p>
          <p>Annotation from annotator: 2
Example 21:
Text: It’s wrong to think
responsible for somebody else’s life
Annotation from annotator: 2
Example 22:
Text: You should not get angry when your
chronically ill relative needs help.</p>
          <p>Annotation from annotator: 3 population they think would agree with that
rule of thumb. 4 means "universal". 3 means
Example 23: "common". 2 means "controversial". 1 means
Text: You shouldn’t be bitter and salty "occasional". 0 means "almost no one".
about petty things.</p>
          <p>Annotation from annotator: 3
Given the previous dataset description,
your goal is to predict how one of the
Example 24: annotators of the previous dataset would
Text: It’s good to play games of chance annotate an example from that dataset.
where you can win money. You will be given 4 samples of how that
Annotation from annotator: 2 particular annotator has responded to other
examples and be shown the distributional
Example 25: label of how all annotators have annotated
Text: Relationships are expected to have the target example, and will then complete
clear communication. the prediction for the target example as
Annotation from annotator: 3 that annotator would.</p>
          <p>Example 26: Here’s the samples of how the particular
Text: You shouldn’t try to get sympathy from annotator has responded to other examples:
people just because you’re in the hospital. 1.</p>
          <p>Annotation from annotator: 2 EXAMPLE: It’s good to be honest when you
give your kids the sex talk.</p>
          <p>ANSWER: 4
Example 27:
Text: You shouldn’t be friends with someone
that talks badly about you behind your back. 2.</p>
          <p>Annotation from annotator: 3 EXAMPLE: It’s good to want others to
celebrate your joys with you.</p>
          <p>ANSWER: 3
Example 28:
Text: It’s okay to carry your pet in public
as a status symbol
Annotation from annotator: 1
Example 29:
Text: It’s not okay to be in a romantic
relationship with a teacher.</p>
          <p>Annotation from annotator: 3
3.</p>
          <p>EXAMPLE: It is dangerous for a doctor to
misdiagnose a patient.</p>
          <p>ANSWER: 4
4.</p>
          <p>EXAMPLE: It’s bad for a child to be in
danger.</p>
          <p>ANSWER: 4
Example 30:
Text: You can’t expect everyone to like
you.</p>
          <p>Annotation from annotator: 3
Here’s how the distributional label of how
all annotators have annotated the target
example:
Target Example: 0.00% of people responded with 0
Text: People shouldn’t betray the trust of 18.00% of people responded with 1
those they’re close to. 76.00% of people responded with 2
Annotation from annotator: 6.00% of people responded with 3</p>
        </sec>
        <sec id="sec-15-5-6">
          <title>F.3.7. GPT-3 Non-imputed Distributional Full</title>
        </sec>
        <sec id="sec-15-5-7">
          <title>Prompt</title>
          <p>Here’s a description of a dataset:
In this dataset, people are given rules
of thumb, such as "You should be accepting
of other people’s bodies." People then rate
the rule of thumb as to how much of the world
How would the particular annotator annotate
the target example?
EXAMPLE: It’s expected that families will
not get along well with their in-laws.
ANSWER:</p>
        </sec>
        <sec id="sec-15-5-8">
          <title>F.3.8. GPT-3 Imputed Distributional Full Prompt</title>
          <p>Here’s a description of a dataset:
In this dataset, people are given rules
of thumb, such as "You should be accepting
of other people’s bodies." People then rate
the rule of thumb as to how much of the world
population they think would agree with that
rule of thumb. 4 means "universal". 3 means
"common". 2 means "controversial". 1 means
"occasional". 0 means "almost no one".</p>
          <p>Here’s a description of a dataset:
In this dataset, people are given rules
of thumb, such as "You should be accepting
of other people’s bodies." People then rate
the rule of thumb as to how much of the world
Given the previous dataset description, population they think would agree with that
your goal is to predict how one of the rule of thumb. 4 means "universal". 3 means
annotators of the previous dataset would "common". 2 means "controversial". 1 means
annotate an example from that dataset. "occasional". 0 means "almost no one".
You will be given 4 samples of how that
particular annotator has responded to other Given the previous dataset description,
examples and be shown the distributional your goal is to predict how one of the
label of how all annotators have annotated annotators of the previous dataset would
the target example, and will then complete annotate an example from that dataset.
the prediction for the target example as You will be given 4 samples of how that
that annotator would. particular annotator has responded to other
examples and 49 sample of how others have
Here’s the samples of how the particular annotated the target example, and will
annotator has responded to other examples: then complete the prediction for the target
1. example as that annotator would.
EXAMPLE: It’s good to be honest when you
give your kids the sex talk.</p>
          <p>ANSWER: 4
Here’s the samples of how the particular
annotator has responded to other examples:
1.
2. EXAMPLE: It’s good to be honest when you
EXAMPLE: It’s good to want others to give your kids the sex talk.
celebrate your joys with you. ANSWER: 4
ANSWER: 3
2.
3. EXAMPLE: It’s good to want others to
EXAMPLE: It is dangerous for a doctor to celebrate your joys with you.
misdiagnose a patient. ANSWER: 3
ANSWER: 4
3.
4. EXAMPLE: It is dangerous for a doctor to
EXAMPLE: It’s bad for a child to be in misdiagnose a patient.
danger. ANSWER: 4
ANSWER: 4
the target example?
EXAMPLE: It’s expected that families will
not get along well with their in-laws.</p>
          <p>ANSWER:</p>
        </sec>
        <sec id="sec-15-5-9">
          <title>F.3.9. GPT-3 Individual Full Prompt</title>
          <p>4.</p>
          <p>EXAMPLE: It’s bad for a child to be in
danger.</p>
          <p>ANSWER: 4
Here’s how the distributional label of how
all annotators have annotated the target
example:
0.00% of people responded with 0
10.00% of people responded with 1
85.00% of people responded with 2
5.00% of people responded with 3
2. 13.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
3. 14.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
4. 15.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
5. 16.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 1
6. 17.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
7. 18.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 3 ANSWER: 2
8. 19.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 1 ANSWER: 2
9. 20.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 3 ANSWER: 2
10. 21.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 1 ANSWER: 2
11. 22.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 1 ANSWER: 2
12. 23.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
24. 35.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
25. 36.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
26. 37.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
27. 38.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
28. 39.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
29. 40.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
30. 41.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 1
31. 42.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 2
32. 43.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 3 ANSWER: 2
33. 44.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
ANSWER: 2 ANSWER: 1
34. 45.</p>
          <p>EXAMPLE: It’s expected that families will EXAMPLE: It’s expected that families will
not get along well with their in-laws. not get along well with their in-laws.
give your kids the sex talk.</p>
          <p>ANSWER: 4
46.</p>
          <p>EXAMPLE: It’s expected that families will 2.
not get along well with their in-laws. EXAMPLE: It’s good to want others to
ANSWER: 2 celebrate your joys with you.
ANSWER: 3
47.</p>
          <p>EXAMPLE: It’s expected that families will 3.
not get along well with their in-laws. EXAMPLE: It is dangerous for a doctor to
ANSWER: 2 misdiagnose a patient.</p>
          <p>ANSWER: 4
48.</p>
          <p>EXAMPLE: It’s expected that families will 4.
not get along well with their in-laws. EXAMPLE: It’s bad for a child to be in
ANSWER: 1 danger.</p>
          <p>ANSWER: 4
49.</p>
          <p>EXAMPLE: It’s expected that families will Here’s how the plurality of annotators
not get along well with their in-laws. labeled the target example:
ANSWER: 2 1.</p>
          <p>EXAMPLE: It’s expected that families will
How would the particular annotator annotate not get along well with their in-laws.
the target example? ANSWER: 2
EXAMPLE: It’s expected that families will
not get along well with their in-laws.</p>
          <p>ANSWER:</p>
        </sec>
        <sec id="sec-15-5-10">
          <title>F.3.10. GPT-3 Majority-Voted Full Prompt</title>
        </sec>
        <sec id="sec-15-5-11">
          <title>F.3.11. ChatGPT Non-imputed Soft Label Full</title>
        </sec>
        <sec id="sec-15-5-12">
          <title>Prompt</title>
          <p>Here’s a description of a dataset:
In this dataset, people are given rules
of thumb, such as "You should be accepting
of other people’s bodies." People then rate
the rule of thumb as to how much of the world
population they think would agree with that
rule of thumb. 4 means "universal". 3 means
"common". 2 means "controversial". 1 means
"occasional". 0 means "almost no one".</p>
          <p>Example 390
Text: It’s hurtful for a parent to enroll
a child in classes without permission.</p>
          <p>Soft labels:
0.00% of responsive annotators labeled the
text with 0
Given the previous dataset description, 14.00% of responsive annotators labeled the
your goal is to predict how one of the text with 1
annotators of the previous dataset would 54.00% of responsive annotators labeled the
annotate an example from that dataset. text with 2
You will be given 4 samples of how that 28.00% of responsive annotators labeled the
particular annotator has responded to other text with 3
examples and be shown what the plurality 4.00% of responsive annotators labeled the
of annotators gave as a label, and will text with 4
then complete the prediction for the target
example as that annotator would.</p>
          <p>Example 261
Text: It’s okay to enjoy powdered milk.</p>
          <p>Here’s the samples of how the particular Soft labels:
annotator has responded to other examples: 0.00% of responsive annotators labeled the
1. text with 0
EXAMPLE: It’s good to be honest when you 2.00% of responsive annotators labeled the
How would the particular annotator annotate
the target example?
EXAMPLE: It’s expected that families will
not get along well with their in-laws.</p>
          <p>ANSWER:
text with 1 text with 3
16.00% of responsive annotators labeled the 2.00% of responsive annotators labeled the
text with 2 text with 4
52.00% of responsive annotators labeled the
text with 3 Example 261
30.00% of responsive annotators labeled the Text: It’s okay to enjoy powdered milk.
text with 4 Soft labels:
0.00% of responsive annotators labeled the
Example 20 text with 0
Text: It’s good to find proof of cheating 1.00% of responsive annotators labeled the
you perceive to be happening. text with 1
Soft labels: 8.00% of responsive annotators labeled the
0.00% of responsive annotators labeled the text with 2
text with 0 64.00% of responsive annotators labeled the
0.00% of responsive annotators labeled the text with 3
text with 1 27.00% of responsive annotators labeled the
16.00% of responsive annotators labeled the text with 4
text with 2
64.00% of responsive annotators labeled the Example 20
text with 3 Text: It’s good to find proof of cheating
20.00% of responsive annotators labeled the you perceive to be happening.
text with 4 Soft labels:
0.00% of responsive annotators labeled the
Example 84 text with 0
Text: People shouldn’t betray the trust of 0.00% of responsive annotators labeled the
those they’re close to. text with 1
Soft labels: 8.00% of responsive annotators labeled the
0.00% of responsive annotators labeled the text with 2
text with 0 78.00% of responsive annotators labeled the
0.00% of responsive annotators labeled the text with 3
text with 1 14.00% of responsive annotators labeled the
0.00% of responsive annotators labeled the text with 4
text with 2
44.00% of responsive annotators labeled the Example 84
text with 3 Text: People shouldn’t betray the trust of
56.00% of responsive annotators labeled the those they’re close to.
text with 4 Soft labels:
Target Text: You should always question 0.00% of responsive annotators labeled the
someone’s character when they say something text with 0
inappropriate. 0.00% of responsive annotators labeled the
Soft labels: text with 1
0.00% of responsive annotators labeled the</p>
        </sec>
        <sec id="sec-15-5-13">
          <title>F.3.12. ChatGPT Imputed Soft Label Full Prompt text with 2</title>
          <p>47.00% of responsive annotators labeled the
Example 390 text with 3
Text: It’s hurtful for a parent to enroll 53.00% of responsive annotators labeled the
a child in classes without permission. text with 4
Soft labels: Target Text: You should always question
0.00% of responsive annotators labeled the someone’s character when they say something
text with 0 inappropriate.
7.00% of responsive annotators labeled the Soft labels:
text with 1
67.00% of responsive annotators labeled the
text with 2
24.00% of responsive annotators labeled the</p>
        </sec>
        <sec id="sec-15-5-14">
          <title>F.3.13. ChatGPT Non-imputed One of Three Full</title>
        </sec>
        <sec id="sec-15-5-15">
          <title>Prompt</title>
          <p>Description of the dataset:
In this dataset, people are given rules
of thumb, such as "You should be accepting
of other people’s bodies." People then rate
the rule of thumb as to how much of the world
population they think would agree with that
rule of thumb. 4 means "universal". 3 means
"common". 2 means "controversial". 1 means
"occasional". 0 means "almost no one".</p>
          <p>Example 3:
Text: Employers are expected to let
employees go when business is down.</p>
          <p>How annotator A has labeled some examples: Annotation from annotator: 4.0
Example 1:
Text: You should be focused on learning
while at school.</p>
          <p>Annotation from annotator: 3.0
Example 4:
Text: People shouldn’t assume it’s okay to
take something just because it’s not nailed
down.</p>
          <p>Annotation from annotator: 4.0
Example 1:
Text: You should only call the cops for
emergencies.</p>
          <p>Annotation from annotator: 3.0
Example 2:
Text: It is dangerous for a grandparent to
risk a fall.</p>
          <p>Annotation from annotator: 3.0
How might annotator A label this example?
Output the annotation that seems most
likely.</p>
          <p>Text: You should always question someone’s
character when they say something
inappropriate.</p>
          <p>Annotation from annotator:</p>
        </sec>
        <sec id="sec-15-5-16">
          <title>F.3.14. ChatGPT Imputed One of Three Full</title>
        </sec>
        <sec id="sec-15-5-17">
          <title>Prompt</title>
          <p>Example 2:
Text: It is dangerous for a doctor to
misdiagnose a patient.</p>
          <p>Annotation from annotator: 3.0
Example 3:
Text: It’s understandable if you can’t
keep your political leanings out of your
religion lectures.</p>
          <p>Annotation from annotator: 3.0
Example 4:
Text: You shouldn’t date a bad partner.</p>
          <p>Annotation from annotator: 3.0
Description of the dataset:
In this dataset, people are given rules
How annotator B has labeled some examples: of thumb, such as "You should be accepting
Example 1: of other people’s bodies." People then rate
Text: You shouldn’t be with someone who the rule of thumb as to how much of the world
sleeps with your kids. population they think would agree with that
Annotation from annotator: 4.0 rule of thumb. 4 means "universal". 3 means
"common". 2 means "controversial". 1 means
"occasional". 0 means "almost no one".</p>
          <p>Example 2:
Text: You shouldn’t have sex with people if
you have herpes.</p>
          <p>Annotation from annotator: 3.0
How annotator A has labeled some examples:
Example 1:
Text: It’s hurtful for a parent to enroll
Example 3: a child in classes without permission.
Text: You shouldn’t be crazy about things. Annotation from annotator: 2.0
Annotation from annotator: 2.0
Example 2:
Example 4: Text: It’s okay to enjoy powdered milk.
Text: If you don’t want to crash, you Annotation from annotator: 3.0
shouldn’t drive recklessly.</p>
          <p>Annotation from annotator: 3.0
Example 3:
Text: It’s good to find proof of cheating
How annotator C has labeled some examples: you perceive to be happening.
Annotation from annotator: 3.0</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-16">
      <title>G. PCA Results</title>
      <p>Example 4: One of the fundamental aspects of imputation methods is
Text: People shouldn’t betray the trust of how they treat and interpret data. In the 2-dimensional
those they’re close to. scatter plot based on the first two principal components
Annotation from annotator: 3.0 of imputed datasets, clear variations can be observed
across diferent imputation techniques. This visualization
How annotator B has labeled some examples: underscores the unique characteristics of each
imputation method. Refer to Figure 6 for a detailed comparison.</p>
    </sec>
    <sec id="sec-17">
      <title>H. Variance and Disagreement</title>
      <sec id="sec-17-1">
        <title>Post-imputation, a notable observation is the drop in</title>
        <p>variance, as shown in Figure 7. This phenomenon can
be attributed to the fact that most imputation methods
tend to approximate missing values based on observed
patterns in the data, leading to a convergence of values
around certain estimates.</p>
      </sec>
    </sec>
    <sec id="sec-18">
      <title>I. Soft Label Analysis Extra</title>
    </sec>
    <sec id="sec-19">
      <title>Examples</title>
      <p>Example 1:
Text: It’s hurtful for a parent to enroll
a child in classes without permission.
Annotation from annotator: 2.0
Example 2:
Text: It’s okay to enjoy powdered milk.
Annotation from annotator: 3.0
Example 3:
Text: It’s good to find proof of cheating
you perceive to be happening.</p>
      <p>Annotation from annotator: 3.0
Example 4:
Text: People shouldn’t betray the trust of
those they’re close to.</p>
      <p>Annotation from annotator: 3.0</p>
      <sec id="sec-19-1">
        <title>The computation to determine whether an example has</title>
        <p>Example 2: is “low", “medium", or “high" disagreement was done
inText: It’s okay to enjoy powdered milk. dividually for each fold of the data. When given a fold
Annotation from annotator: 2.0 of data, we first compute the proportion of people who
disagreed with the majority-voted label. (Note that ties in
Example 3: the majority-voted label do not impact this computation,
Text: It’s good to find proof of cheating since the same number of people will disagree regardless
you perceive to be happening. of which label is chosen among the tied options.) Then,
Annotation from annotator: 3.0 we assign a threshold for “low" and “high" disagreement:
any examples with disagreement equal to or lower than
Example 4: the “low" threshold are considered to have “low"
disagreeText: People shouldn’t betray the trust of ment, while any examples with disagreement equal to or
those they’re close to. greater than the “high" threshold are considered to have
Annotation from annotator: 4.0 “high" disagreement. The number of examples in each
category is a sum across all five folds of that dataset of
How might annotator A label this example? examples that matched the threshold for that category.
Output the annotation that seems most The choice of thresholds must satisfy three rules: (1)
likely. The high threshold must be higher than the low threshold
Text: You should always question someone’s (2) There must be at least some examples in each category
character when they say something (3) The variance among the number of examples in each
inappropriate. category must be minimized. When looking at Table 7,
Annotation from annotator: it may seem odd that the number of examples in each
category is so varied, given the explicit minimization of
25
20
15
10
5
0
-400 -200 0 200 400 600 800
-10 0 10 20 30 40 50 60
-20 -10 0 10 20 30 40 50 60
Original SChem</p>
        <p>Kernel Imputed SChem</p>
        <p>NCF Imputed SChem</p>
        <p>Multitask Imputed SChem
-40
-20 0
-5
Kernel Imputed SChem
5 Labels</p>
        <p>NCF Imputed SChem
5 Labels</p>
        <p>Multitask Imputed SChem
5 Labels
10
20 30 40
-150 -100 -50 0 50 100 150 200
-150 -100 -50 0
50 100 150
-200 -100
Original GHC</p>
        <p>Kernel Imputed GHC</p>
        <p>NCF Imputed GHC
40
30
20
10
0
-10
-30 -20 -10 0
600
400
200
0
-200
-400
60
40
20
0
-20
-40
150
100
50
0
-50
-100
100
80
60
40
20
0
-20</p>
        <p>Original SChem</p>
        <p>5 Labels
-25 0
25 50 75 100 125 150
150
100
50
0
-50
30
20
10
0
-10
-20
12
10
8
5
3
0
-3
-5
40
30
20
10
0
-10
-20
-30
Multitask Imputed Politeness
400
300
200
100</p>
        <p>0
-100
-200
-300
-400 -200 0 200 400 600
Kernel Imputed Politeness
400
300
200
100</p>
        <p>0
-100
-200
-300
60
40
20
0
-20
-40
-60
-400 -200
0
50
100
150
200
-200 -100
0
100
200
0
200
400
variance in the rules. However, this occurs because there
are many examples that have the exact same level of
disagreement; rather than split these examples into two
diferent categories, we opted to ensure that examples
with the same level of disagreement were always labeled
with the same level of disagreement.</p>
        <p>Figure 8: Examples from the SBIC dataset.</p>
        <p>Figure 9: Examples from the Politeness dataset.</p>
        <p>Figure 10: Examples from the SChem5Labels dataset.</p>
        <p>Figure 11: Examples from the Sentiment dataset.</p>
        <p>Figure 12: Examples from the GHC dataset.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>