<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Edward Said at Touché: Human Value Detection Using Transformers and Upsampling</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aisha Nur Aydin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shaden Shaar</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claire Cardie</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Conclusion Stance Premise</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>In this paper, we tackle both subtasks of the proposed shared task Human Value Classification at Touché- that aims to classify dialogue speech into one of 19 human values determined by Schwartz's Refined Theory of Basic Individual Values. We fine-tune models like DeBERTa and RoBERTa with F1-loss to handle multi-label settings. We additionally test diferent sampling strategies to accommodate for data imbalance. We found that by training on the English-translated utterances, we beat the baselines by at least 2 F1 points across both subtasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;human-value-classification</kwd>
        <kwd>Touché</kwd>
        <kwd>CLEF</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>There are nineteen human values in Schwartz’s Refined Theory of Basic Individual Values [ 3]. There is
a wide range of values in this taxonomy. Additionally, since human values can be addressed implicitly, it
can be challenging to identify which value is being used in a given text [4]. The human value detection
task for 2023 had three components that made up each argument: a conclusion, a stance, and a premise.
Here is an example argument from the 2023 Human Values Task:</p>
      <p>We should ban human cloning
in favor of
We should ban human cloning as it will only cause huge issues when you have
a bunch of the same humans running around all acting the same.</p>
      <p>This year, the task takes a sentence as an input and outputs its human value, whether it is (even
partially) attained, constrained, or neither. A value is attained if the text supports it and is constrained
if it hinders it. This required us to treat the problem as a multi-class, multi-label classification problem
with 38 labels consisting of 19 human values that are attained or constrained.</p>
      <p>An example sentence with a human value is the following. The input is:
"Young women examining their options for third level education have been urged to consider
careers in science, technology, engineering or maths (STEM)."
The value referred to in this sentence (the output) is Achievement Attained.</p>
      <p>These sentences are provided by the ValueML dataset , which contains text and their respective
human values from articles and political text. The data consists of 3000 texts in over eight languages.
20% of the data is used for validation, 20% is part of testing, and the remaining 60% is for training.</p>
    </sec>
    <sec id="sec-3">
      <title>3. System Overview</title>
      <p>For all our experiments, we chose to combine all languages in the dataset into one by keeping the
English utterances only. We use RoBERTa Large, specifically "FacebookAI/roberta-large" as the base
model for our submission. We first experimented with the original RoBERTa and DeBERTa models
but found the F1 scores on the validation dataset lower than the BERT baseline provided by the task
organizers. We then fine-tuned RoBERTa-Large and DeBERTa-Large and observed improved results.
There was an improvement, but the overall F1 score appeared to be afected by the class imbalance.
Human Values that appeared less in the dataset had noticeably lower F1 scores than those that seemed
more.</p>
      <p>We attempted multiple configurations of upsampling but ended up using an upsampling technique
where the human values with lower performance metrics were chosen to be upsampled by a factor of
four.</p>
      <p>To determine which human values to upsample, we fine-tuned the RoBERTa large model on the
training data and analyzed the evaluation results. If the F1 score was 0.15 or less for subtask 1, we
upsampled the attains and constrains form of that human value. We also looked at the metrics for
subtask 2 to determine whether we should upsample only one of the attains or constrains of that value.
If the recall for one of the constrains or attains was 50% or less than its counterpart, we only increase
the underperforming one. For example, if the recall of the attains version of a human value was 50% or
less than the constrains version, we only choose to upsample the attains version and vice versa. Based
on the metrics and the unevenness of the dataset, we upsampled these values by a factor of four:
• Self-direction: thought constrained
• Self-direction: action constrained
• Humility attained
• Humility constrained
• Face attained
• Face constrained
• Benevolence: caring constrained
• Benevolence: dependability constrained
• Universalism: tolerance attained
• Universalism: tolerance constrained
• Conformity: interpersonal attained
• Conformity: interpersonal constrained
• Tradition constrained
• Power: dominance constrained</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Setup</title>
      <p>To fine-tune the model, we set the learning rate to 2e-5, used a warm-up ratio of 0.2, set the batch size
to 8, and used four epochs. We put the random seed to 42, used the AdamW optimizer, and used a
linear scheduler. We used one A100 GPU to run the experiments for the final submitted models. Our
experiments used pre-trained RoBERTa [5] and DeBERTa [6] models. For evaluation, we use Precision,
Recall, and the macro F1-score. We fine-tuned four models: RoBERTa and DeBERTa large on the
upsampled data, and RoBERTa and DeBERTa large on the regular data, with no upsampling. For our
ifnal submission, we chose the model fine-tuned with the upsampled data using RoBERTa large as the
base. This had the best metrics among all four models on the validation dataset. All four models can be
found on Hugging Face2.
:iitttrcegoodunhh :iiit-tfrcceaoodnn iltaon isnm teeevnm :irceeaodnnm :ssrrrceeeou :iltsrreaoypn :iilttsrceaoy iiton :iltsrreoyum :iilttfsrrreeaooypnnm iilty :ilrcceeagvonn :iilltceeeeavboyddpnn :ilssrrcceeaonnm :iltssrreeaaunm :illtssrrceeeaaonm
EN llA l-feS leS itSum eodH ichA oPw oPw ceaF ceSu ceSu radT fonC onC uHm eenB eenB ivnU ivnU ivnU</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>The upsampling methods resulted in a 4% increase for the macro F1 score of subtask 1 and a 2% increase
for the macro F1 score of subtask 2. Of the upsampled human values, "Benevolence: caring" and
"Tradition" were the only ones that did not have an improved F1-score for subtask 1. For the other
upsampled values, there were improved F1-scores, some marginal and others significant. For subtask
2, the upsampled values performed only marginally better except for "Self-direction: thought" and
"Humility", which all had worse metrics, and "Benevolence: dependability" and "Universalism: tolerance",
which stayed the same.</p>
      <p>After the task deadline, we looked at the performances of the other three models on the test data.
Those models were DeBERTa large fine-tuned on upsampled data and DeBERTa large and RoBERTa
large fine-tuned on the original training data. The results comparing all of these models are on Table 1
and Table 2. The RoBERTa large model with upsampling performs better than the other three models in
subtask 1, but ends up being the worst-performing model on subtask 2. The best performing model for
subtask 2 is DeBERTa large without upsampling.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>We were able to beat the BERT baselines by incorporating the F1-Loss function and up-sampling
lower-performing categories. This entails that the data benefits from knowledge transfer obtained from
the diferent categories. In the future, we would also like to consider diferent languages as just using
the translation could have led to missed cultural nuances expressed in language.
J. Odijk, S. Piperidis (Eds.), Proceedings of the Thirteenth Language Resources and Evaluation
Conference, European Language Resources Association, Marseille, France, 2022, pp. 6948–6958.</p>
      <p>URL: https://aclanthology.org/2022.lrec-1.751.
[3] S. H. Schwartz, J. Cieciuch, M. Vecchione, E. Davidov, R. Fischer, C. Beierlein, A. Ramos, M. Verkasalo,
J.-E. Lönnqvist, K. Demirutku, et al., Refining the Theory of Basic Individual Values, Journal of
personality and social psychology 103 (2012). doi:10.1037/a0029393.
[4] J. Kiesel, M. Alshomary, N. Mirzakhmedova, M. Heinrich, N. Handke, H. Wachsmuth, B. Stein,
SemEval-2023 Task 4: ValueEval: Identification of Human Values behind Arguments, in: R. Kumar,
A. K. Ojha, A. S. Doğruöz, G. D. S. Martino, H. T. Madabushi (Eds.), 17th International Workshop on
Semantic Evaluation (SemEval 2023), Association for Computational Linguistics, Toronto, Canada,
2023, pp. 2287–2303. doi:10.18653/v1/2023.semeval-1.313.
[5] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov,</p>
      <p>RoBERTa: A robustly optimized BERT pretraining approach (2019). arXiv:1907.11692.
[6] P. He, X. Liu, J. Gao, W. Chen, DeBERTa: Decoding-enhanced BERT with disentangled attention
(2020). arXiv:2006.03654.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kiesel</surname>
          </string-name>
          , Ç. Çöltekin,
          <string-name>
            <given-names>M.</given-names>
            <surname>Heinrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alshomary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Longueville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Erjavec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Handke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kopp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ljubešić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Meden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mirzakhmedova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Morkevičius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Reitis-Münstermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Scharfbillig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Stefanovitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wachsmuth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , Overview of Touché 2024:
          <article-title>Argumentation Systems</article-title>
          , in: L.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Quénot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Schwab</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Soulier</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. M. D. Nunzio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Galuščáková</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fifteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Darwish</surname>
          </string-name>
          ,
          <article-title>Cross-lingual emotion detection</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Béchet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Blache</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Cieri</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Goggi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Isahara</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Maegaard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mariani</surname>
          </string-name>
          , H. Mazo,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>