<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Can ChatGPT Be Useful for Distant Reading of Music Similarity?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arthur Flexer</string-name>
          <email>arthur.flexer@jku.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HCMIR23: 2nd Workshop on Human-Centric Music Information Research</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Johannes Kepler University Linz</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We explore whether large language models (ChatGPT) can be used as a 'distant reading' tool to estimate music similarity between songs from textual information, complementing experiments with human participants. We compare degrees of rater agreement to previous results from a listening test, showing that correlation of ChatGPT with human raters is significantly lower than the average human inter-rater agreement, but nevertheless still of moderate positive size. We discuss whether an approach based on the largely opaque ChatGPT model can be scientifically valid and to what extent it allows transparent evaluation of music information research experiments.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;music similarity</kwd>
        <kwd>large language models</kwd>
        <kwd>transparency</kwd>
        <kwd>evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        chatbot imitating a human conversational partner and is also based on a ‘Generative Pre-trained
Transformer’ (GPT). More specifically it is based on GPT-3.5, which itself is a fine-tuned version
of GPT-3 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A problem common to all members of the GPT family (including ChatGPT) is that
the exact models, training sets, parameters, etc are not known. A non peer reviewed report [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
by the developing team about the latest version (GPT-4) even states that "[...] no further details
about the architecture (including model size), hardware, training compute, dataset construction,
training method, or similar" can be given due to "safety implications" and the "competitive
landscape" of LLM research.
      </p>
      <p>
        Most related to our work is an experiment using ChatGPT to rate musical instrument sounds
on a set of 20 semantic scales [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. ChatGPT’s results are only partially correlated with human
ratings, most pronounced for clearly defined dimensions of musical sounds such as brightness
(bright–dark) and pitch height (deep–high) with Pearson correlations above 0.80. This is related
to work on distilling psychophysical information from text by aligning GPT-4 results with human
auditory experience [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Other applications of LLMs to music include lyrics summarization [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
and usage as ranking models for music recommendation [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Applying ChatGPT to music
data is also reminiscent of earlier approaches to compute music similarity using textual sources,
most notably web-based data [12], but also semantic music tags [13] or lyrics [14].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Methods</title>
      <sec id="sec-2-1">
        <title>2.1. Human evaluation of music similarity</title>
        <p>
          We will compare our ChatGPT results to a previous study that conducted a series of listening
tests with human participants [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. In this study the age of participants ranged from 26 to 34
years with an average of 28.2. The sample consisted of three females and three males, which we
call graders S1 to S6 from here on. The 5 × 18 songs belonged to five genres (for a full list see
section A of the appendix of the original article): (i) American Soul from the 1960s and 1970s
with only male singers singing; (ii) Bebop, the main jazz style of the 1940s and 1950s, with
excerpts containing trumpet, saxophone and piano parts; (iii) High Energy (Hi-NRG) dance
music from the 1980s, typically with continuous eighth note bass lines, aggressive synthesizer
sounds and staccato rhythms; (iv) Power Pop, a Rock style from the 1970s and 1980s, with
chosen songs being guitar-heavy and with male singers; (v) Rocksteady, which is a precursor
of Reggae with a somewhat soulful basis. It is worth noting that one aim of the original study
was to find less known songs by making sure that every song had under 50.000 accesses on
Spotify. Genres were validated via respective Wikipedia artist pages as well as by listening to all
songs. For presentation in the listening test, 15 seconds of a representative part of every song
(usually the refrain) were chosen and participants were asked to “assess the similarity between
the query song and each of the five candidate songs by adjusting the slider” (ranging from 0
to 100 %) and “to answer intuitively since there are no wrong answers”. Based on randomly
chosen 15 query songs, comparisons of five pairs had to be made for every query group yielding
a total of 15 × 5 = 75 pairs, with every song appearing exactly once in the whole questionnaire
(15 as query songs, 75 as candidate songs).
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. ChatGPT evaluation of music similarity</title>
        <p>For our experiments conducted on the 5th and 6th of April 2023 we used the "Free Research
Preview" of the ChatGPT Mar 23 Version. The service came with a warning that "ChatGPT
may produce inaccurate information about people, places, or facts" and the information that
"ChatGPT is fine-tuned from GPT-3.5, a language model trained to produce text. ChatGPT
was optimized for dialogue by using Reinforcement Learning with Human Feedback (RLHF)
– a method that uses human demonstrations and preference comparisons to guide the model
toward desired behavior".</p>
        <p>For the exact same 15 × 5 = 75 song pairs as used in the human listening test we asked
ChatGPT the following question: "On a scale of 1 to 100, how similar is the song [s_i] by
[artist_A] to the song [s_j] by [artist_B]?". It is worth noting that ChatGPT sometimes needed
persuasion to provide an answer at all, stating e.g. that "As an AI language model, I do not have
the ability to directly listen to music or interpret subjective qualities such as similarity between
songs", or that any answer would be "merely speculation". However, the following additional
input sentences provided by us in ensuing dialogues always resulted in ChatGPT answering
with a similarity score: "Please just make a guess based on the information you have already",
"Please try anyway", "Then please just speculate". This kind of persuasion was necessary for 8
out of 75 questions, mostly at the beginning of ChatGPT sessions. Due to restriction of the free
ChatGPT version experiments had to be split over three separate sessions.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>
        In accordance with the previous study [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], we analyse the degree of inter-rater agreement by
computing the Pearson correlations1   between graders S1 to S6 as well as   between
graders S1 to S6 and ChatGPT for the 75 pairs of query/candidate songs. Please note that the
human listening test had been conducted twice at time points t1 and t2 with a two week time
lag [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The 15 plus 15 correlations   (t1 and t2) between the six graders range from 0.59 to
0.86, with an average of 0.74 (see Table 1 for individual correlations). The 6 plus 6 correlations
  (t1 and t2) between the six graders and ChatGPT are considerably lower, with a range from
0.39 to 0.72 and an average of 0.58. The diferences in correlation between   and   are
statistically significant (t(40)=6.05, p=0.00).
      </p>
      <p>When ChatGPT answers the similarity questions it always also provides some form of
explanation. Sometimes they are very brief: "Based on the available information, I would
speculate [...]". Often genre, instrumentation or era of recording are being debated: "[...] given
that both artists were active in the same time period and were part of the Jamaican music scene, it
is possible that there may be some similarities in terms of instrumentation, rhythm, or vocal style"
or "They are from diferent musical genres, diferent eras, and have diferent rhythms, melodies,
instrumentation, and lyrics". Some explanations are quite detailed: "Both songs are characterized
1One of the advantages of Pearson correlation is that it implicitly normalizes for diferent rating styles because it
is invariant under separate changes in location and scale in the two correlated variables. For example a Pearson
correlation is perfect if two raters give identical answers, but also if all answers of one of the raters are always
shifted by e.g. 10 units. Other measures of rater agreement like Fleiss’ Kappa are only defined for the categorical
scale, but applicable alternatives for the interval scale like Krippendorf’s alpha exist.
time point t1</p>
      <p>S3 S4
0.74 0.72
0.72 0.75
1.00 0.70
1.00
time point t2</p>
      <p>
        S3 S4
0.73 0.77
0.73 0.74
1.00 0.69
1.00
by smooth, soulful vocal performances and feature catchy melodies and memorable hooks.
Additionally, both songs deal with themes of love and relationships, which further underscores
their similarities". ChatGPT has been criticized for sometimes ‘hallucinating’ [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] facts that
sound plausible but are actually incorrect. We verified that ChatGPT’s argumentation seems
to be correct basically all the time by searching and reading respective online sources (e.g.
Wikipedia or Discogs), or, in case these cannot be found, by listening to the audio.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion and Conclusion</title>
      <p>Although the mean Pearson correlation of ChatGPT with human raters is significantly lower
(0.58) than the average inter-rater agreement (0.74), it still remains at a relatively high level.
Even the ranges of correlations   and   are overlapping, hence some of the raters agree
to a higher degree with ChatGPT than with some of the other human raters.</p>
      <p>
        Many of the explanations provided by ChatGPT are about music genre or instrumentation,
which can be seen as an indirect indication of genre. Interestingly, previous self-reports by
human raters [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] have already indicated that genre is an important aspect when rating similarity
of songs. When comparing ChatGPT similarity scores within genres to scores obtained when
query and candidate songs are of diferent genres, on average within genre scores are always
higher than between genre scores. This efect is most pronounced for ‘Soul’ and least for
‘Power Pop’. It would be interesting to explore whether correlations between ChatGPT and
human raters would decrease if all song material were part of the same genre. For human raters,
agreement for such a single single genre study indeed was lower [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] since genre of songs seems
a major factor in judging similarity between songs.
      </p>
      <p>Although it seems that music similarity, as measured via human listening tests, can to a
certain degree be recovered form textual data by using ChatGPT as a distant reading tool, the
black box nature of LLMs and especially ChatGPT remains a problem. Since the exact training
data and modeling approach are unknown (see sec. 1), it remains unclear what enabled ChatGPT
to provide answers that at least partially are in alignment with human feedback to listening
tests. One very likely possibility is that respective webpages about artists and songs have
been part of ChatGPT’s training data, allowing ChatGPT to reproduce this content when being
queried accordingly. Another possibility is that ChatGPT is actually able to reason about musical
concepts like genre, instrumentation, harmony, etc. How ChatGPT arrives at a similarity score
for a pair of songs remains completely mysterious. Because the full nature of ChatGPT has not
been disclosed we are left to speculate about these matters. The future will show whether such
opaque tools will nevertheless become part of science’s repertoire or if they will be confined to
practical applications where measurable success is suficient and scientific rigour not necessary.
This could mean that music similarity measured via ChatGPT could e.g. be useful for music
recommendation systems but not for research on human semantic music similarity concepts.</p>
      <p>As a final comment it will also be interesting to see how open source alternatives to ChatGPT
like LLaMA [15], Alpaca [16] or Open-Assistant (https://github.com/LAION-AI/Open-Assistant)
will change assessment of the usefulness of large language models for distant reading.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This research was funded in whole, or in part, by the Austrian Science Fund (FWF) [P 36653].
For the purpose of open access, the author has applied a CC BY public copyright licence to any
Author Accepted Manuscript version arising from this submission.
[12] P. Knees, E. Pampalk, G. Widmer, Artist classification with web-based data, in: Proceedings
of the 5th International Conference on Music Information Retrieval, 2004.
[13] D. Turnbull, L. Barrington, D. Torres, G. Lanckriet, Towards musical
query-by-semanticdescription using the cal500 data set, in: Proceedings of the 30th annual international
ACM SIGIR conference on Research and development in information retrieval, 2007, pp.
439–446.
[14] B. Logan, A. Kositsky, P. Moreno, Semantic analysis of song lyrics, in: IEEE International</p>
      <p>Conference on Multimedia and Expo (ICME), volume 2, IEEE, 2004, pp. 827–830.
[15] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière,
N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, G. Lample, Llama: Open
and eficient foundation language models, 2023. arXiv:2302.13971.
[16] R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, T. B. Hashimoto,
Stanford alpaca: An instruction-following llama model, https://github.com/tatsu-lab/
stanford_alpaca, 2023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Moretti</surname>
          </string-name>
          ,
          <source>Conjectures on world literature, New left review 2</source>
          (
          <year>2000</year>
          )
          <fpage>54</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Moretti</surname>
          </string-name>
          , Graphs, maps, trees
          <article-title>: abstract models for a literary history</article-title>
          ,
          <source>Verso</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Subbiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Neelakantan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          , et al.,
          <article-title>Language models are few-shot learners</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] OpenAI, Gpt-4
          <source>technical report</source>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>08774</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Flexer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lallai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Rašl</surname>
          </string-name>
          ,
          <article-title>On evaluation of inter- and intra-rater agreement in music recommendation</article-title>
          ,
          <source>Transactions of the International Society for Music Information Retrieval</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ) (
          <year>2021</year>
          )
          <fpage>182</fpage>
          -
          <lpage>194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schedl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Flexer</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Urbano,</surname>
          </string-name>
          <article-title>The neglected user in music information retrieval research</article-title>
          ,
          <source>Journal of Intelligent Information Systems</source>
          <volume>41</volume>
          (
          <year>2013</year>
          )
          <fpage>523</fpage>
          -
          <lpage>539</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Siedenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Saitis</surname>
          </string-name>
          ,
          <article-title>The language of sounds unheard: Exploring musical timbre semantics of large language models</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2304</volume>
          .
          <fpage>07830</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Marjieh</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sucholutsky</surname>
          </string-name>
          , P. van Rijn,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jacoby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Grifiths</surname>
          </string-name>
          ,
          <article-title>Large language models predict human sensory judgments across six modalities</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2302</volume>
          .
          <fpage>01308</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Jiang, G. Xia,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dixon</surname>
          </string-name>
          ,
          <article-title>Interpreting song lyrics with an audio-informed pre-trained language model</article-title>
          ,
          <source>in: Proceedings of the 23rd International Society for Music Information Retrieval Conference</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>McAuley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Large language models are zero-shot rankers for recommender systems</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2305</volume>
          .
          <fpage>08845</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>