<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Can LLMs solve generative visual analogies?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shrey Pandit</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gautam Shrof</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ashwin Srinivasan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lovekesh Vig</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>BITS Pilani</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K.K. Birla Goa Campus</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>TCS Research</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>New Delhi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India</string-name>
        </contrib>
      </contrib-group>
      <fpage>30</fpage>
      <lpage>32</lpage>
      <abstract>
        <p>Recent experiments with large language models (LLMs) have provided some evidence that these models can perform abstract analogical reasoning [1], including textual puzzles similar to Raven's progressive matrices. We consider a visual analogical reasoning task that was solved using neuro-symbolic techniques in [2], and investigate how LLMs fare on this task. The task involves learning a sequence of transformations by which a sample input/output pair of images are related so as to analogously transform a test input. Note that unlike the analogical reasoning tasks in [1], this task involves generating an output as opposed to selecting from a set of choices. We evaluated various LLMs including GPT-4, GPT 3.5-turbo (ChatGPT), and GPT3 on this task for difering lengths of the sequence of transformations relating the input and output. Our results suggest that GPT-4 performs the best overall, while GPT 3.5-turbo and GPT3 perform strongly on shorter program lengths. At the same time, the performance of LLMs for this task falls far short of the neuro-symbolic approach used earlier, and we speculate as to why this may be the case, at least as of now.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large language models</kwd>
        <kwd>GPT-4</kwd>
        <kwd>Visual analogy</kwd>
        <kwd>Neural analogical reasoning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Experiments</title>
      <p>We applied LLMs to solve such a visual analogy task, with the image translated into symbolic
form. We use the trained models provided by OpenAI API and give a few solved examples in
the prompt to help the LLMs learn.</p>
      <p>Prompting We prompt the LLMs with a set of rules for the task, which includes information
on the allowed positional shifts, the set of permissible states, and the non-wrapping state of the
grid. As a hint, we also specify the expected program length in the prompt. We also provide a
set of solved examples to guide the LM’s learning process. Figure 1a illustrates a representative
example of the prompt.
(a) Prompt with rules, solved
examples and test input.</p>
      <p>(b) Chain-of-thought vs. simple</p>
      <p>prompting
,
)</p>
      <p>) Task</p>
      <p>Solution
(c) Given example input-output
image pairs, generate the
analogous output for the
given query image.</p>
      <p>GPT-3</p>
      <sec id="sec-2-1">
        <title>GPT-3.5-turbo GPT-4</title>
        <p>GPT3
Simple Prompting
Chain-of-thought
prompting</p>
      </sec>
      <sec id="sec-2-2">
        <title>Tokens Accuracy</title>
      </sec>
      <sec id="sec-2-3">
        <title>Tokens Accuracy 2847 48%</title>
        <p>
          Chain-of-thought prompting: Previous works such as [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] have shown that language-model
performance increases drastically in reasoning tasks when given chain-of-thought prompts. In
our context we provide a chain-of-thought by providing, with each step of the solved example,
the current state on the grid, positional shift, next position on grid, and the state of the output
bufer.
        </p>
        <p>Changing the program length: We experimented with diferent program lengths (3 &amp; 5);
empirically, increasing the program length makes the task more dificult.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results and Conclusions</title>
      <p>Referring to Table 1 we observe the following: (i) Chain-of-thought improves performance
over simple prompting, which was expected. (ii) Further, chain-of-thought prompting is also
better (albeit slightly) than providing more examples for similar token lengths, see Table 2. (ii)
Sampling more examples improves performance, also expected. (iii) Analogies involving longer
sequences (programs) are more dificult as expected; however we observe a drastic drop in
performance for GPT3 and GPT 3.5-turbo but only a marginal drop is observed for GPT4. While
the input prompts do afect the LLMs’ performance, it’s important to mention that a uniform
prompt template was used for all the analyzed LLMs in this study.</p>
      <p>
        Overall the performance of LLMs for our simple visual analogy task fall far short of the
neuro-symbolic techniques used in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We note that [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] relied search over possible sequences
that could successfully transform a text input to its output. LLMs do not explicitly search over
potential outputs. We speculate that incorporating elements of explicit search may enable LLMs
to perform better at generative analogies.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>W.</surname>
          </string-name>
          et. al,
          <article-title>Emergent analogical reasoning in large language models</article-title>
          ,
          <source>arXiv:2212.09196</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>S.</surname>
          </string-name>
          et. al,
          <article-title>Solving visual analogies using neural algorithmic reasoning</article-title>
          ,
          <source>AAAI (Student Abstract)</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>J. W.</surname>
          </string-name>
          et. al,
          <article-title>Chain of thought prompting elicits reasoning in large language models</article-title>
          ,
          <source>NeurIPS</source>
          <year>2022</year>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>