<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Workshop on AI Evaluation Beyond Metrics (EBeM)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>José Hernández-Orallo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lucy Cheke</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joshua Tenenbaum</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tomer D. Ullman</string-name>
          <email>tullman@fas.harvard.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fernando Martínez-Plumed</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danaja Rutar</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Burden</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryan Burnell</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wout Schellaert</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for the Study of Existential Risk, University of Cambridge</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>EBeM'22: Workshop on AI Evaluation Beyond Metrics</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Harvard University, Department of Psychology</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="US">United States of America</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Leverhulme Centre for the Future of Intelligence, University of Cambridge</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Massachusetts Institute of Technology, Department of Brain and Cognitive Sciences</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country>United States of</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Valencian Research Institute for Artificial Intelligence (VRAIN), Universitat Politècnica de València</institution>
          ,
          <addr-line>Valencia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We summarize the IJCAI-22 Workshop on Artificial Intelligence Evaluation Beyond Metrics, held at the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI-ECAI) on July 24.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Traditional approaches to AI evaluation lack the necessary robustness to analyse the capabilities
of complex AI systems. Many AI systems solve a task or excel at a particular benchmark,
but then fail at other tasks or instances that putatively represent the same capability. This
problem becomes more blatant as performance in many AI benchmarks ramps up rapidly, while
concomitant improvement in capability is moderate, or limited to specific areas (e.g., language
models), which are assessed informally and with low robustness. This generates a large disparity
between the “hype” of AI achievement and the reality, and contributes to a feeling of distrust
about what AI is really capable of. It also makes it dificult to predict the potential operational
applications of AI in the future, especially as AI systems become less task-specific and more
general-purpose, as it is happening with language models.</p>
      <p>For metrics to be really useful they have to be meaningful. Meaningful assessment must
measure well defined capabilities that translate into predictable skills and applications. Here
many lessons can be taken from psychology. Current performance indicators are used to
compare AI systems for the same benchmark or domain, and leaderboards are used to extrapolate
progress in the field. But because most of these benchmarks are not driven by theory, they have
very limited construct validity or generalisability. For this reason, there is still much uncertainty
over how to assess and monitor the state, development, uptake and impact of AI as a whole,
including its future evolution, progress and capabilities.</p>
      <p>The IJCAI-ECAI-22 Workshop on Artificial Intelligence Evaluation Beyond Metrics (EBeM
2022) seeks to challenge the widespread yet limited approach of evaluating the performance
of intelligent systems with aggregated metrics over a benchmark or distribution of tasks. In
particular, we discuss further alternative approaches that draw on ideas and recent progress in
cognitive and developmental psychology, psychometrics, software testing, and other areas. The
following are the main topics of the workshop:
• Evaluation methods founded on cognitive, developmental or comparative psychology
• Measurement of skills, capabilities, or cognitive abilities
• Evaluation methods based on software testing or other engineering practices
• Meta-analysis or comparisons of evaluation instruments
• The role of evaluation in AI development, policy making, and modeling of social impact
• Measurements of generality or common-sense
• Capture and use of evaluation data
• Analysis of the task space and its relation to corresponding capabilities
• The role of causality in evaluation
• Topics complementary to evaluation such as documentation or auditing
• Alternative evaluation methods with added benefits
• Discussion and progress in hard to evaluate scenarios</p>
      <p>These topics aim to achieve a holistic view of AI evaluation, targeting people from cognitive
science, comparative psychology, neuroscience, psychometrics, philosophy of science and
technology, measurement theory, policy, etc., with the ultimate goal of encouraging
crossdisciplinary approaches and theoretical and experimental analysis of how AI evaluation should
be done at present and in the future.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Program</title>
      <p>The Program Committee (PC) received 16 submissions. Each paper was peer-reviewed by at
least two PC members (with an average of 3 reviewers), by following a single-blind reviewing
process. The committee decided to accept 13 full papers, of which 3 papers elected not to be
published in this proceedings. The EBeM 2022 program was organized in to four sessions, with
two invited speakers, two panels, and a special session by the OECD.</p>
      <p>• Technical Session 1
– Invited speaker: Amanda Seed
– Paper presentations
• Technical Session 2
– Paper presentations
– Panel: Cognitive Evaluation with the Animal AI Environment
• Technical Session 4
– Invited speaker: Adina Williams
– Special OECD Session: Artificial Intelligence and the Future of Skills (AIFS)
– Paper presentations
– Panel: Evaluating pre-trained, generative, and prompted systems</p>
    </sec>
    <sec id="sec-3">
      <title>3. Acknowledgements</title>
      <p>We thank all researchers who submitted papers to EBeM 2022 and congratulate the authors
whose papers were selected for inclusion into the workshop program and proceedings. We
thank all invited speakers and panel members, and our distinguished PC members for reviewing
the submissions and providing useful feedback to the authors:
• Sean Holden - University of Cambridge
• Sebastian Gehrmann - Google Research
• Songul Tolan - European Commission, JRC
• Tadahiro Taniguchi - Ritsumeikan University
• Vicky Charisi - European Commission, JRC</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>