<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Learned Sparse Representations of Text: A Story of Indexing and Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Keynote</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Franco Maria Nardini</string-name>
          <email>francomaria.nardini@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>He is a member of the editorial board of ACM TOIS and a PC member of SIGIR</institution>
          ,
          <addr-line>ECIR, SIGKDD, CIKM</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>In recent years, transformer-based large language models (LLMs) have fundamentally reshaped the way large textual collections are indexed and retrieved. A key driver of this transformation is the use of LLMs to learn high-dimensional, contextual sparse representations of input text. These representations are rapidly gaining popularity for several reasons: 1) they perform competitively with learned dense representations, 2) they are grounded in the LLM's vocabulary, enabling interpretability by design, and 3) they can be eficiently leveraged with a well-established data structure - the inverted index - to support fast maximum inner product search. In this talk, we will review recent advancements that enable eficient indexing and retrieval based on these sparse representations. We will then discuss current limitations and emerging challenges in this rapidly evolving area. Franco Maria Nardini is a Senior Researcher with ISTI-CNR in Pisa, Italy. His research interests focus on Web Information Retrieval and Machine/Deep Learning. He authored over 100 papers in peer-reviewed international journals, conferences, and other venues. He has served as General Co-Chair of ECIR the ECIR 2014 Best Demo Paper Award. He coordinated activities in several EU and IT research projects. WSDM, IJCAI, and ECML-PKDD. He currently teaches ”Information Retrieval” in the Computer Science and AI Master's Degrees of the University of Pisa.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>LGOBE
rOcid
CEUR</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>