<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Topological Data Analysis of Navigation Paths ⋆ within Digital Libraries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bayrem Kaabachi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simon Dumas Primbaul</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bibliothèque nationale de France (BnF)</institution>
          ,
          <addr-line>Quai François Mauriac, 75706 Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Biomedical Data Science Center (BDSC), Centre Hospitalier Universitaire Vaudois (CHUV)</institution>
          ,
          <addr-line>CH-1002 Lausanne</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Laboratory for the history of science and technology (LHST), Swiss Federal Institute of Technology (EPFL)</institution>
          ,
          <addr-line>CH-1015 Lausanne</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>OpenEdition (UAR 2504, CNRS/EHESS/AMU/AU)</institution>
          ,
          <addr-line>22 rue John Maynard Keynes, 13013 Marseille</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <fpage>111</fpage>
      <lpage>134</lpage>
      <abstract>
        <p>The digitization of library resources and services have opened up physical informational spaces to new dimensions by allowing users to access a wealth of documents in ways that di昀er from browsing bookshelves traditionally organized according to the ”tree of knowledge”. How do readers of digital library orient themselves within big corpora? What landmarks do they use to navigate masses of digital documents? Taking Gallica as a case study-the digital heritage platform of the French national library-, this paper presents an experimental research on the navigation practices of its users. Using methods from topological data analysis, we inferred from Gallica's server logs an informational space as it is roamed by readers. Coupled with user interviews, this mixed-methods study allowed us to identify a set of ”regimes of navigation” characterizing how readers deploy various strategies to browse the digital library's corpus. From directed search to wandering to crawling, these regimes answer di昀erent needs and show that a single corpus can, in turns, be apprehended as a heritage collection, a database, a set of documents, and a mass of information.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;digital library</kwd>
        <kwd>navigation practices</kwd>
        <kwd>topological data analysis</kwd>
        <kwd>information retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>1.1. Research question</title>
        <p>
          The birth and development of digital libraries–broadly understood as curated collections of
electronic documents accessible online on dedicated platforms with tools for search and
consultation [
          <xref ref-type="bibr" rid="ref5">6</xref>
          ]–have radically transformed research activities. From any computer with
Internet access, the wealth of information available allows for the simultaneous consultation of
a great variety of resources and for their continual rearrangement into renewed information
landscapes. Consequently, traditional practices observed in physical places of knowledge such
as libraries and archives–searching catalogues, browsing through shelving, taking notes–have
been supplemented with a series of digital practices developed by researchers to browse
websites and databases–searching by keywords, 昀椀ltering results, navigating through links.
        </p>
        <p>While historically science was, and to a great extent still is, arranged into arborescent
taxonomies, scholars have in practice always negotiated with this normative order of knowledge,
resulting in a complex landscape much thicker than a tree-like structure. In contributing to the
digitization of information practices, digital libraries are further redrawing this landscape and
its relation to established orders of knowledge. The appropriation of digital libraries by their
users opens the “tree of knowledge” to a multiplicity of other dimensions allowing for leaps
from one “leaf” to another, as well as a continuous recon昀椀guration of branches of knowledge.
Once we become aware that the practical landscape of knowledge is evolving into a much more
dynamic arrangement, the use of spatial concepts to understand research practices from the
point of view of scholars becomes essential.</p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2. Literature review</title>
        <p>Research on digital information practices attracted data science in the late 2000s, focusing more
speci昀椀cally on “interaction traces” (understood as the sequential records of a user’s interactions
with one or across multiple platforms). Aiming at collecting, quantifying, and modelling
online traces, the investigation of interaction traces helped emphasize the role of navigation, as
opposed to search, in seeking information: on the one hand, “the perfect search engine is not
enough” therefore prompting user to brows2e2[], while, on the other hand, existing
recommendation algorithms do not favour navigability, even hindering1i4t][.</p>
        <p>
          Yet, inheritors of a rather mechanical model supported by “information foraging theory”,
these studies are still too o昀琀en dedicated to practical optimisation and problem-solving—aimed
at enhancing the ranking of pages found through “post-query navigation5”][, [
          <xref ref-type="bibr" rid="ref19">20</xref>
          ], predicting
the next page based on statistical regularities19[], [
          <xref ref-type="bibr" rid="ref12">13</xref>
          ], or suggesting the most beaten “trails”
[
          <xref ref-type="bibr" rid="ref22">24</xref>
          ]. Furthermore, when users are brought to the fore by data science and UX studies, the
emphasis is usually put on “directed searchi”.e–., shortest, directed, and local paths–, thereby
neglecting a whole part of navigation strategies: crawling, exploratory searches, serendipity...
        </p>
        <p>
          When eschewing this way more traditional ethnographic approaches based on interviews
and on-site observations, navigation is deemed worthy of practice and study only for its end
products or its “waypoints”2[5], not as a process in itself, and it is assessed according to
relevance only, rather than discovery or originality. Most quantitative studies on navigation
therefore bypass altogether a practice that their authors nonetheless underlined as
fundamental: search engines, however accurate and powerful, never fully satisfy users who tend to rely
more on step-by-step contextual navigation. Only recently studies have been conducted on
user navigationper se, especially on Wikipedia1[
          <xref ref-type="bibr" rid="ref7">8</xref>
          ].
        </p>
        <p>
          Within the realm of computational humanities, the recent development of topological data
analysis (TDA) calls for a reappraisal of navigation through the use and honing of tools
speci昀椀cally devoted to the qualitative study of shape. The shape of the World Wide Web is an issue as
old as the Internet itself and still widely debated to this d1a2y][, [
          <xref ref-type="bibr" rid="ref9">10</xref>
          ]. Traditionally, the Web
has been modelled as a graph. Although this has proven to be a powerful tool to study digital
communication, graphs are intrinsically limited to model pairwise interactions. The recent
success of topological methods in studying data, and the parallel establishment of topological data
analysis as a 昀椀eld [
          <xref ref-type="bibr" rid="ref8">9</xref>
          ], have con昀椀rmed the utility of viewing data through a higher-dimensional
analogue of graphs. Providing the tools to model navigation as a rich and thick process, rather
than as a problem to solve, TDA o昀ers the possibility to understand it in a less systematic, more
descriptive way, if coupled with humanistic methods addressing navigation in its lived
practical thickness. Although TDA has been used to analyse scienti昀椀c collaborations17[], nothing
properly topological has been endeavoured for navigational practices yet.
1.3. Case study and methodology
to understand how digital library users navigate online content, we chose to study Gallica
(https://gallica.bnf.f)r, the online platform of the French national library (BnF). Gallica
preserves and provides access to ten million documents in the public domain, freely available
either for download or for consultation on the dedicated online reader. The collection gathers
a wide variety of document types (printed books, press, manuscripts, musical scores, maps,
images, videos, objects...) in more than ten languages, deriving either from the BnF preservation
policies and digitisation campaigns, from other libraries, or fromdtéhpeôt légal.
        </p>
        <p>
          Previously, Nouvelleet al. [
          <xref ref-type="bibr" rid="ref15">16</xref>
          ] already provided an exploration of Gallica 2016 server logs,
focusing on the modeling of user sessions as Markov chains and providing 昀椀rst results about
typical modes of engagement with resources, their provenance, or the mediation e昀ect of
Gallica’s blog posts. More recently, Trabel2si3][ used o昀-the-shelf process mining techniques on
these same logs to model users’ paths between search pages, document viewing, blogs...
        </p>
        <p>
          Our approach for the whole project extends Beaudoueint al.’s mixed methods [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
by weaving together a socio-ethnography of users’ practices (le昀琀 of 昀椀g. 1)–based on
semistructured episodic interviews with seven Gallica users–and a computational ethnography
(right of 昀椀g. 1)–a topological analysis of server logs understood as interaction traces. These
two methods need to be cautiously dovetailed to yield interpretable results: the models used
to reconstruct and cluster reading paths from the logs were based on users’ testimonies, while
the resulting visualisations have been submitted to interviewees for validation or as probes to
incentivise discussion.
        </p>
        <p>The main sociological results of this 昀椀rst study are presented in8][and a semiotic analysis
of the visualisations generated is presented in7][.</p>
        <p>The present paper endeavours to shed light on the Python pipeline used to process Gallica
server logs as part of a computational ethnography of users’ navigation practices within digital
library. The code is available under a GNU licensehotntps://github.com/Kaabachi/TDA-Galli
ca.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Dataset and Sessionization</title>
      <p>In this study, we work with the browsing logs of Gallica over the month of April 2101A6.
description of the features contained in the dataset can be seen in Tabl1e.
1Compliance with GDPR, both for the qualitative and the quantitative approach, was approved by decision HREC
No. 056-2020 of EPFL ethical committee.</p>
      <p>Notably, certain 昀椀elds such as country, city, or referring website may lack speci昀椀city due to
incomplete information in the requests. These 昀椀elds are subsequently populated with a ”null”
value in cases of information absence.</p>
      <p>The focus of the project is on exploring user interactions and navigation patterns within
Gallica’s digital platform. As such, the raw data requires augmentation with relevant
documentspeci昀椀c details, considering the inherent limitations of the original dataset.</p>
      <p>To enhance the data, we utilize ARK-type requests. These requests act as document
identi椀昀ers within Gallica’s system but do not provide semantically interpretable information about
the respective documents. To address this, we systematically extract the unique ARK name
associated with each request. Following this, we employ Gallica’s dedicated service, designed to
extract bibliographical information corresponding to a document, to query the digital platform,
as illustrated in Figur2e.</p>
      <p>To render our data analysis more informative, we supplement our dataset through the
execution of these queries. This approach yields additional features of the accessed documents,
such as the title, the year of publication, the primary language, and the prevailing theme. This
enhancement of our dataset deepens our understanding of user interactions within Gallica.</p>
      <p>In our e昀ort to accurately model a user’s journey, akin to navigating a traditional library,
we place signi昀椀cant emphasis on the topic, or discipline, of each document accessed by the
user. Gallica utilizes the Dewey Decimal Classi昀椀cation system, a widely accepted numerical
taxonomy with established hierarchical categories. At a broad level, documents are classi昀椀ed
into ten primary classes, based on the 昀椀rst digit, re昀氀ecting fundamental disciplines or 昀椀elds of
study. Delving deeper, the second digit enables a more nuanced categorization. For instance,
610 is assigned to general works on medicine and health, 6i1s1associated with human anatomy,
612 designates human physiology, and 613corresponds to personal health and safety. This
granular classi昀椀cation o昀ers a comprehensive and systematic structure for navigating the array
of documents.</p>
      <p>
        Once we obtain the core data needed for the study, we separate requests made by distinct
users to properly characterize their interactions with the digital platform. Previous work by
Nouvelletet al. [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ] introduced methods to properly distinguish users through the parsing of
Gallica logs. In this work, we adopt a similar approach to analyzing users’ behavior through
the concept of “sessions”. A session refers to a single, continuous period during which a user
is actively engaging with Gallica’s web platform. At the end of this process, we aim to 昀椀lter
the logs so that we only look at active human use2r.s
      </p>
      <p>The technique we used to achieve this relies on applying several 昀椀lters and transformations
to the dataset. We 昀椀rst introduced an inactivity threshold of 60 minutes; the absence of any
user query a昀琀er this duration is interpreted as session termination. Subsequent queries by the
same user a昀琀er this interval are considered distinct sessions. We then assign those queries
di昀erent session numbers according to the hashed IP address linked to it. Finally, we drop the
sessions that contain no ARKs at all as they contain no topical information. Note that this is a
major limit of this study as only printed documents and prints have ARKs.</p>
      <p>
        Following the aforementioned transformations, we analyze the duration of each session,
speci昀椀cally focusing on sessions that contain a minimum of three ARKs (Archival Resource
2Specifying the human part is important, as we want to avoid potential interference from internet bot crawler data.
Keys). A predominance of Gallica users tend to consult only a single document as indicated
in prior research1[
        <xref ref-type="bibr" rid="ref5">6</xref>
        ][23]. Consequently, sessions involving the consultation of just one
document hold less relevance for our study, which is centered around exploring user transition
behaviors. Therefore, these sessions are not included in our analysis. Additionally, we
introduce an upper limit, 昀椀ltering out sessions that incorporate more than 50 consecutive ARKs.
This step ensures we manage extreme cases and outliers in our dataset that could potentially
skew the analysis.
      </p>
      <p>The selection criteria of a minimum of three ARKs and a maximum of 50 ARKs in a session
thus ensure a focused and meaningful investigation of user behavior patterns. Future work
might consider di昀erent criteria based on the research question at hand or employ di昀erent
statistical approaches to handle sessions with varying numbers of ARKs.</p>
      <p>= [</p>
    </sec>
    <sec id="sec-3">
      <title>3. From word embedding to TDA: An Integrated Model</title>
      <p>to model users’ pathways through Gallica, we adopted three di昀erent approaches, each
capturing a unique dimension of user interactions.</p>
      <p>The 昀椀rst approach employed the word2vec algorithm, treating each Dewey class as a word.
We created a corpus where the sentences were the transitions between classes in a user session.
The concept behind this representation was to learn the relationships between classes akin
to context relationships between words in natural language processing. While this approach
was useful in understanding immediate relationships and similarities between themes, it didn’t
consider the global structure or topology of theme interactions and the chronological order of
class visits.</p>
      <p>The second method was a network approach where we constructed a graph with classes
as vertices and transitions between classes as edges. Using betweenness centrality, we
identi昀椀ed popular classes based on their position and frequency of appearance in users’ journeys.
This method addressed some limitations of the word2vec approach by incorporating the
directionality of transitions between classes and providing a global view of class interactions.
Nevertheless, it only captured one type of global structure and was limited in its ability to
detect subtler topological patterns.</p>
      <p>The third, and most novel approach, was to employ Topological Data Analysis (TDA). With
TDA, we capture and quantify high-dimensional structural information about the users’
journey and e昀ectively map out a ”topological 昀椀ngerprint” of their interactions with classes. We
use persistent homology, a tool in TDA, to track the creation and destruction of connected
components, loops, and voids o昀ering a unique multiscale perspective of the data. The TDA
approach allows us to observe the existence of intricate patterns and structures that other
methods might overlook. The combination of these three methods provides a comprehensive view
of user behavior and class interactions in Gallica.</p>
      <sec id="sec-3-1">
        <title>3.1. Word2vec representation</title>
        <p>In this part, we aim to develop a metric that distinguishes one class from another, analogous
to the organisation of a physical library.</p>
        <p>The establishment of this metric relies on a word embedding representation, namely
word2vec. We regard each session as a phrase, where every visited class represents a word.
A corpus is constructed with each sentence signifying the class transitions in a session,
resembling the output produced by a Bag-Of-Words model applied to a document.</p>
        <p>Corpus



1   
2   
3   
1,   
5,   
6,   
2,   
2,   
9,   
3,   
7,   
7,   
1,   
8,   
2,   
4...
4...
1...</p>
        <p>Employing word2vec, we discern the relationships among the classes as it embeds words
in a lower-dimensional vector space. The result is a collection of word vectors with similar
meanings for vectors proximate in vector space, and dissimilar meanings for vectors distant in
the space, based on their context. Our implementation uses the skip-gram model. This model
selects word pairs by moving a window across the text data and trains a one-hidden-layer
neural network. For a window of si zeand a centered word  , we predict the context words
(or in our case themes){  }, ( −  ≤  ≤  + ,  ≠ ) .</p>
        <p>The cost function for one target word minimizes the negative log-likelihood of the target
word vector given the associated predicted word, formulated as follows:
ℒ
(, ) =</p>
        <p>∑
−≤≤+,≠
− log (  |  )
(3)</p>
        <p>As a result, we identify classes closest to each others according to our constructed corpus.
The top-N most similar keys are determined by computing the cosine similarity between a
simple mean of the projection weight vectors of the given keys and the vectors for each key in
the model. Positive keys contribute positively towards the similarity, negative keys contribute
negatively.3</p>
        <p>As an illustration, let us consider the classes that we found closest to ”Bible”:
Closest classes to ”Bible” Similarity
History, geographic treatment, biography of Christianity 0.758
Other Literatures 0.611
Italian, Romanian and related languages 0.589.</p>
        <p>Christian practice and observance 0.545
Earth sciences and geology 0.540
Religion 0.536</p>
        <p>Epistemology 0.535</p>
        <p>To conclude on this 昀椀rst approach, the use of word2vec embedding allowed us to construct
a high-dimensionality metric space within which Dewey classes are represented closer to each
other when they are more frequently consulted sequentially by users, and reciprocally. Later
on, this space will allow us to draw navigational path and characterize them according to class
interactions.
3To visualize the output of the word2vec model, we proceed with a projection of our vectors from a 200-dimensional
space to a 3-dimensional one using t-SNE, a t-distributed stochastic neighbor embedding method (see AppenCdix
and https://kaabachi.github.io/TDA-Gallica)/.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Network representation</title>
        <p>To extract further insights from our data, we propose an alternative representation via a social
network-like structure. All classes are added to the graph as vertices, with each transition from
one class to another in sessions represented as edges.</p>
        <p>From this representation, we employ graph theory to calculate the betweenness centrality of
centrality of a nod e is given by the expression:
each class and thereby comprehend the ”in昀氀uential” classes in the network. The betweenness
  ( ) =
∑
≠≠
  ( )


,
(4)
where  is the total number of shortest paths from nodesto node and   ( ) is the number
of those paths that pass through . This formula measures for a given no deits importance in
terms of circulation through the network: the higher, the more the node connects other nodes
and is used as a step to navigate from one class to another.</p>
        <p>The most ”popular” classes we obtained can be seen in Tabl2e.</p>
        <sec id="sec-3-2-1">
          <title>Top 5 classes according to betweenness centrality metric</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>Classes</title>
        </sec>
        <sec id="sec-3-2-3">
          <title>History of Europe</title>
        </sec>
        <sec id="sec-3-2-4">
          <title>French and related literatures</title>
        </sec>
        <sec id="sec-3-2-5">
          <title>Latin and Italic literatures</title>
        </sec>
        <sec id="sec-3-2-6">
          <title>The arts</title>
        </sec>
        <sec id="sec-3-2-7">
          <title>News media, journalism, and publishing</title>
        </sec>
        <sec id="sec-3-2-8">
          <title>Betweenness centrality</title>
          <p>directed algorithm. This algorithm arranges the nodes of a graph in two- or three-dimensional
space such that all edges are roughly of equal length, and crossing edges are minimized (see
椀昀g. 3).</p>
          <p>This second approach shows that the metric space of Dewey classes is not 昀氀at. Beyond
the fact that some classes are very central and others peripheral due to being more or less
consulted, the network approach also allowed us to start identifying which classes may act
as ”pivots,” helping users to transition between dissimilar or farther classes. This inquiry was
pursued parallel to the next subsection and is exposed in appendBi.x</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Topological Representation</title>
        <p>Although the previous approaches outlined in sectio3.n1 and 3.2 provided us valuable insights
on how users interface with a digital library, the approaches were limited when it comes to
providing a deeper understanding of the higher-level view of these sessions and how users
navigate the website.</p>
        <p>In this work, we use Topological Data Analysis (TDA) to reveal the topological features that
might have been obscured in the original vector space. We supplement the analysis done using</p>
        <sec id="sec-3-3-1">
          <title>Fruchterman-Reingold [11] force-directed algorithm</title>
          <p>word2vec and combine both methods to provide a more comprehensive understanding of users’
interactions with Gallica.</p>
          <p>This novel approach borrows concepts from the realm of topological algebra to enable the
extraction of deeper features that describe the topology of users’ navigation paths in this
digital library. The synergy and interaction between both the more traditional machine learning
approach and the topological data analysis one can be seen in 昀椀gur4e. The TDA applied in
this study was executed using the giotto-tda library, an open-source Python toolbox designed
for topological machine learnin2g1[].</p>
          <p>The input to this pipeline is a 昀椀nite set of points called point clouds, where each point
represents a Dewey class that a user has interacted with in a session, and the points are characterized
by a metric that denotes the distance or similarity between them. The metric space that de昀椀nes
those points is the result of earlier word2vec transformations applied on the classes visited by
users.
[</p>
          <p>, 
We denote a session as a point cloud, where each point is a pair of coordinates

 ] , signifying the embeddings for each visited class in the session. The relation
between these three components can be expressed as follows:
Topological</p>
          <p>Data
Analysis
Pipeline</p>
          <p>Compute
sessions
Construct a
simplicial
complex using
Viertoris-Rips</p>
          <p>Complex</p>
          <p>Establish a
metric space
using word2vec</p>
          <p>Create
2dimensional
point clouds
using T-SNE</p>
          <p>Compute clusters on
the resulting feature
matrix using K-means
and silhouette method</p>
          <p>Extract the
closest points to
the centroids for</p>
          <p>visualization</p>
          <p>Vary the
parameter α during
the construction to
study the changes
in topology</p>
          <p>Extract
topological
invariants - Betti
Numbers - for</p>
          <p>each
construction</p>
          <p>Encode the resulting
Betti Numbers in a
persistence diagram
corresponding to a
session.</p>
          <p>Transform the
persistence
persistence diagram
into a feature matrix
using persistence
entropy</p>
          <p>Point Cloud ⟺ (Class1 , Class1 )...(Class  , Class  )
(5)
where each pair [Clas s , Class  ] corresponds to a speci昀椀c class visited during a session.</p>
          <p>The fundamental principle behind TDA1[5]is to construct a continuous shape, also known
as a simplicial complex atop the point clou4d.We construct simplicial complexes from our
dataset using a speci昀椀c method known as the Vietoris-Rips construction. The input set of our
metric space is derived from the result of a Word2Vec transformation on the document themes,
producing a multidimensional vector spa(c ,e) . For a non-negative real numbe r, we de昀椀ne
the Vietoris-Rips complex  () as the set of simplices[ 0, … ,   ] such that (  ,   ) ≤  for all
,  .</p>
          <p>De昀椀nition 1. (Vietoris-Rips Complex). Let  be a subset of a metric space ( , ) , where  is
the result of a Word2Vec transformation applied to the document themes. For a non-negative real
number  , the Vietoris-Rips complex    ( ) is de昀椀ned as follows: The vertices are points in  ,
and for each subset { 0, … ,   } of  , we include a  -simplex [ 0, … ,   ] if and only if (  ,   ) ≤ 
for all ,  .</p>
          <p>This construction has the e昀ect of building higher-dimensional simplices (triangles,
tetrahedra, and their higher-dimensional analogues) on top of the dataset whenever groups of points
are closer to each other than the speci昀椀ed distance. This method provides a versatile way to
uncover the hidden geometric and topological structure embedded in the dataset.</p>
          <p>Within this pipeline, we extract valuable topological information from the simplicial complex
using a method called persistent homology. We vary theparameter during the Vietoris-Rips
construction to study the changes of topology in a scale-invariant perspective. We then extract
the corresponding topological invariants - also called Betti Numbers. Those betti numbers are
depicted in a persistence diagram that forms a ”barcode” as can be seen in 昀椀gur5e where the
length of this barcode is indicative of the persistence of the homological feature.</p>
          <p>We then transform the persistence diagrams, which depict the topological structure of the
data set, into a feature matrix. The ensuing matrix, captures the entropy distribution of points
4Simplicial are sets of simplices that adhere to certain conditions, enabling a broad representation of geometric and
topological properties of the data.
across each persistence diagram, thus summarising the underlying topology in a compact and
insightful format. This is done using persistence entropy.</p>
          <p>Finally, we subject resulting matrix to k-means clustering. We employ this unsupervised
learning method to separate the transformed data into eight distinct clusters, each
representing unique topological features discovered through TDA. The selection of eight clusters was
guided by the examination of silhouette scores. These metrics provided a robust estimation
of the clustering quality, with higher values (close to 1) indicating well-separated and
cohesive clusters. The clustering process achieved an overall silhouette score of 0.70, indicating a
reasonable separation between clusters.</p>
          <p>To summarize, the use of topological data analysis allowed us to cluster users’ navigational
paths according to their geometric features once drawn within the metric space of Dewey
classes previously constructed. The core idea here was to identify typical user navigation
behaviours on Gallica based on the shape of their path through the available corpus. The results
are presentend in the following section.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>Clustering navigation paths according to their shape within a disciplinary space inferred from
the embedding of Dewey classes allowed us to identify a number of ”navigation regimes”
depending on a number of variables and accounting for a variety of orienteering practices within
the informational space that is Gallica’s corpus. We identi昀椀ed four main regimes through
clustering and a 昀椀琀h one by a simple re昀氀exive argument (see table below). Note that the regimes
described here were interpreted by crossing users’ testimonies with the shape and properties
of the paths closest to the clusters’ centroids. These regimes are ideal-typical in the sense
that they represent archetypal user behaviour and are seldom observed in their purest form.
Rather, users juggle between regimes depending on their needs, their skills, and their strategies,
thereby exhibiting navigation patterns that lie farther from the centroids, albeit with dominant
features pertaining to speci昀椀c regimes.</p>
      <p>• Directed search: very short paths spanning a limited number of Dewey classes, directed
searches using very speci昀椀c keywords and 昀椀lters in the search engine are made by users
who know precisely what they are looking for, either a speci昀椀c reference, or a piece of
information, and want to retrieve it through the quickest, shortest, and most e昀케cient
path–ideally a straight line from query to result;
• Constitution or consultation of a corpus: when a user tries to gather all available
references and information on a given 昀椀eld that is well de昀椀ned by a combination of
variables including a time period, a geographical area, one or multiple topics, one or
multiple types of documents..., they tend to consult more documents for a longer time to
make an exhaustive survey within a very speci昀椀c discipline–usually less than 2 Dewey
classes;
• Star-shaped radiation: as was well described by our interviewees, some users start
from one or two main Dewey classes and try to establish links with other perspectives
on the same subject by foraying into adjacent classes–these paths tend to be even longer
and to spread to more classes either in a comparative manner or in search for a diversity
of views;
• Wandering: although some interviewees forget to mention this regime or actively
discard it (and then contradict themselves), wandering through Gallica’s corpus is quite a
common practice that knows no disciplinary boundary and usually takes users along
longer navigation paths spanning a greater number of Dewey classes, either for mere
entertainment or to actively set the conditions for serendipity and the discovery of
unexpected documents, information, ideas, or concepts;
• Crawling: 昀椀nally, and although our method did not allow us to catch this kind of regime
(due to the fact that we decided to catch only human behaviour), we did crawl Gallica’s
corpus with the use of an algorithm to retrieve, reliably and in signi昀椀cant quantities, the
metadata of the requested documents.
Regime
Duration
Mean nb of doc.</p>
      <p>Dewey reach
Documentary unit
Values
Regime
Duration
Mean nb of doc.</p>
      <p>Dewey reach
Documentary unit
Values
Regime
Duration
Mean nb of doc.</p>
      <p>Dewey reach
Documentary unit
Values</p>
      <p>Directed search</p>
      <p>Very short</p>
      <p>6 to 7
Limited (2 to 3)</p>
      <p>Information
or document
Relevance,
e昀케ciency,
椀昀ndability</p>
      <p>Constitution of a corpus</p>
      <p>Long
7 to 10
Very limited (1 to 2)</p>
      <p>Corpus,
references</p>
      <p>Exhaustivity
Star-shaped radiation</p>
      <p>Short to very long</p>
      <p>20
Extended (10)</p>
      <p>Topic
or Field
Comparison,
diversity</p>
      <p>Wandering
Medium to long</p>
      <p>More than 20
Very extended (more than 10)</p>
      <p>Unexpected documents,
ideas or concepts</p>
      <p>Serendipity,
discoverability</p>
      <p>Finally, we also identi昀椀ed paths of users getting lost and then 昀椀nding the right path again.
For example, one use looking for casts in the history of art may end up consulting a document
on dental casts among the list of results. Further research, both quantitative and qualitative,
will help us identify other meaningful navigation regimes.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>The main conclusion of this study is a proof of concept of ”navigation paths”. Given that Gallica
is 昀椀rst and foremost a service dedicated to providing digital documents to its users, mostly
through the use of a single search engine (thereby fostering its consultation as a mere ”database”
to query), it was not obvious at 昀椀rst glance that readers would engage in navigating or browsing
through the platform, from document to document. Yet, the wealth of spatial, mechanical,
entertainment, play metaphors mobilised by our interviewees to try and objectify the way
they use Gallica, coupled with the robustness and meaning of the navigation paths that we
modeled based on their testimonies, we observed that a signi昀椀cant amount of users engage in
longer sessions, consulting multiple documents, sometimes at length, in diverse classes. These
informational practices are made possible by the size and diversity of Gallica’s corpus that
allows for comparisons, discoveries, serendipity...</p>
      <p>Indeed, further ongoing study shows that these users somewhat ”hack” the principles of the
search engine by iterating queries in a manner that allows them to manage the noise present in
the list of results. Thereby, using the main tool at their disposal, they construct their navigation
path step-by-step according to di昀erent strategies either taking the scenic route or radiating
from one or several poles. Nonetheless, all our interviewees explained that they were almost
always working ”in paralleli.”e,. juggling between tabs, platforms, and physical content. If this
issue was partly addressed by using a word2vec embedding that takes into account the
nonlinearity of navigation paths, we have to acknowledge that the paths studied are but a narrow
view on much thicker and complex research practices, amounting the ”digital work昀氀ow2”][to
a kind of ”bricolage”1[].</p>
      <p>From a methodological point of view, this study shows how a qualitative and a quantitative
approaches can be precisely dovetailed so as to enrich each other. Indeed, the users’ discourses
were extremely important in de昀椀ning the model used to extract navigation paths, as well as
they played the role of safeguards in the manipulation of data. Reciprocally, the results and
visualisations obtained through quantitative analysis allowed us to reinterpret the users’
interviews and better understand their practices. Furthermore, the use of tools from topological
data analysis, which is quite novel in the 昀椀eld of computational humanities and social sciences,
proved particularly relevant to document and enact the informational space through which
digital library users navigate.</p>
      <p>Finally, these results will help digital libraries take into account the experience of their users
in very meaningful ways. First, the well-known imbalance of the Dewey classi昀椀cation is now
better documented with network analysis and an understanding of pivotal literature that can
help build new taxonomies, or even folksonomies, e.g. made of tags selected by users. Second,
the navigation regimes can help librarians design new navigation tools that could foster the
discoverability of lesser known content, thereby facilitating serendipity. Eventually, a whole
new set of information architectures could be devised and co-designed with users depending
on their needs and their research practices.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgments</title>
      <p>This exploratory study was funded by a CROSS grant awarded by EPFL and Unil.</p>
      <p>The second phase of the research began in October 2022 thanks to a Mark Pigott Fellowship
in Digital Humanities, awarded by BnF.</p>
      <p>We would like to thank Jeanne Fernandez of the DHLab for giving us a masterclass on TDA
at the onset of this project.</p>
      <p>Fruchterman
Online Library .</p>
      <p>M. Trabelsi. “Modélisation des processus utilisateurs à partir des traces d’exécution,
application aux systèmes d’information faiblement structurés”. PhD thesis. La Rochelle: La
Rochelle Université, 2022.</p>
    </sec>
    <sec id="sec-7">
      <title>A. ARK (Archival Resource Key)</title>
      <p>The implementation of ARKs on Gallica’s website is based on several principles, one of which
allows for the collection of metadata from an ARK. The structure of an ARK can be depicted
as follows:</p>
      <p>In our speci昀椀c case, the assigning authority number for Gallica’s website is consistently
12148. Consequently, the majority of the work revolves around utilizing the NAA-assigned
name to query the website and obtain the desired metadata.</p>
      <p>The subpipeline used to enrich the user sessions with additional metadata by querying
Gallica’s API with the corresponding ARKs is detailed in 昀椀gur7e below.
Gallica Raw Log</p>
      <p>Select only ARK
requests</p>
      <p>Extract
Additional</p>
      <p>Features
Features</p>
      <p>Advanced Features
IP - Country - City - Date - Request, Protocol
Answer code - Length - Referring website</p>
      <p>Document Title - Document Language - Publication</p>
      <p>Year - Theme
Perform OAI</p>
      <p>query</p>
      <p>Gallica Full Dataset</p>
    </sec>
    <sec id="sec-8">
      <title>B. Pivotal literature</title>
      <p>to better understand how users actually navigate between Dewey classes, we endeavoured to
model sessions as Markov chains. This allowed us to identify the type of literature mobilized
by readers to jump from branches to branches within the tree of knowledge.</p>
      <sec id="sec-8-1">
        <title>B.1. Sessions as Markov chains</title>
        <p>Sessions were modeled as Markov chains, a proven technique in literature to understand human
navigation. The key characteristic leveraged here is the memoryless property of Markov chains,
providing a robust framework for comprehending the structural mechanisms driving a user’s
journey in a digital library where each theme equates to a state. For a 昀椀rst-order Markov model,
the transition probability from one theme (state) to another is de昀椀ned as:
 ( ℎ
 , ℎ 
) = ( +1 =  ℎ
 |  =  ℎ
)</p>
        <p>The above equation signi昀椀es that the probability of transitioning to theme ’j’ from theme ’i’
is dependent solely on the present theme ’i’ and not on any previous themes. The transition
probability is estimated from the overall statistics of our dataset. If we ledtenote the number
of transitions from state ’i’ to state ’k’, then the estimated transition probability is given by:
 ̂ℎ
 , ℎ 
=
  ℎ</p>
        <p>This equation indicates that the estimated probability of transitioning from theme ’i’ to
theme ’j’ is the ratio of the count of transitions from ’i’ to ’j’ to the total number of
transitions from ’i’ to any other theme.
1/3
(6)
(7)
1
 ℎ
1
 ℎ
2
 ℎ
3
 ℎ
4
1
1/3
2/3
2/3
This model sheds light on the relation of users to the practical architecture of information in
Gallica. In the introduction, we were wondering whether the digitization of libraries actually
allowed users to jumped from one branch, or one leaf, of the ”tree of knowledge” to another,
thereby facilitating interdisciplinarity and a diversity of approaches. Indeed, the main
descriptors available at the French national library, hence on Gallica, to know about the topic of a
document are the Dewey classes. They are arranged into an arborescent structure, that is they
comprise a set of general classes making up for a partition of all possible topics, and are rami昀椀ed
into multiple levels of sub-classes. In physical libraries, the Dewey classi昀椀cation is frequently
used to order the books along the shelves, thereby structuring the whole informational space.</p>
        <p>On Gallica, most interviewees told us they would not use the Dewey classes for their searches,
sometimes explicitly despising them for various reasons. Yet, a quantitative analysis of
navigation paths modeled as Markov chains depicts a di昀erent situation. As can be seen on 昀椀g. 8,
the probabilities to transition from discipline A to discipline B is signi昀椀cant only if B is equal
to A, or if B is contained within the sub-classes of the general class of discipline B. This means
that overall, users tend to stay within the disciplinary boundaries of a single Dewey classes and
some of its sub-classes. This can be explained by a variety or factors, ranging from the users
disciplinary habits, to the majority of ”directed searches” (see below), to the search algorithm
which must favour results from close classes.
Besides the diagonal and dribbles around it, we can also identify some Dewey classes which
are more likely to be either transitioned-from or transitioned-to. Classes most
transitionedto (like 340 Law, 610 Medecine and health, or 840 French and related literatures) denote the
imbalance of the Dewey classi昀椀cation: they are all sub-classes and yet very important and
extensive 昀椀elds. They are classes towards which some users converge and, then, tend to remain
within since they cover ample informational space.</p>
        <p>Most transitioned-from classes (such as 900 History and geography, 940 History of Europe, or
070 News media, journalism and publishing) cover general topics, as well as news, collections,
or encyclopedias. They act as hinges or junctions between diverse other Dewey classes and
allow users to navigate from one branch of the classi昀椀cation to another. They are what we
called ”pivotal literature”.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>C. Word2vec projection in 3 dimensions</title>
      <p>An ongoing research led by one of the authors furthers and widens the exploratory study
presented in this paper. The same mixed-methods framework is currently deployed to document
users’ practices with an emphasis on action. Parallel to the present enquiry about how users
imagine or picture their navigation through Gallica, said research tries to shed light on how
users seize which elements of the interface and the search tools to concretely construct their
navigation paths. Qualitatively, the interviews realised with about 20 users emphasise how
they practically build their navigation paths step by step, by performing simple actions on the
interface – iterating queries, 昀椀ltering results, tweaking parameters... Quantitatively, the
analysis of the server logs, rather than relying on sequence of Dewey classes, is built upon a tree of
possible actions ranging from running a simple search, to turning the page of a document, to
downloading it... The confrontation between both studies shows how users wanting to wander
through Gallica’s corpus have to ”hack” its search engine by iterating slightly di昀ering queries,
thereby constructing their own scenic route step-by-step.</p>
      <p>Future work should focus on generating other descriptors than Dewey classes to qualify with
the relevant granularity the topics or disciplines of the documents consulted on the platform.
This could be done by doing topic modeling on the metadata and OCR of the corpus.
Furthermore, it would be interesting to have access to the navigation data of our interviewees so as to
use as probes their own navigation paths. Finally, it would be ideal to be able to follow users
across multiple tabs and websites to shed light on their cross-platform navigation patterns.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Antonijevic</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Cahoy</surname>
          </string-name>
          . “
          <article-title>Researcher as Bricoleur: Contextualizing humanists' digital work昀氀ows”</article-title>
          .
          <source>In: DH Quarterly 12.3</source>
          (
          <year>2018</year>
          ). url: http://www.digitalhumanities.org /dhq/vol/12/3/000399/000399.htm.l
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Antonijević</surname>
          </string-name>
          . “
          <article-title>Digital Work昀氀ow in the Humanities and Social Sciences: A Data Ethnography”</article-title>
          . In:
          <article-title>Anthropological Data in the Digital Age</article-title>
          . Ed. by
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Crowder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fortun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Besara</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Poirier</surname>
          </string-name>
          . London: Palgrave Macmillan,
          <year>2020</year>
          , pp.
          <fpage>59</fpage>
          -
          <lpage>83</lpage>
          .
          <year>d1o0i</year>
          .:
          <volume>1007</volume>
          /
          <fpage>978</fpage>
          -3 -
          <fpage>030</fpage>
          -24925-0\_4.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Beaudouin</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Denis</surname>
          </string-name>
          .Observer et évaluer les usages de Gallica.
          <source>Ré昀氀exion épisté- mologique et stratégique. Research Report</source>
          . Paris: Télécom ParisTech; BnF; Labex Obvil,
          <year>2014</year>
          . url: https://hal.archives-ouvertes.
          <source>fr/halshs-0107853</source>
          .
          <fpage>0</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
            <surname>Beaudouin</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Garron</surname>
          </string-name>
          , and
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Rolle'tJ.e pars d'un sujet, je rebondis sur un autre' : Pratiques et</article-title>
          usages des publics de Gallica.
          <source>Research Report</source>
          . Paris: Télécom ParisTech; BnF; Labex Obvil,
          <year>2016</year>
          . url:https://hal.archives-ouvertes.
          <source>fr/hal-0170923</source>
          .8 [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bilenko</surname>
          </string-name>
          and
          <string-name>
            <surname>R. W. White.</surname>
          </string-name>
          “
          <article-title>Mining the Search Trails of Sur昀椀ng Crowds: Identifying Relevant Websites From User Activity”</article-title>
          .
          <source>InW:</source>
          ww
          <year>2008</year>
          .
          <year>2008</year>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>60</lpage>
          . doi:
          <volume>10</volume>
          .1145/136 7497.1367505.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Borgman</surname>
          </string-name>
          .
          <article-title>From Gutenberg to the Global Information Infrastructure: Access to Information in the Networked World</article-title>
          . Cambridge, MA: MIT Press,
          <year>2000</year>
          . doi:
          <volume>10</volume>
          .7551/mitpress/31 31.001.0001.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S. Dumas</given-names>
            <surname>Primbault</surname>
          </string-name>
          . “
          <article-title>Documenter la navigation en bibliothèque numérique. Retour sur un chassé-croisé méthodologique entre qualitatif et quantitatif”</article-title>
          . In: (forthcoming
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. Dumas</given-names>
            <surname>Primbault</surname>
          </string-name>
          . “
          <article-title>Naviguer dans les savoirs à l'ère numérique. Pour une ethnographie des pratiques informationnelles sur Gallica”</article-title>
          .ÉItnu:des de communication
          <volume>61</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Edelsbrunner</surname>
          </string-name>
          and J.
          <source>HarerC. omputational Topology: An Introduction. Providence: American Mathematical Society</source>
          ,
          <year>2010</year>
          . url:https : / / www . maths . ed . ac . uk / ~v1ranick /papers/edelcomp.pd.f
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ghitalla</surname>
          </string-name>
          .
          <article-title>Qu'est-ce que la cartographie du web ? Expéditions scienti昀椀ques dans l'univers des données numériques et des réseaux</article-title>
          . Marseille: OpenEdition Press,
          <year>2021</year>
          . do1i:
          <fpage>0</fpage>
          .4000 /books.oep.
          <volume>15358</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <article-title>Graph Drawing by Force-directed Placement - 1991 - So昀琀ware: Practice</article-title>
          and Experience - Wiley https://onlinelibrary.wiley.com/doi/10.1002/spe.4380211102.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Huberman</surname>
          </string-name>
          .
          <article-title>The Laws of the Web: Patterns in the Ecology of Information</article-title>
          . Cambridge, MA: MIT Press,
          <year>2001</year>
          . doi:
          <volume>10</volume>
          .7551/mitpress/4150.001.0001.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Koopmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dallmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hettinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Niebler</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Hotho</surname>
          </string-name>
          . “
          <article-title>On the right track! Analysing and Predicting Navigation Success in Wikipedia”</article-title>
          .
          <source>InP:roceedings of the 30th ACM Conference on Hypertext and Social Media (HT '19)</source>
          .
          <year>2019</year>
          , pp.
          <fpage>143</fpage>
          -
          <lpage>152</lpage>
          . doi:
          <volume>10</volume>
          .1145 /3342220.3343650.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lamprecht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Geigl</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Karas</surname>
          </string-name>
          . “
          <article-title>Improving recommender system navigability through diversi昀椀cation: A case study of IMDb”</article-title>
          .
          <source>In:i-KNOW '15</source>
          .
          <year>2015</year>
          . doi:
          <volume>10</volume>
          . 1145 / 2 809563.2809603.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Murugan</surname>
          </string-name>
          and
          <string-name>
            <surname>D. RobertsonA.</surname>
          </string-name>
          <article-title>n Introduction to Topological Data Analysis for Physicists: From LGM to FRBs</article-title>
          .
          <year>2019</year>
          . arXiv:
          <year>1904</year>
          .11044 [astro-ph,
          <source>physics:hep-th].</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nouvellet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Beaudouin</surname>
          </string-name>
          ,
          <string-name>
            <surname>F. D'Alché-Buc</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Prieur</surname>
            , and
            <given-names>F.</given-names>
          </string-name>
          <article-title>RoueA昀.nalyse des traces d'usage de Gallica: Une étude à partir des logs de connexions au site Gallica</article-title>
          .
          <source>Research Report</source>
          . Paris: Télécom ParisTech; BnF; Labex Obvil,
          <year>2017</year>
          . urhl:ttps://hal.archives
          <article-title>-ouv ertes</article-title>
          .fr/hal-01709264.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Patania</surname>
          </string-name>
          , G. Petri, and
          <string-name>
            <given-names>F.</given-names>
            <surname>Vaccarino</surname>
          </string-name>
          . “
          <article-title>The shape of collaborations”</article-title>
          .
          <source>IEnP:J Data Science</source>
          <volume>6</volume>
          .18 (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .1140/epjds/s13688-017-0114-8.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Piccardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gerlach</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>West</surname>
          </string-name>
          . “
          <article-title>Going Down the Rabbit Hole: Characterizing the Long Tail of Wikipedia Reading Sessions”</article-title>
          .
          <source>InW:ikiWorkshop '22 Proc. of World Wide Web Conference (Companion)</source>
          .
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2203.06932.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sen</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Hansen</surname>
          </string-name>
          . “Predicting. Web Users'
          <article-title>Next Access Based on Log Data”</article-title>
          .
          <source>In: Journal of Computational and Graphical Statistics 12.1</source>
          (
          <year>2003</year>
          ). doi:
          <volume>10</volume>
          .1198/10618600312 75.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Singla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W.</given-names>
            <surname>White</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J.</given-names>
            <surname>Huang</surname>
          </string-name>
          . “
          <article-title>Studying trail昀椀nding algorithms for enhanced web search”</article-title>
          . In: Sigir'
          <fpage>10</fpage>
          .
          <year>2010</year>
          , pp.
          <fpage>443</fpage>
          -
          <lpage>450</lpage>
          . doi:
          <volume>10</volume>
          .1145/1835449.1835524.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G.</given-names>
            <surname>Tauzin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Lupo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tunstall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Caorsi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Reise</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Medina-Mardones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dassatti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Hess.</surname>
          </string-name>
          Giotto-Tda:
          <article-title>A Topological Data Analysis Toolkit for Machine Learning</article-title>
          and
          <string-name>
            <given-names>Data</given-names>
            <surname>Exploration</surname>
          </string-name>
          .
          <year>2021</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.
          <year>2004</year>
          .
          <volume>02551</volume>
          . arXiv:
          <year>2004</year>
          .02551 [cs, math, stat].
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Teevan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Alvarado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Ackerman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Karger</surname>
          </string-name>
          . “
          <article-title>The perfect search engine is not enough: a study of orienteering behavior in directed search”</article-title>
          .
          <source>CInh:</source>
          i
          <year>2004</year>
          .
          <year>2004</year>
          , pp.
          <fpage>415</fpage>
          -
          <lpage>422</lpage>
          . doi:
          <volume>10</volume>
          .1145/985692.985745.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>R. W.</given-names>
            <surname>White</surname>
          </string-name>
          and
          <string-name>
            <surname>J. Huang.</surname>
          </string-name>
          “
          <article-title>Assessing the Scenic Route: Measuring the Value of Search Trails in Web Logs”</article-title>
          . In:Sigir'
          <fpage>10</fpage>
          .
          <year>2010</year>
          , pp.
          <fpage>587</fpage>
          -
          <lpage>594</lpage>
          . doi:
          <volume>10</volume>
          .1145/1835449.1835548.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R. W.</given-names>
            <surname>White</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Singla</surname>
          </string-name>
          . “
          <article-title>Finding our way on the web: exploring the role of waypoints in search interaction”</article-title>
          .
          <source>In:Www</source>
          <year>2011</year>
          .
          <year>2011</year>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>148</lpage>
          . doi:
          <volume>10</volume>
          .1145/1963192.1963267.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>