<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Federated Information Retrieval in Cross-Domain Information Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sylvia Melzer</string-name>
          <email>sylvia.melzer@uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hagen Peukert</string-name>
          <email>hagen.peukert@uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eliana Dal Sasso</string-name>
          <email>eliana.dal.sasso@uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Charles Li</string-name>
          <email>charles.li@uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Asselborn</string-name>
          <email>asselborn@ifis.uni-luebeck.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralf Möller</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universität Hamburg, Centre for Sustainable Research Data Management</institution>
          ,
          <addr-line>Monetastraße 4, 20146 Hamburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universität Hamburg, Centre for the Study of Manuscript Cultures</institution>
          ,
          <addr-line>Warburgstraße 26, 20354 Hamburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Lübeck, Institute of Information Systems</institution>
          ,
          <addr-line>Ratzeburger Allee 160, 23562 Lübeck</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In humanities research projects, scholars examine written artefacts, such as manuscripts, for various purposes based on factors like language, textual content, provenance, codicological aspects, and other characteristics. While humanities scholars can make statements about diferent aspects of self-contained artefacts based on their expertise, there are instances where the statements are made without numerical verification due to limited research data that are available within a project. If a variable, e.g. the size of a book, is requested, classic search engines provide similar but not precise answers. Our thesis proposes that by combining various cross-domain information sources as a federated database system, these missing variables can be supplemented, thereby validating research questions in the humanities. The article proposes a cross-domain information system that enables eficient federated search for comprehensive research in the humanities. The system combines diverse information sources and provides eficient search capabilities by demonstrating an eficient data matching approach called indexing. This article also presents how users can define their queries in natural language by integrating GPT4all to generate SQL queries from natural language queries. The achieved result is a cross-domain information system that facilitates comprehensive research in the humanities by combining diverse information sources and providing eficient federated information retrieval.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Depending on the research interest, it may be that diferent researchers study the same
manuscript with a diferent focus or apply the same research question to written artefacts
pertaining to diferent manuscript traditions. Alongside traditional publications in journals and
monographs, research data about written artefacts can be found independently in digital
resources produced by research institutions, museums, or libraries. Increasing amounts of sources
residing in libraries and archives are digitized and made accessible in an RDR (Research Data
Repository) such as Zenodo [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or adjusted instances of it [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] like at the Universität Hamburg.
      </p>
      <p>
        In the project Beta mas.a¯ h. fet a collection of XML files, based on the TEI ( Text Encoding
Initiative) Guidelines [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], were created that describe textual and physical features of manuscripts
from Ethiopia and Eritrea. These TEI files had been published and are available online 1. In this
machine-readable format users usually do not have an overview of all written artefacts that
have the same property.
      </p>
      <p>
        Apart from that, project-specific web applications were built. In addition, users cannot
perform natural language queries to obtain required data from XML documents. Additional
tools are necessary to make XML data searchable. For this reason, we developed and used the
generic DBoD (DataBasing on Demand) process [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to transform research data from TEI files
to a database instance2, then an information system based on top of the database instance was
created. “An information system is an integrated set of components for collecting, storing, and
processing data and for providing information, knowledge, and digital products.” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
      </p>
      <p>In the project Bookbindings as Instruments of Classification at the Universität Hamburg a
collection of JSON files were created to document the binding technique used in Egypt from the
fourth to the twelfth centuries. JSON is also a machine-readable format, so we also created a
database instance3 using the DBoD process using JSON files as input and created an information
system based on top of the database instance.</p>
      <p>In the project Text-Surrounding-Text the research data about binding techniques in South
India are stored directly in the National library of France 4.</p>
      <p>While digital collections, like the three examples mentioned above, provide valuable data
for studying bookbinding techniques, it is important to note that addressing certain research
questions often necessitates the use of multiple information sources. For instance:</p>
      <p>Are there any similarities between the binding techniques, the object size or written area
dimension of manuscripts from Ethiopia, Eritrea, early Egypt, and South India?</p>
      <p>
        Evaluating and retrieving information from diverse sources and domains, FIR (Federated
Information Retrieval) in cross-domain information systems is a research area that focuses on
advancing the scholars of the humanities, both technically and methodologically, by integrating
diferent sources of data and evaluating them based on various criteria such as accuracy, currency,
and relevance. Access to the three diferent databases is realized through federated search.
Federated search is a technique for searching multiple collections simultaneously with a single
query. To make the search more eficient, we use the indexing method from Melzer et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In
the process, EpiDoc (Epigraphic Documents in TEI XML) files [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] (a customized version of TEI)
were used as input. In this article, we use TEI and JSON instead of EpiDoc. As a result of the
indexing process, we have a similarity score for each of the data sets, whereby only parts of
      </p>
      <sec id="sec-1-1">
        <title>1https://github.com/BetaMasaheft/Manuscripts</title>
        <p>2https://heurist.fdm.uni-hamburg.de/html/heurist/?db=CSMC_UWA_BETAMASAHEFT
3https://heurist.fdm.uni-hamburg.de/html/heurist/?db=CSMC_UWA_RFE09
4https://tst-project.github.io/mss/Sanscrit_1129.xml
the data, the so-called index candidates, are used for comparison so that the complexity of the
calculation does not increase.</p>
        <p>To define queries in natural language, we use the pre-trained transformer model GPT4all
which generate SQL (Structured Query Language) queries from natural language queries. The
results are presented in a single result page. We implemented a cross-cultural bookbinding
information system as a prototype to demonstrate how to search in multiple databases with one
query defined in natural language.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>In the literature, there are several works on the topic of FIR. Federated search can be challenging
in terms of retrieving relevant information for the user. We present a few approaches, each
describing a diferent focus.</p>
      <p>
        Shokouhi and Si [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] have provided a foundational definition of federated search and delved
into its potential applications and challenges. Their work notably sheds light on the persistent
issue of maintaining up-to-date representation sets, proposing innovative methods to address
this challenge. In addition to the definition of federated search, their contributions have been
pivotal in understanding and mitigating issues related to the timeliness and accuracy of search
results in federated systems.
      </p>
      <p>
        Building on the concept of federated search, Demeester et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] conducted a study focused on
the use of snippets, rather than entire webpages, to predict the relevance of a given page. This
approach, which examines the content at a more granular level, proves to be particularly valuable
in the context of federated search. By ofering insights into deeper details of page content, their
work contributes to enhance precision and eficiency of federated search algorithms. Eficiency
of query processing will also be an issue in our work.
      </p>
      <p>
        Federated search, while promising, can be inherently challenging when it comes to retrieving
pertinent information for users. FedCDR [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] introduces a novel approach known as federated
cross-domain recommendation. This innovative method addresses the delicate balance between
providing users with tailored recommendations while safeguarding their private data. FedCDR’s
contributions are instrumental in ensuring that federated search remains user-centric and
privacy-conscious. This aspect of safeguarding of private data should definitely be addressed in
productive systems.
      </p>
      <p>
        Furthermore, Melzer et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] present a methodology designed to simulate federated
databases, ofering a means of conducting feasibility studies before committing to the
implementation of real federated databases. This approach enables researchers and organizations
to experiment and assess the viability of federated database projects, reducing the risk of
investing substantial resources in endeavors that may ultimately prove unfeasible.
      </p>
      <p>
        In the area of federated search, the diversity of data sources often poses a significant challenge
due to the heterogeneity of data formats and structures. Addressing this concern, Melzer et
al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] introduce a novel indexing process tailored to matching data from XML files and the
relational representation of research data so that eficient searches across heterogeneous data
sets are given.
      </p>
      <p>
        The process of building information systems on demand is described in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This innovative
approach empowers humanities scholars by allowing them to construct information systems
without the arduous task of manually transferring data.
      </p>
      <p>Finally, recent advances in NLP (Natural Language Processing), particularly the use of models
such as GPT (Generative Pre-trained Transformer), have shown promise in simplifying the
creation of SQL queries.5 Using GPT-based NLP techniques, users can formulate queries in
natural language, which are then automatically translated into SQL queries that retrieve relevant
information from various federated databases. This innovative approach not only streamlines
the query process, but also enables a wider range of users, including those who do not have
extensive SQL knowledge, to efectively use the full potential of federated information systems.
Therefore, we will integrate this functionality into a cross-domain information system.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Bookbinding</title>
      <p>Binding is the process by which stacked sheets or quires are secured along one edge with needle
and thread or other materials such as loose-leaf rings, binding posts, twin-loop spine coils,
plastic spiral coils, and plastic spine combs. The bound stack of leaves can then be enclosed in a
cover. Bookbinding is a skilled craft that requires measuring, cutting, and gluing, and combines
skills from the trades of paper making, textile and leather-working crafts, model making, and
graphic design. There are various types of bookbinding techniques, they are imparted by
tradition, evolve across time taught from one generation to the next, and assume distinctive
traits according to the area to which they belong. The presence of recurring patterns in the
structures allows to group the bindings accordingly, thus identifying macro-areas corresponding
to diferent binding traditions (Coptic, Ethiopian, Islamic, Byzantine, etc.). Modern binding
methods are numerous, such as perfect binding, case binding, saddle stitch binding, PUR
binding, singer sewn binding, section sewn binding, Coptic stitch binding, wiro and comb
binding. [12, 13] Three diferent bookbinding techniques are described in the following.</p>
      <sec id="sec-3-1">
        <title>3.1. Ethiopian Bookbinding</title>
        <p>When writing was adopted by the Semites who settled in the area between the northern
highlands of the Horn of Africa and the Red Sea. The existence of an extensive Christian
literature going back to the fourth century CE implies the use of manuscripts. The Ethiopian
language and script used for centuries as the literary language of the Christian kingdom of
Ethiopia are very similar to those used in the fourth century. Ethiopian bookbinding is one of
the material expressions of the ancient manuscript culture of Ethiopia and Eritrea, which is the
research field of the Beta mas.a¯ h. fet project. The expression ‘Ethiopian bookbinding’ identifies a
set of structural features shared by the bindings of Christian manuscripts produced in Ethiopia
and Eritrea. These include chainstitch sewing (mostly) on paired sewing stations, slit-braid
endbands, and wooden boards, which may be covered with leather and lined with colourful
textiles. In Ethiopic manuscripts, the writing support is usually parchment, produced without
making use of lime baths. [14]</p>
        <sec id="sec-3-1-1">
          <title>5https://github.com/soumyansh/NLP-To-SQL</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Coptic Bookbinding</title>
        <p>The expression Coptic bookbinding is commonly used to refer to the binding techniques prevalent
in Egypt in the Late Antique and Early Medieval eras. Coptic bookbinding is a historical
expression, deeply rooted in the literature, which refers to the binding tradition prevalent in
Egypt during the Late Antique and Early Medieval periods. Coptic book structures vary, and
include single quires attached directly to the leather cover using tackets; multi-quire codices
sewn with chainstitch and furnished with wooden boards, or laminated papyrus boards with
leather covers.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Pothi and Codex Binding in South India</title>
        <p>In South India, a traditional manuscript — or pothi — consists of a stack of palm leaves, in
landscape format, inscribed with a stylus, and bound together with a string thread through
holes in the folios. These folios were often protected with wooden board covers. But with the
arrival of Portuguese traders and missionaries in the 16th century, a new manuscript format
became increasingly common: the codex. The early codices from South India and the way in
which Western bookbinding techniques were learnt and applied by local craftsmen have hardly
been researched so far. By the 19th century, new hybrid formats had begun to emerge across
India: Sanscrit 1232, preserved at the National Library of France, is a fascinating codex-pothi
hybrid, a lithograph printed in horizontal pothi format but collated in sections of two bifold
each. This data, on early modern South Indian bookbinding, has been collected by the Texts
Surrounding Texts project, a catalogue of Indian manuscripts from the National Library of
France and the Staats- und Universitätsbibliothek Hamburg.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. The Need for a Cross-Cultural Bookbinding Information System</title>
        <p>To date, there has been little work on comparing bookbinding practices across cultures.
Codicological expertise does not necessarily translate from one field to another; an expert in Coptic
bookbinding would not know how to approach an Indian manuscript, or vice versa. As a result,
research projects usually focus on a specific culture and a specific time period. To compare
practices across cultures, we would need to, firstly, understand which data can be compared,
and secondly, to collate that data by extracting it from diferent, heterogeneous databases. In
the three aforementioned databases that will be used as the foundation for the Cross-Cultural
Bookbinding Information System, we have initially selected three features to be compared: the
number of sewing stations, leaf width, and leaf height. For the first time, we will be able to
compare bookbinding techniques as they spread across space and time, and as they crossed
boundaries of language, religion, and material tradition.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Matching Bookbinding Data</title>
      <p>In general, matching data sets involves comparing two or more data sets to identify similar
elements. The process of matching data involves several steps (see Figure 1). According to
[15], the first step is data pre-processing, which involves preparing the data sets for matching.
This includes cleaning, formatting, and standardizing the data sets to ensure compatibility and
efective comparison. The second step is indexing, which is a strategy to pre-select potential
matches and leads to a reduction in the number of matches. Indexing usually involves identifying
the key variables that will be used to match the data sets. The third step is comparison, where the
data sets are compared to identify matches. The fourth step is classification, where the matching
records are classified as match or non-match. The final step is evaluation, which involves
validating the matched data sets and reviewing the results for accuracy and completeness. This
may involve checking for errors, inconsistencies, or missing data and making any necessary
adjustments.</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] an improvement of the indexing procedure using XML (EpiDoc) and relational
representations of research data as input, is presented.
      </p>
      <p>Pre-Processing The following projects have diferent relational representations to describe
manuscripts and bookbinding techniques.</p>
      <p>• The Beta mas.a¯ h. fet project has the following relation representation. The column names
are “Title”, “Editor(s)”, “PubPlace”,“Manuscript Item(s)”, “idno”, “Material”, “Deco Note(s)”,
“Hand Description”, “Binding”, ”Orig. Date”, . . ..
• The Coptic Bookbinding project has the following relation representation. The column
names are “CLM”, “TM”, “Shelfmarks”, “Leaf width”, “Leaf height”, “Board height”, “Board
width”, ”Spine width”, “Type of sewing”, “No. of sewing stations”, “Fold pattern”, . . ..
• The Texts Surrounding Texts project has the following relation representation. The column
names are “Title”, “Shelfmark”, “Format”, “Technology”, “Material”, ”Leaf width”, ”Leaf
height”, “Leaf depth”, “Binding”, . . ..</p>
      <p>
        It can be seen that not all column names have the same name. The bindings are described
under “Deco Note(s)” in Beta mas.a¯ h. fet, this data can be found under “Binding” in the other both
projects. However, while the “number of sewing stations” is described under “‘Deco Note(s)” in
Beta mas.a¯ h. fet, this data is found under “No. of sewing stations” in Coptic Bookbinding. At this
point, a mapping function must therefore be defined (by the humanities scholars) so that one
knows which column names are mapped to one another. If the mapping rules are not known,
one can also use large language models (LLMs), as also shown in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], to obtain them. However,
it should be noted that the schemes should be given as input so that the results of the LLMs can
be used.
      </p>
      <p>In this article, we explain the indexing process using the two projects: Beta mas.a¯ h. fet and
Coptic Bookbinding.</p>
      <p>Indexing Indexing includes identifying the key variables for an eficient data matching
process. For existing relational databases, it can be assumed that the column names belong to
the key variables and are used for their project-specific analysis. Therefore, the column names
are regarded as key variables.</p>
      <p>To identify the matching candidates, we use the XML and the JSON scheme used in the
projects (where the raw data is stored).</p>
      <p>Formally: If a set  of XML tags and  a set of JSON nodes, where the sets  and  are from
diferent schemes, are mapped to the same element, then that element is a matching candidate
to be added to the matching candidate set .</p>
      <p>Let  = {1, . . . } and  = {1, . . .  } be sets of XML tags or JSON nodes, and let  be a
function which represents a mapping from  to :  :  → , then the matching candidates
 are given by:
 = { ∈  : ∃ ∈  with  () = }
(1)
In our example, the “JSON” schema belongs to set A and the TEI schema Beta mas.a¯ h. fet to set B.
Table 1 displays the column names used in the respective projects and the corresponding XML
tags and JSON nodes. The matching candidates are  = {leaf width, leaf height}.</p>
      <p>In our project, however, we also need the “number of sewings” that are not considered in the
matching candidates.</p>
      <p>I. e. the comparison of the “number of sewings” in this example is done via height and width.
In order for “number of sewings” to be a matching candidate, it should be noted here that a
standard should be applied semantically correctly in the various projects or an adjustment could
be made in the pre-processing.</p>
      <p>Matching In Table 2 each matching data (width and height) of both projects (Coptic
Bookbinding, Beta mas.a¯ h. fet) were assigned an id. The table also present some more data (CLM/ID)
to have a better overview of the data.</p>
      <p>The content of the matching candidates are compared are compared for equality (:=1) or
inequality (:=0). We use this simple comparison because only values need to be compared.
For words, texts or dates, other comparison approaches such as the Soundex algorithm [16],
Levenshtein distance [17] or suitable artificial intelligence (AI) algorithm can be used instead.
If we consider all matching candidates (separate comparison of width and height data), the
identified record pairs are: (3, 2), (4, 2).</p>
      <p>The fact that only one record pair was identified is due to the simple number matching.
With the leaf width and height, one could also allow smaller deviations if it fits the content.
For a simple illustration of the indexing process, we will first continue with the one identified
matching candidate.</p>
      <p>Comparison The comparison process in schema matching indicates the degree of similarity
between two record pairs to determine whether they are a match or not. In general, for the
comparison process all fields are considered. In Table 3 the column names are: width (mm),
height (mm), and No.sewing. Consider that “decoNote” and “sewingstationsno” represent both
“No.sewing.”</p>
      <p>
        The comparison function (,  ) maps the content of each column value of  and  in the
range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ], where 0 indicates no similarity and 1 indicates a perfect match. The comparison
function can be defined using diferent similarity metrics depending on the characteristics of
the schema elements and the matching criteria.
      </p>
      <p>The following comparison function can be used to rank the candidate matches based on their
similarity scores:
number of attributes - 1
∑︁
=0
simall(,  ) =
((),  ()),
(2)
where an attribute is a column name and  is the position of the column.</p>
      <p>The classification of each compared record pair can be based on either the full comparison
vectors or on the summed similarities. Based on the summed similarity score, a match is defined
as:
match =
{︃1 sim ≥ 
0 otherwise
(3)</p>
      <p>In the context of the project, a good value for  is between the “number of attributes” divided
by 2 and the total “number of attributes” to achieve matching results between approximately
50% and below 100%. Formally:
number of attributes</p>
      <p>≤  &lt; number of attributes. (4)
2
If  = “number of attributes” (100% similarity), then it could indicate a duplicate.</p>
      <p>This matching algorithm can compare the data in an ofline process. This algorithm can then
be implemented in FIR in such a way that the category, such as “No. of sewings” is created and
the most similar data sets are displayed to the user.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Federated Information Retrieval</title>
      <p>FIR, also known as distributed information retrieval or federated search, is a technique used
to search multiple data sources simultaneously. It allows users to retrieve information from
various content locations with just one query and one search interface. Federated search has
revolutionized how user search and retrieve information online, making it easier for researchers
to manage data and search for data. Implementing a federated search engine can be challenging,
especially when integrating the system with heterogeneous databases. Federated search is
an eficient option for mid-to-low funnel users who know exactly what they need and can
search through a large body of data from one location with one query, reaching their goal with
fewer eforts.</p>
      <p>Architecture In recent years, the Sqlite database has become more and more common as a
way to share research data. For example, the website of the Texts Surrounding Texts Project is a
front-end that queries a read-only Sqlite database hosted on GitHub. This architecture means
that the project automatically has its own, open API — any researcher, any website can also
access the database using SQL queries, without requiring any authentication. A cross-cultural
bookbinding information system takes advantage of this openness by connecting directly to
the Texts Surrounding Texts Sqlite database and extracting bookbinding data from it, which
is then collated with bookbinding data from the Ethiopian, Eritrean, and Coptic databases. In
demonstrating our federated search application, we hope to encourage more and more research
projects to make their databases openly accessible in this way, so that researchers can more
easily cross-reference data from multiple sources.</p>
      <p>Querying The use of natural language queries instead of SQL for accessing databases has
been an area of active research in recent years. One approach involves the use of transformer
models such as GPT to generate SQL queries from natural language queries. We used the
GPT4All6 library with the “wizardlm-13b-v1.1-superhot-8k”7 model to generate SQL queries
from natural language queries. The source code for this implementation is based on the code of
“soumyansh” on GitHub [18].</p>
      <p>The basic idea of this querying approach is to pass information about the database, in our
case the names of the table together with the column names, together with the prompt given by
the user. An example prompt can be seen in Figure 2. Since we only want to allow “SELECT”
statements to be executed automatically on the databases, it is given as part of the prompt to
the GPT. After the prompt has been generated, it is passed to the chosen GPT model using the
GPT4All library (see Figure 3). Depending on the hardware resources, chosen model and query,
execution time is around 45 to 60 seconds. Once the GPT has generated an output, it is further
passed on to the functions responsible to generate the webpage. This querying approach can be
generalized to  databases by repeating the process  times. While this makes it easy to apply
the same principle to an undefined number of databases, it also increases execution time per
database added. Further work needs to be done to make the process faster when using a large
number of databases to query.</p>
      <sec id="sec-5-1">
        <title>6https://gpt4all.io/index.html</title>
        <p>7https://huggingface.co/TheBloke/WizardLM-13B-V1-1-SuperHOT-8K-GGML/resolve/main/wizardlm-13b-v1.
1-superhot-8k.ggmlv3.q4_0.bin</p>
        <sec id="sec-5-1-1">
          <title>Federated Bookbinding Information System By extracting bookbinding data from the</title>
          <p>three databases pertaining to three diferent manuscript traditions, we can begin to compare
how the codex format was adapted by diferent cultures at diferent periods of time.</p>
          <p>In our prototype (see Figure 4) it can already be seen that desired database entries from
diferent database instances can be viewed in one view. This joint representation makes it easier
to answer the research question and to prove it with concrete values. The additional linking
to similar documents through the matching algorithm improves the overview of information.
The prototype still needs to be further refined over time, as not all queries have been answered
correctly so far. Additionally, it takes a few minutes to execute queries. While a user only
sends one natural language query to the system, it internally generates a separate SQL query
per database. This makes it possible to generate queries to databases with diferent table as
well as column names. During our testing, the system seemed to give reasonable results. The
performance will be formally evaluated at a later stage, which presents an opportunity for
improvement. This cross-cultural bookbinding information system was created with little efort.
Although work still needs to be put into a productive system for correctly responding to all
user requests. Using the indexing process, we can now ofer similar documents to each record.
Which they are for our example will be presented in the next subsection.</p>
          <p>Federated Search Results In an ofline process, the bookbinding matching process can be
applied. In Table 5 you can see the results if the three columns width, height, and number of
sewing stations (cf. Figure 4) are defined as the relational structure. According to Equation 4,
we receive the data sets that fulfil 1.5 ≤  &lt; 3.</p>
          <p>For the “number of sewing” category, the precision score is perfect, with a value of 1. This
implies that all the instances identified as belonging to the “number of sewing” category were
indeed accurate, leaving no room for false positives. In contrast, for the “width” category, the
precision score is 0.535, indicating that approximately 53.5% of the items classified as “width”
were true positives, while the remaining 46.5% were false positives. This suggests some room
for improvement in reducing false positives within the “width” category. The “height” category
exhibits a precision score of 0.465, indicating that about 46.5% of the items identified as “height”
were true positives, while 53.5% were false positives. Similar to the “width” category, there is
potential for enhancing precision within the “height” category to reduce false positives.</p>
          <p>Even though the values for precision are not very high for some values, it can be seen in
the following that the choice for  in this example is well chosen. If a diferent value is taken
for theta, the results are much worse. That is for 1 ≤  &lt; 3 as follows: For the “number of
sewing stations,” the precision score is 0.99 and therefore high. This implies that the process
for determining the number of sewing stations is remarkably accurate, with only a 1% margin
for error. However, the precision values for “width” and “height” paint a diferent picture. The
precision score of 0.001 for “width” indicates a notable lack of precision in this measurement.
Similarly, the precision score of 0.007 for “height” also suggests a measurement process that
falls short in terms of precision.</p>
          <p>In data analysis and research, the choice of parameters like  is just one piece of the puzzle.
Equally important is the alignment of these parameters with the overarching research question.
In the context of the Coptic Bookbinding project, it is obvious that expanding the data set and
considering additional suggestions has proven beneficial.</p>
          <p>Researchers can improve the robustness of their analysis and enhance the overall quality
of results. Incorporating more data points and seeking suggestions from similar data sets can
provide a broader context and lead to more meaningful insights. This approach not only helps
in fine-tuning the parameters but also contributes to a deeper understanding of the subject
matter and the research objectives.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Outlook</title>
      <p>In this article, we present a cross-cultural bookbinding information system that supports FIR
with low efort. We demonstrate how federated search can be used to retrieve information from
various digital resources produced by research institutions, museums, or libraries to answer
cross-domain research questions. Our system integrates GPT4All to generate SQL queries
from natural language queries, enabling users to search for similarities between the binding
techniques, object size, or written area dimension of manuscripts from Ethiopia, Eritrea, early
Egypt, and South India. We use a data matching method to make the search for finding similar
data sets more eficient. With the bookbinding information system, we have succeeded in
substantiating statements with numbers by combining diferent sources.</p>
      <p>In the future, we plan to expand the system to include more digital collections and data
sources. We also plan to improve the system’s search capabilities by integrating the similarity
score calculation in our prototype. Additionally, we aim to integrate the system with other
research tools and platforms to provide a more comprehensive and seamless research experience
for scholars in the humanities.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>The research for this contribution was funded by the Deutsche Forschungsgemeinschaft (DFG,
German Research Foundation) under Germany’s Excellence Strategy - EXC 2176 ’Understanding
Written Artefacts: Material, Interaction and Transmission in Manuscript Cultures’, project no.
390893796.
[12] Matt Marzullo , WHAT’S IN A BIND? 4 TYPES OF BOOK BINDING - PROS AND CONS,
https://blog.ironmarkusa.com/4-types-book-binding, 2021. Accessed 28 July 2023.
[13] Wikipedia, Bookbinding, https://en.wikipedia.org/wiki/Bookbinding, 2023. Accessed 28</p>
      <p>July 2023.
[14] Universität Hamburg , Background, https://www.betamasaheft.uni-hamburg.de/about/
background.html, 2017. Accessed 28 July 2023.
[15] P. Christen, Data Matching: Concepts and Techniques for Record Linkage, Entity
Resolution, and Duplicate Detection, Springer Publishing Company, Incorporated, 2012.
[16] J. Jacobs, Finding words that sound alike. The SOUNDEX algorithm., Byte 7 (1982) 473–474.
[17] F. P. Miller, A. F. Vandome, J. McBrewster, Levenshtein Distance: Information Theory,
Computer Science, String (Computer Science), String Metric, Damerau-Levenshtein Distance,
Spell Checker, Hamming Distance, Alpha Press, 2009.
[18] soumyansh, NLP-To-SQL, https://github.com/soumyansh/NLP-To-SQL, 2023. GitHub
repository, Accessed 01 September 2023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E. O. F. N.</given-names>
            <surname>Research</surname>
          </string-name>
          , OpenAIRE, Zenodo,
          <year>2013</year>
          . URL: https://www.zenodo.org/. doi:
          <volume>10</volume>
          . 25495/7GXK-
          <fpage>RD71</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Universität</given-names>
            <surname>Hamburg</surname>
          </string-name>
          , Research Data Repository, Available: https://www.fdr.uni-hamburg. de/,
          <source>2022. Accessed March</source>
          <volume>09</volume>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Text</given-names>
            <surname>Encoding</surname>
          </string-name>
          <string-name>
            <surname>Initiative</surname>
          </string-name>
          ,
          <article-title>P5: Guidelines for Electronic Text Encoding</article-title>
          and Interchange,
          <source>Version 4.0</source>
          .0, https://tei-c.org/Vault/P5/4.0.0/doc/tei-p5-doc/en/html/,
          <source>2020. Accessed 29 June</source>
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Schif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Melzer</surname>
          </string-name>
          , E. Wilden,
          <string-name>
            <given-names>R.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <article-title>TEI-Based Interactive Critical Editions</article-title>
          , in: S.
          <string-name>
            <surname>Uchida</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Barney</surname>
          </string-name>
          , V. Eglin (Eds.),
          <source>Document Analysis Systems</source>
          , Springer International Publishing, Cham,
          <year>2022</year>
          , pp.
          <fpage>230</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Zwass</surname>
          </string-name>
          , Vladimir, information system,
          <source>Encyclopedia Britannica</source>
          , https://www.britannica. com/topic/information-system,
          <year>2023</year>
          . Accessed
          <issue>28</issue>
          <year>July 2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Melzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klettke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Weise</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Harter-Uibopuu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <article-title>EpiDoc Data Matching for Federated Information Retrieval in the Humanities</article-title>
          , in: 1st International Workshop on AI in Digital Humanities,
          <source>Computational Social Sciences and Economics Research at part of the 18th Conference on Computer Science and Intelligence Systems (FedCSIS)</source>
          ,
          <source>Proceedings of the 2023 Federated Conference on Computer Science and Intelligence Systems</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1063</fpage>
          -
          <lpage>1068</lpage>
          . URL: https://annals-csis.org/proceedings/2023/pliks/1515.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Elliott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Bodard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mylonas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Stoyanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tupman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vanderbilt</surname>
          </string-name>
          , et al.,
          <article-title>EpiDoc Guidelines: Ancient documents in TEI XML (Version 9)</article-title>
          ., Available: https://epidoc.stoa. org/gl/latest/., (
          <year>2007</year>
          -
          <fpage>2022</fpage>
          ).
          <source>Accessed January 22</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Shokouhi</surname>
          </string-name>
          , L. Si, Federated Search, Found.
          <source>Trends Inf. Retr</source>
          .
          <volume>5</volume>
          (
          <issue>2011</issue>
          )
          <fpage>1</fpage>
          -
          <lpage>102</lpage>
          . URL: https://doi.org/10.1561/1500000010. doi:
          <volume>10</volume>
          .1561/1500000010.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Demeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Trieschnigg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Develder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          ,
          <article-title>Snippet-Based Relevance Predictions for Federated Web Search</article-title>
          , in: P. Serdyukov,
          <string-name>
            <given-names>P.</given-names>
            <surname>Braslavski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. O.</given-names>
            <surname>Kuznetsov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rüger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Agichtein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Segalovich</surname>
          </string-name>
          , E. Yilmaz (Eds.),
          <source>Advances in Information Retrieval</source>
          , Springer Berlin Heidelberg,
          <year>2013</year>
          , pp.
          <fpage>697</fpage>
          -
          <lpage>700</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>W.</given-names>
            <surname>Meihan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rigall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xiaodong</surname>
          </string-name>
          , X. Cheng-Zhong,
          <article-title>Fedcdr: Federated cross-domain recommendation for privacy-preserving rating prediction</article-title>
          ,
          <source>in: Proceedings of the 31st ACM International Conference on Information &amp; Knowledge Management, CIKM '22</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2022</year>
          , p.
          <fpage>2179</fpage>
          -
          <lpage>2188</lpage>
          . URL: https://doi.org/10.1145/3511808.3557320. doi:
          <volume>10</volume>
          .1145/3511808.3557320.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Melzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thiemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <article-title>Modeling and Simulating Federated Databases for early Validation of Federated Searches using the Broker-based SysML Toolbox</article-title>
          , in: IEEE International Systems Conference, SysCon
          <year>2021</year>
          , Vancouver, BC, Canada, April 15 - May 15,
          <year>2021</year>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>