<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Graph Technologies for the Analysis of Historical Social Networks Using Heterogeneous Data Sources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sina Menzel</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marian Dörk</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark-Jan Bludau</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julián Moreno-Schneider</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Georg Rehm</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Leitner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vivien Petras</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DFKI - Deutsches Forschungszentrum für Künstliche Intelligenz GmbH</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>FH Potsdam - University of Applied Sciences</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Humboldt-Universität zu Berlin</institution>
        </aff>
      </contrib-group>
      <fpage>124</fpage>
      <lpage>149</lpage>
      <abstract>
        <p>Over the last decades, cultural heritage institutions have provided extensive machine-readable data, such as bibliographic and archival metadata, full-text collections, and authority records containing multitudes of implicit and explicit statements about the social relations between various types of entities. In this paper, we discuss how approaches to the creation and operation of advanced research infrastructure for historical network analysis (HNA) based on heterogeneous data sources from cultural heritage institutions can be examined and evaluated. Based on our interdisciplinary research, we describe challenges and strategies with a special focus on the issue of data processing, sketch out the advantages of human-centered project design in the form of a preliminary co-design workshop, and present an iterative approach to data visualization.</p>
      </abstract>
      <kwd-group>
        <kwd>In</kwd>
        <kwd>Tara Andrews</kwd>
        <kwd>Franziska Diehr</kwd>
        <kwd>Thomas Efer</kwd>
        <kwd>Andreas Kuczera and Joris van Zundert (eds</kwd>
        <kwd>)</kwd>
        <kwd>Graph Technologies in the Humanities - Proceedings 2020</kwd>
        <kwd>published at http</kwd>
        <kwd>//ceur-ws</kwd>
        <kwd>org</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The study of historical events is relevant to many disciplines in the digital
humanities, with the analysis of relationships between agents often being
crucial for the understanding and explanation of social, political, and cultural
phenomena. Given that historical research is heavily dependent on
information from the respective time period, the combination of as many historical
sources as possible is essential for the reconstruction of historical networks
– and this is where the method of historical network analysis (HNA) comes
into play. Derived from social network analysis, HNA is characterized by the
same dependency on numerous historical sources that ideally support each
other
        <xref ref-type="bibr" rid="ref18">(Jansen and Wald, 2007)</xref>
        .
      </p>
      <p>One limiting factor in HNA can be a lack of awareness with regard to
the availability of suitable research data. At the same time, over the past
decades, cultural heritage institutions have produced very large amounts of
machine-readable and, in many cases, standardized and well-organized data
in the form of bibliographic and archival metadata, full-text collections, and
sets of authority or reference records. These datasets contain a plethora of
implicit and explicit statements about social relations, which can in turn be
exploited for HNA research. However, systematically combining multiple
data sources (not to mention extracting and visualizing the complex resulting
networks) currently requires extensive knowledge in graph theory as well as
time-consuming manual work carried out by the individual researcher. One
reason for this is the heterogeneity of the data sources made available by
cultural heritage institutions, for example, in terms of data formats.</p>
      <p>The research project SoNAR (IDH): Interfaces to Data for Historical
Social Network Analysis and Research1 addresses this issue. We examine and
evaluate approaches to the development and operation of HNA-supporting
research infrastructure based on heterogeneous cultural heritage data. In
this paper, we present a number of preliminary insights related to the
process of modeling and transforming heterogeneous data sources, and to the
design of user-centered visualization for historical social networks. By
sharing our approach and its accompanying challenges, we aim to contribute to
the ongoing discussion on the suitability of bibliographic big data for HNA
and the development of corresponding research technologies.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>The following section gives an overview of previous research, and discusses
projects related to graph modeling and visualization approaches within the
digital humanities from the perspective of historical network analysis.
2.1</p>
      <sec id="sec-2-1">
        <title>Related Projects</title>
        <p>
          In recent years, open knowledge graphs have frequently been used as an
alternative to a document-based approach
          <xref ref-type="bibr" rid="ref3">(Auer and Mann, 2019)</xref>
          . Several
large-scale initiatives such as EOS,2 Europeana,3 and CLARIN4 provide
researchers in the digital humanities with access to cultural data. Meanwhile,
the issue of decentralized and heterogeneous bibliographic data sources is
being addressed by projects such as Culturegraph
          <xref ref-type="bibr" rid="ref35">(Vorndran, 2018)</xref>
          and
DARIAH-DE5 in the digital humanities, Lynx6 in the legal domain, and, to
a certain extent, ELG7 in language technology
          <xref ref-type="bibr" rid="ref30">(Rehm et al., 2020)</xref>
          . Most of
these initiatives connect to infrastructures of cultural heritage institutions,
often hosted by libraries or archives.
        </p>
        <p>Even though these initiatives provide, among other things, access to new,
previously unidentifiable or implicit information, they do not primarily
focus on the extraction of network data. Therefore, HNA researchers are often
left to create their own individual graphs after gathering data that is suitable
to address their research question(s), in many cases using open source
software tools such as Gephi,8 Palladio,9 or VennMaker,10.</p>
        <p>Along with the establishment of network analysis as a method in historical
research, there has been an increase in joint research projects that are focused
on the extraction of historical networks within the social sciences and the
humanities.</p>
        <p>
          For example, the project Six Degrees of Francis Bacon11 applies statistical
methods to the base data with the goal of inferring relations that permit the
reconstruction and visualization of historical social networks in Early
Modern Britain. The project allows for the expansion and curation of the data
through collaborative annotation by the users
          <xref ref-type="bibr" rid="ref36">(Warren et al., 2016)</xref>
          . The
histoGraph12 project follows a similar approach by oefring users an
opportunity to collaboratively explore and research historical social networks by
means of extensive multimedia collections, with a special focus on
crowdsourced indexation
          <xref ref-type="bibr" rid="ref26">(Novak et al., 2014)</xref>
          . In a joint project involving several
European research institutions, Issues with Europe – A Network Analysis
of the German-Speaking Alpine Conservation Movement (1975-2005)13 is
currently examining the disputes over European alpine transit policy, while
the Austrian project APIS – Mapping historical networks has been working
on the extraction and visualization of networks from more than 18,000
records in the Austrian Biographical Encyclopedia.14 Finally, the German
project Gesellschaftliche Wissensproduktion in der Aufklärung – Text- und
netzwerkanalytische Diskursrekonstruktion considers full texts of more than
300 periodicals published in Halle, Germany between 1688 and 1815, and
combines the methods of topic modeling with historical network analysis in
order to systematically analyze public discourse during the Age of
Enlightenment
          <xref ref-type="bibr" rid="ref29">(Purschwitz, 2018)</xref>
          .
        </p>
        <p>These are only a few examples of the ongoing eofrts to provide users with
direct access to networks in existing data collections. In our project, we are
working with data sources that have not been modeled for HNA before. Our
generic data approach is closely connected to similar projects, like the North
American cooperative SNAC – Social Networks in Archival Context15 and
the French project PIAAF,16, which both have a strong focus on archival
metadata and full texts.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Network Visualization</title>
        <p>
          As far as the visualization of data for HNA is concerned, many interfaces
have been developed over the years that oefr explorative, web-based
network visualization tools for historical network analysis. Examples include
the above-mentioned Six Degrees of Francis Bacon
          <xref ref-type="bibr" rid="ref36">(Warren et al., 2016)</xref>
          and
histoGraph
          <xref ref-type="bibr" rid="ref26">(Novak et al., 2014)</xref>
          , as well as Visualizing the Republic of
Letters
          <xref ref-type="bibr" rid="ref7">(Chang et al., 2009)</xref>
          , Kindred Britain,17 and Deutsche Biographie.18
        </p>
        <p>
          Graph visualization is an extensive field in itself, which is accompanied
by a substantial body of literature on issues such as graph-related algorithms
          <xref ref-type="bibr" rid="ref12 ref17 ref4">(e. g. Gibson et al., 2012; Jacomy et al., 2014; Behrisch et al., 2016)</xref>
          , task
taxonomies for graph visualization
          <xref ref-type="bibr" rid="ref1 ref19 ref21">(e. g. Lee et al., 2006; Ahn et al., 2013;
Kerracher et al., 2015)</xref>
          , state-of-the-art visualization interaction techniques and
13https://www.uibk.ac.at/projects/issues-with-europe/index.html.en
14Österreichisches Biographisches Lexikon, https://apis.acdh.oeaw.ac.at
15https://snaccooperative.org
16Pilote d’interopérabilité pour les autorités archivistiques françaises https://piaaf.demo.
logilab.fr
17http://kindred.stanford.edu
18https://www.deutsche-biographie.de
developments
          <xref ref-type="bibr" rid="ref27 ref33 ref34">(e. g. van Ham and Perer, 2009; von Landesberger et al., 2011;
Pienta et al., 2015)</xref>
          , as well as the use of visual facilitators for the construction
of graph queries
          <xref ref-type="bibr" rid="ref28">(e. g. Pienta et al., 2017)</xref>
          . Nevertheless, existing research and
taxonomies mostly address the wider field of graph visualization. More
often than not, visualizations and digital practices are not specifically adapted
to the requirements of HNA research or established data practices in the
humanities, and are ill-suited to address issues such as uncertainty, subjectivity,
or observer-dependence
          <xref ref-type="bibr" rid="ref9">(Drucker, 2011)</xref>
          .
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Human-Centered Design</title>
        <p>
          A key element in the examination and development of a new research
infrastructure designed for human-computer interaction is how well it meets the
needs of the people it is intended to assist. This human-centered approach
is closely related to Grounded Theory, which generates inductive results by
means of sociological methods
          <xref ref-type="bibr" rid="ref13">(Glaser and Strauss, 1967)</xref>
          .
        </p>
        <p>
          <xref ref-type="bibr" rid="ref16">Isenberg et al. (2008)</xref>
          adapted Grounded Theory for the evaluation of
information visualizations. They suggest iterative evaluation throughout the
process of system development using several points of qualitative inquiry to
ensure the focus of a system’s intended use, including field research to
examine potential contexts of human interaction with the system. In
keeping with this argument for grounded evaluation, the neuralgic points for
evaluation in our project are based on Munzner’s nested model for
visualization design and validation
          <xref ref-type="bibr" rid="ref25">(Munzner, 2009)</xref>
          , which allows for iterative
improvement of the prototypes. The stages of evaluation include the
assessment of possible use cases, and the investigation of the problems and data
of a particular user domain at the top level. In order to better address such
issues, it is becoming more and more common to include domain experts
in the creation process of digital humanities-related projects. This kind of
co-creation is precisely what
          <xref ref-type="bibr" rid="ref8">Chen et al. (2014)</xref>
          attempted to foster with a
workshop, wherein the participants were asked to create collages to make
sense of a photo archive with the aim of creating collection-sensitive
interfaces.
          <xref ref-type="bibr" rid="ref14">Henry and Fekete (2006)</xref>
          used a similar participatory approach in the
development of a tool for the exploration of social networks: they invited
social science researchers to create paper prototypes, which in turn led to a
list of domain requirements for their tool and resulted in a prototype with
novel features. A thorough evaluation of such co-creation methods,
conducted in a co-design process with social science researchers, found that domain
experts in general appreciate their additional empowerment in the process
and the domain-customized results based on their specific needs.
Nevertheless, regarding their personal involvement and necessary time commitment,
some participants did not perceive their personal involvement as beneficial
for the facilitation of their own research
          <xref ref-type="bibr" rid="ref24">(Molina León and Breiter, 2020)</xref>
          .
Besides the use of co-design techniques, there is also a shift from perceiving
visualizations as mere tools for humanities-related research towards the
acknowledgment of visualization and visualization processes as a methodology
and facilitator of cross-disciplinary research in and of itself
          <xref ref-type="bibr" rid="ref15">(Hinrichs et al.,
2019)</xref>
          . While we have noticed increased attention to the method of HNA,
to the best of our knowledge, there has so far been little investigation of the
modeling and visualization of (bibliographical) big data for this purpose.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data Sources</title>
      <p>The interdisciplinary project SoNAR (IDH), which studies the potential of
large heterogeneous data collections for HNA, includes partners from the
ifelds of historiography, information visualization, and artificial intelligence,
as well as computer and information science. This variety of disciplines
opens diefrent perspectives on the requirements and challenges connected
to the use of heterogeneous (meta)data for HNA. What distinguishes our
approach is the synchronous operation of all components of the project, so
that the design of the data technology, the development of a model research
design for HNA, and the development of innovative visualization and
interface approaches with the involvement of HNA experts are all intertwined
and influence one another.</p>
      <p>
        The project is based on heterogeneous source data from authority files,
bibliographic records, and full texts. The data is available in various
XMLbased formats such as MARC21
        <xref ref-type="bibr" rid="ref20">(Kruk et al., 2005)</xref>
        , EAD
        <xref ref-type="bibr" rid="ref2">(Allison-Bunnell,
2016)</xref>
        , and METS/ALTO19
        <xref ref-type="bibr" rid="ref6">(Cantara, 2005)</xref>
        :
• The Integrated Authority File (GND)20 represents and describes
8,295,047 entities (people, corporations, conferences, geographical
areas, technical terms, and works);
• The German National Library (DNB)21 provides descriptions of
bibliographic resources. The dataset has 19,926,573 records of books,
magazines, newspapers, sheet music, music recordings, audio books
etc.;
• The German Union Catalogue of Serials (ZDB)22 describes
newspapers, magazines, serial titles, yearbooks, etc. and contains 1,908,334
records;
19http://www.loc.gov/standards/alto
20https://www.dnb.de/EN/Professionell/Standardisierung/GND/gnd_node.html
21https://www.dnb.de/EN/Home/home_node.html
22https://zdb-katalog.de/index.xhtml
• The Kalliope Union Catalog (KPE)23 is a collection of personal
papers, manuscripts, and publishers’ archives, which consists of 26,752
records;
• The Newspaper Information System (ZeFYS)24 represents 2,596,641
digitized pages of historical newspapers and full texts;
• The Exile Press25 represents German-language exile journals between
1933 and 1945 and consists of 5,336 digitized pages.
      </p>
      <p>Since the source data – describing entities (authority files) and resources
(bibliographic files) – is encoded in various formats, these formats must first
be analyzed in order to enable the design of an appropriate data model and
allow their transformation into a uniform, generic format. Full texts are
prepared for automatic enrichment (i. e. named entity recognition and linking)
and converted to a corresponding format.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Data Processing</title>
      <p>In this section, we will give an overview of the data transformation and graph
modeling process, and outline the challenges that we have encountered along
the way.</p>
      <p>
        The technical goal of our project is the integration of the various source
datasets into a common research infrastructure. We currently use the graph
database Neo4j,26 which is well suited to the eficient storage and
highperformance analysis of large amounts of highly networked information
        <xref ref-type="bibr" rid="ref10 ref22">(Efer, 2016; Matschinegg and Nicka, 2018; Wintergrün, 2019)</xref>
        . Entities are
modeled as nodes and relations as edges with absolute and relational features.
      </p>
      <p>There are a total of 9 entity types extracted from the source data:
1. Person PerName;
2. Corporate body CorpName;
3. Place or geographic name GeoName;
4. Conference or event MeetName;
5. Subject heading TopicTerm;
6. Work UniTitle;
7. Temporal information ChronTerm;
8. Information about ISIL27 IsilTerm;
9. Resource Resource.
23https://kalliope-verbund.info/en/index.html
24http://zefys.staatsbibliothek-berlin.de/index.php?id=start&amp;L=1
25https://www.dnb.de/EN/Sammlungen/DEA/Exilpresse/exilpresse_node.html
26https://neo4j.com
27International Standard Identifier for Libraries and Related Organizations</p>
      <p>Six entity types (i. e., person PerName, corporate body CorpName, place or
geographic name GeoName, conference or event MeetName, subject heading TopicTerm,
and work UniTitle) are taken from the corresponding classes of the authority
ifles. Bibliographic entities are represented as Resource. We added two types to
this list: ChronTerm, which describes temporal information encoded in entity
types from authority files; and IsilTerm, which is used to identify the
libraries related to other entity types. Each entity has general features, such as a
unique source identifier, URI, name, link, etc., and specific features, such as
age, gender, coordinates, etc. Furthermore, there are also nine relation types
that correspond to entity types, such as RelationToPerName, RelationToCorpName,
RelationToGeoName. Relations between entities include information about the
relation source, relation source type, information about temporal validity,
and additional information (Figure 1).
4.1</p>
      <sec id="sec-4-1">
        <title>Data Model</title>
        <p>While the relations between entities are explicitly described in authority files,
relations between actors such as persons or corporate bodies that are
identiifed or defined in the resource are only implicitly encoded in bibliographic
ifles. Our aim is to automatically infer these implicit relations with the
assistance of a set of strict guidelines (e. g. a connection between two persons can
be assumed if both are co-authors of a scientific publication), and to make
them available as explicitly encoded data. In order to derive corresponding
relation types, the role of actors regarding a specific resource (e. g. as author,
editor, or addressee) and the resource type (bibliographic files of primary
sources of the Kalliope Union Catalog and of secondary sources of the
German National Library and the German Union Catalog of Serials) are to be
taken into account. Using this approach, we were able to infer additional
relations (but marked them as computed), for instance between co-authors,
co-publishers, and authors/addressees, to further enrich the data.</p>
        <p>In order to prepare full texts for analysis, named entities are automatically
recognized, disambiguated, and linked to their associated authority files (e. g.
the Integrated Authority File or Wikidata28). Next, relations between
detected entities are automatically recognized, added to the graph database, and
connected with their respective full texts, represented as nodes.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Challenges and Solutions</title>
        <p>Overall, the authority and bibliographic files used by us contain
approximately 30 million records that describe entities and resources in detail. As was
to be expected, normalization of the data revealed a number of errors and
inconsistencies. In this section, we would like to describe some particularly
problematic areas in more detail and suggest possible solutions.</p>
        <p>We have modeled and transformed data for the graph database in such
a way that identifiers are used as coordinates for relations between entities.
In the Integrated Authority File, entities with old identifiers were found, so
that an appropriate connection of two entities was not possible. The first
challenge was to detect old identifiers and replace them with valid ones in
order to enable error-free representation. All replacements were written in a
log file. However, during a consistency check we also found relations to
entities within the source data that were without identifiers. Since such entities
could not be clearly assigned to existing entities with identifiers, ambiguous
relations of this kind had to be ignored.</p>
        <p>Information that was encrypted in internal codes in the Integrated
Authority File, the German National Library, and the German Union
Catalogue of Serials (in format MARC21) was also checked for codes of general
and specific entity types, codes of relation types between an agent and a
resource, and country codes. Further examinations were performed on the
consistency of entity names, resource titles, and identifiers. Again, all errors
or inconsistencies were written in a log file.</p>
        <p>Building on the conclusions that we were able to draw from testing Neo4j,
we decided to adapt the data model to our needs. In order to simplify
searching and filtering according to temporal dimension, time information from
the source data was adjusted. First, while retaining the source data, we
additionally separated time intervals, noted as “begin” and “end.” Second, in
order to facilitate more performant visualization and querying of the data, we
added a feature to resource descriptions that reflects the year of publication
(in addition to the publication date). Thirdly, diefring time expressions in
MARC21 and EAD were normalized.</p>
        <p>We also decided to change gender-specific names of professions. These
are represented in the Integrated Authority File as two diefrent entities with
their own identifiers, male and female. Conceptually, however, what we are
dealing with is a single entity with two versions, so these versions must be
merged in the graph database and represented as a node. One challenge is
to adequately display all information from the two versions without making
the search more dificult. In this case, we are currently still looking for a
suitable solution.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Co-Design Workshop</title>
      <p>
        In accordance with the principles of grounded evaluation
        <xref ref-type="bibr" rid="ref16">(Isenberg et al.,
2008)</xref>
        , we aim to closely integrate domain experts into the data modeling
and visualization process. The presence of HNA experts in our project team
means that all internal decisions that are made take the domain perspective
into account. Additionally, the inclusion of external domain experts is
another integral part of our research design. Conducting studies with
researchers from various fields allows us to iteratively improve the project’s outcome.
      </p>
      <p>
        At the beginning of the project, it was important to us to stimulate
discussions on the potential of bibliographic (meta)data for HNA, and on the
requirements for the visualization of historical networks. Following the
approach proposed by
        <xref ref-type="bibr" rid="ref8">Chen et al. (2014)</xref>
        and
        <xref ref-type="bibr" rid="ref14">Henry and Fekete (2006)</xref>
        , we
organized a co-design workshop that included domain experts in order to help
identify key aspects and gain new insights into historical network research
and visualization.
5.1
      </p>
      <sec id="sec-5-1">
        <title>Procedure</title>
        <p>The workshop consisted of ten participants, including four historical/social
network practitioners as domain experts, two project-internal information
visualization designers/engineers, two members of our project-internal
evaluation team, one member of our team of data scientists (responsible for the
data transformation), and an external participant who had a background in
design and previous experience with the co-design format. The
interdisciplinary composition of the group was intended to enrich the discussion by
oefring a multitude of perspectives on the topic of HNA through the lens
of HNA experts, with fresh insights being provided by participants from
other (project-relevant) fields. Since we aim to develop an infrastructure for
HNA that can be used by researchers from all disciplines working with this
method, the participation of experts from fields other than history was
especially welcome.</p>
        <p>
          The workshop was scheduled for three hours in total. As suggested by
          <xref ref-type="bibr" rid="ref11">Fekete and Plaisant (2002)</xref>
          , we started of with a brief presentation of various
recent developments in the field of network visualization, including some of
the more novel and experimental approaches.
        </p>
        <p>
          We started the process of conceptualizing network visualizations with a
short, hands-on visualization exercise, during which the participants were
asked to visualize a very small social network (ten nodes) based on a data
matrix we provided. After this warm-up, we gave a short introduction to the
goals of our project and the data we are using. The participants were then
asked to create a collage depicting possible approaches to HNA research with
our specific data and project in mind (see Figure 2). For the collages, we
supplied a variety of materials (e. g. construction paper, pencils and
markers, sticky notes). While
          <xref ref-type="bibr" rid="ref8">Chen et al. (2014)</xref>
          provided visual material from
their photographic collection, our data is more abstract and less visual. To
compensate for this, we printed out and distributed further visual material
including an empty map, various icons (e. g. as representations of network
nodes), and a small number of scans from our full-text data sources. We then
introduced several questions to help initiate the creative process (e. g. “How
would you like to move through the data?” and “What role do data
dimensions such as time, space, or semantic relationships play?”), but encouraged
the participants to feel free to disregard them. The task we had in mind was
not to create wireframe sketches for a concrete user interface, but to envision
desired functionalities as well as general approaches and entrance points to
HNA research and our data.
        </p>
        <p>After about 30 minutes, each of the collages was discussed. First, the
participants not involved in making a collage were asked to interpret and speculate
about what they were seeing. Afterwards, the creators of the collages were
asked to give explanations and discuss their approach with the group. In
this step, the almost inevitable misinterpretations were meant to foster
further discussion and novel ideas. In the final step, each participant was asked
to give a closing statement recapitulating the most important insights from
the process and the most prominent topics or themes in the discussion.</p>
        <p>For further analysis and documentation, the entire workshop was
audiorecorded and photographed. The audio recordings were transcribed and
encoded in a tool for qualitative data analysis. This allowed us to assess various
qualitative aspects of the workshop discussions at a later date. As pointed
out above, the goal of our workshop was not to create functional wireframes
or concrete interaction principles, but to stimulate discussion, foster
sensibility towards the domain and data, and highlight important domain-specific
research aspects and challenges. The following section will discuss some of
the most relevant insights concerning our visualization process.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Results</title>
        <p>We noticed two diefrent types of statement. On a more abstract level, the
participants expressed various information needs that commonly arise in the
process of their research. In some cases, however, the conversation and the
collages yielded very concrete ideas regarding possible features of an HNA
infrastructure that would address these needs. As mentioned before, the
latter were not regarded as direct assignments to be fulfilled in the visualization
process, but rather as indicators for the participant’s general receptiveness
towards various properties of the user interface of an HNA infrastructure.
Table 1 and 2 summarize the main aspects of the workshop discussions in
the form of needs and features.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Number of Mentions</title>
      </sec>
      <sec id="sec-5-4">
        <title>Persons</title>
        <p>The most pressing topic in the discussions was the envisioned user
approaches and use cases the infrastructure is expected to support. Seven of
the ten participants expressed the hope that the visualizations could
generate new perspectives, thereby creating forms of access to the data that would
hardly be available based on non-machine-supported cognitive work. In this</p>
      </sec>
      <sec id="sec-5-5">
        <title>Need</title>
        <sec id="sec-5-5-1">
          <title>New perspectives</title>
          <p>Uncertainty
Data potential
Graph density
Entry points
Data explanations
30
25
27
26
17
16
7
7
4
4
4
4</p>
        </sec>
        <sec id="sec-5-5-2">
          <title>Timeline</title>
          <p>Tie metrics
Other filters
Export and citation
Location filter
Source linking
22
18
16
13
8
6
4
4
3
4
3
4
context, one participant explicitly emphasized the potential of visualizations
to raise new questions:</p>
          <p>What kind of relationships you are looking for in the data is something
you often notice in the very moment that you look at the pile for the
ifrst time. 29</p>
          <p>Since the participants were aware of the fact that we are confronted with
a very large amount of data, which can hardly be presented in its entirety
(see Section 3), a discussion of possible entry points emerged. There was
consensus on the importance of filter options, most importantly time filters:
Without timelines, the visualizations are of no use to me – neither for
analysis, nor for the presentation of results.</p>
          <p>In addition to timelines, other filters (e. g. node type and node source)
were considered a prerequisite for data exploration. Three participants also
mentioned the importance of location filters (e. g. through a map view).</p>
          <p>Participants with more HNA experience explicitly stressed the essential
role of a multi-layered approach. The capacity to display the evolution of
relations (e. g. through time and location) was described as the distinctive
feature of HNA when compared to the non-historical analysis of social
networks. The sole option of static display was considered insuficient.</p>
          <p>Along with possible entry points, another important topic raised in the
discussion was data complexity, with introductions and explanations
regarding the underlying data being identified as particularly crucial. Some
participants suggested addressing this issue with the help of concrete use cases that
could give potential users a more specific idea of the possibilities aofrded by
the HNA infrastructure.</p>
          <p>29All quotes translated from German into English.</p>
          <p>About half of the participants cited the ability to quantify network
characteristics as graph metrics during the research process as one of their main
motivations for using HNA methods. This includes indicators such as the
clustering coeficient, closeness centrality, degree distribution, degree
centrality, and betweenness centrality. Four participants also considered
density within a selected sample of nodes to be a relevant indicator for a given
dataset’s potential for network analysis. After the first cluster of possible
approaches had been discussed, one participant highlighted the added value of
graph metrics when it comes to the identification of anomalies in the data:
What all these things are actually about is that we are looking for
patterns!</p>
          <p>Some participants also stressed the potential of tie metrics to
accommodate a variety of relation types, and expressed the desire to have the weight of
edge properties visualized:</p>
          <p>It is of course a big diefrence whether you are a family member […] or
whether you are a correspondence partner or whether you met at a
congress during a coefe break. These are all relationships, but of course
they have diefrent weights in their interpretation. This is, for example,
something we would like to see in the visualization.</p>
          <p>This statement is representative of another central topic discussed in the
workshop, namely the visual marking of missing or uncertain information
in the data which can, for example, be the result of inconsistencies in the
metadata fields (see Section 4.2). The design expert considered this to be a
major desideratum:</p>
          <p>I think this is not done enough in current visualizations to show
uncertainties of data.</p>
          <p>With regard to the scientific standards of HNA research, a final major
issue was the export and citation of the visualizations. This, of course,
requires unambiguous and persistent provenance links to the source of each
data point, as well as timestamps of the corresponding data import.</p>
          <p>
            Many of the results of our co-design workshop match the challenges in
information visualization discussed in the pertinent literature. In the
following section, we will draw on these results to describe our prototyping
approach and process.
The dataset of our project is comprised of a number of elements that go well
beyond what can be perceptually or cognitively grasped at a glance. When
it comes to encoding, for example, the sheer amount of nodes and relations
poses technological as well as visual challenges
            <xref ref-type="bibr" rid="ref11 ref31">(Fekete and Plaisant, 2002;
Shneiderman, 2008)</xref>
            . While some potential users of our technology might
have a very specific research question in mind, others might be inclined
towards a more serendipitous approach
            <xref ref-type="bibr" rid="ref32">(Thudt et al., 2012)</xref>
            , or may wish to use
such an infrastructure in order to formulate research questions. Our aim is
to provide access points for a broad variety of motivations and research
questions, including ones that we cannot as of yet anticipate. Therefore, the
conceptualization of a visual representation as an access point to our data in the
form of a data exploration interface can be described by a wide and diverse
range of challenges and dificulties:
• How can tens of millions of nodes and hundreds of millions of edges
be visualized?
• What are possible and meaningful entrance points to the data?
• How can we deal with uncertainty, missing data, and varying data
sources?
• How can we deal with multiple data dimensions?
• How can we provide a technology that is complex and open enough for
a broad range of undefined research questions, but simple enough for
casual use?
• How can we be transparent with regard to the algorithms used?
• How can users move between overviews, detail views, and egocentric
views?
          </p>
          <p>
            Even though our workshop, our conversations with domain experts, and
existing task taxonomies
            <xref ref-type="bibr" rid="ref1 ref19 ref21">(e. g. Lee et al., 2006; Kerracher et al., 2015; Ahn
et al., 2013)</xref>
            have already yielded a multitude of potential tasks, needs, and
requirements that should be addressed in our graph technology, we see the
prototyping process as a form of research through design (Zimmerman et al.,
2007) that is not only capable of confirming these requirements, but also of
unveiling new ones. Moreover, in contrast to the above-mentioned task
taxonomies, we are engaging with humanities-related data and research
questions – a field, where traditional visualization approaches are often deemed
incompatible with the nature of the objects of inquiry
            <xref ref-type="bibr" rid="ref9">Drucker (2011)</xref>
            .
Along with the data modeling process and the co-creation approaches
described above, our visualization process can thus be described as a form of
rapid, experimental, and iterative prototyping process and data exploration.
          </p>
          <p>
            Compared to the potentially shortest path to a finished ‘tool,’ our method
resembles a curiosity-driven ‘sandcastling’
            <xref ref-type="bibr" rid="ref15">(Hinrichs et al., 2019)</xref>
            . We
understand experimental approaches and detours in the visualization process
itself as a methodology of knowledge production. By following this route,
visualizations are not necessarily created with the goal of implementing them
in a final prototype or concept. Rather, they become a method for
exploring the data or individual facets of the data, a tool for investigating the
basic challenges of data or their encoding, or a visual facilitator for
encouraging cross-disciplinary communication and the development of novel and
thought-provoking approaches
            <xref ref-type="bibr" rid="ref15">(Hinrichs et al., 2019)</xref>
            .
          </p>
          <p>From the beginning, the entire project has been conducted in an
interdisciplinary and concurrent mode, without any delays between its individual
steps; data processing, case study development, visualization, and evaluation
all occur alongside one another. Initially, the data was neither processed for
visualization, nor was it accessible via some form of API, which only allowed
us to work with small subsets of selected data. While this made it dificult
to anticipate all of the facets and challenges associated with handling the full
extent of the data, working with data subsets early on gave us the ability to
exert iterative influence on the data processing and the data model.</p>
          <p>
            Instead of trying to combine all potential features and ideas into a single
prototype, our approach focuses on small, separate problems and ideas
through a multitude of rough prototypes. Many of our design studies or
prototypes have been developed in close collaboration with our own HNA
specialists, and/or draw extensively on input from our workshop or other
external sources of expertise, whereas others are more experimental in nature,
and are often the product of spontaneous impulses. For the most part, the
following examples were designed with the data visualization library D3.js
            <xref ref-type="bibr" rid="ref5">(Bostock et al., 2011)</xref>
            , which permitted the development of customized
visualizations.
          </p>
          <p>Figure 3, for instance, shows two small design studies from the beginning
of the project, without using real data: the one to the left is a visualization of
levels of relation uncertainty, while the one to the right represents the testing
of an interaction concept with the goal of reducing complexity by merging
multiple edges and allowing users to fan them out on demand.</p>
          <p>As an example of the influence exerted by visualization on the data model,
an early prototype which clusters persons in a small subset of the data based
on related topic terms – in most cases occupational titles – revealed that these
titles are frequently gendered30 in our base data (GND), which means that
men and women are often not related to the same topic term, even though
they practice the same profession. This unexpected diefrentiation in the
data is highly relevant when it comes to search queries and visualization,
since it is quite possible that some researchers do not diefrentiate by gender,
and only use the male form that was traditionally considered to be generic.
One eefct of this diefrentiation in the data can be seen in another
interactive prototype (see Figure 4), where it is possible to select a specific year in the
data with a slider, visualizing top topic terms related to persons who were
alive in the selected timespan (female-gendered topic terms are colored in
orange). The goal of this prototype was to explore the potential of overviews
to reveal aspects of the data that might, at a later point, act as entry points
30Many German occupational titles are gendered and exist in a male and a female form,
as with the English ‘actor’ and ‘actress.’
for specific search interests.</p>
          <p>
            Another experimental prototype (see Figure 5) of a small subset of our
data also focuses on topic terms and the temporality of the data; an aspect
frequently highlighted as important by some of our HNA experts in the
workshop. Here, the dimensionality reduction technique UMAP
            <xref ref-type="bibr" rid="ref23">(McInnes
et al., 2020)</xref>
            was used to map persons with similar topic term relations in
close proximity to each other, eefctively forming clusters for certain
occupational domains (e. g. authors). A timeline on the right displays the general
distribution of all nodes, while a list next to it contains all connected topic
terms, ordered by occurrence. Scrolling enables users to move through the
temporal dimension of the network, creating the impression of a time tunnel.
Nodes belonging to a selected year are displayed in yellow. Temporally close
nodes in the past appear more distant from the viewer and are marked in red
tones, while those that lie in the future are colored in green and blue tones,
and appear to be closer. One insight gained with the help of this prototype
was that our data model and processing approach once again needed to be
adjusted to make the data more accessible for use in visualizations, especially
with regard to temporal filtering.
          </p>
          <p>In some cases, as with Figure 6, we developed prototypes out of curiosity
for very specific research questions, for example: “Are network
communities in the data subset mostly composed of contemporary nodes or do
communities stretch over multiple generations?” Here, the prototyping process
allowed us to test specific algorithm implementations and design strategies,
while at the same time being able to obtain deeper insights into the data.</p>
          <p>While our research is still in progress, the experiences mentioned above
illustrate the benefits of staying curious and open to experimentation
throughout the analysis and visualization process. Even though many ideas and
concepts are inspired by existing research in the field and, of course, the expertise
of our domain specialists, we see additional value in experimenting with the
data and generating a multitude of visual representations, even if this means
knowingly taking detours. It is precisely these more experimental pathways
that can lead to new ideas for tools, or generate fresh insights into the data.
The prototypes are non-incremental steps towards a final concept, iteratively
informed by feedback from our domain experts and other potential future
operators.
7</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Projects</title>
      <p>The converging of multiple heterogeneous data sources containing millions
of nodes and edges for a graph-based research infrastructure that enables
historical social network analysis creates a plethora of multidisciplinary
challenges:
• Dificulties associated with the merging of heterogeneous data sources
• Performance of a system regarding the given scope and further scaling
of the data
• Creation of domain-customized interfaces, which are open and flexible
with regard to unforeseen research questions
• Integration of domain knowledge into the process
• Visualization of millions of data points to provide explorable access
points in addition to search interfaces</p>
      <p>We address these challenges by focusing on tight, interdisciplinary
collaboration and constant evaluation during the whole research and development
process. Our approach brings together historical network specialists, data
visualization researchers, data scientists, and experts on the evaluation of
information infrastructure, an important example being the initial co-design
workshop with additional external HNA practitioners and other domain
experts. Building on the contextual data gathered during the co-design
workshop, we will continue to follow a human-centered approach towards data
modeling and visualization design.</p>
      <p>In our next step, we aim to take a closer look at the individual processes
behind historical network research in one-on-one interviews with domain
experts concerning their approaches to HNA research. Our plans for the
future also include the merging of multiple visualization concepts into one
prototype, which will join global overviews of our data with local views of
specific individual networks inside it. Furthermore, we will make use of our
data and our interface to provide exemplary use cases on a variety of historical
topics in collaboration with our HNA experts.</p>
      <p>Finally, we are experimenting with linked data as an alternative to Neo4j.
Here, the source data would be modeled in the form of subject–predicate–
object expressions and stored in GraphDB.31 This approach would simplify
the integration of Linked Open Data datasets (Wikidata, DBpedia,32
GeoNames,33 etc.), and would provide more sophisticated inference possibilities.
In preliminary comparisons of the two approaches, GraphDB also shows
better performance, but employing it would mean that the source data must be
remodeled in order to display relation features such as relation type, relation
source type, and temporal validity.</p>
      <p>In this paper, we have described the process of examining the potential
of remodeling and merging (bibliographic) big data from cultural heritage
institutions into one single gathering point optimized for the use in
historical network analysis. It is our hope that by providing insights into emerging
challenges and outlining possible solutions, we can encourage additional
research and scholarly exchange in and with similar HNA-related projects.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We would like to thank the participants of our co-design workshop and our
project partners Heiner Fangerau, Katrin Getschmann, Thorsten Halling,
Hans-Jörg Lieder, Gerhard Müller, Clemens Neudecker, David Zellhöfer,
and Josefine Zinck. This research is part of the research project SoNAR
(IDH) and is funded by the DFG – German Research Foundation (project
no. 414792379).</p>
      <p>The Discovery of Grounded Theory.
Zimmerman, J., Forlizzi, J., and Evenson, S. (2007). CHI ’07: Proceedings
of the SIGCHI Conference on Human Factors in Computing Systems.
In Proceedings of the SIGCHI conference on Human factors in computing
systems, pages 493–502. DOI: 10.1145/1240624.1240704.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Ahn</surname>
          </string-name>
          , J.-w.,
          <string-name>
            <surname>Plaisant</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Shneiderman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>A Task Taxonomy for Network Evolution Analysis</article-title>
          .
          <source>IEEE transactions on visualization and computer graphics</source>
          ,
          <volume>20</volume>
          (
          <issue>3</issue>
          ):
          <fpage>365</fpage>
          -
          <lpage>376</lpage>
          , DOI: 10.1109/TVCG.
          <year>2013</year>
          .
          <volume>238</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Allison-Bunnell</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Review of Encoded Archival Description Tag Library - Version EAD3</article-title>
          .
          <source>Journal of Western Archives</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>3</lpage>
          , DOI: 10.26077/af62-
          <fpage>2a86</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Towards an Open Research Knowledge Graph</article-title>
          .
          <source>The Serials Librarian</source>
          ,
          <volume>76</volume>
          (
          <issue>1-4</issue>
          ):
          <fpage>35</fpage>
          -
          <lpage>41</lpage>
          , DOI: 10.1080/0361526X.
          <year>2019</year>
          .
          <volume>1540272</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Behrisch</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bach</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henry Riche</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , et al. (
          <year>2016</year>
          ).
          <article-title>Matrix Reordering Methods for Table and Network Visualization</article-title>
          . In Computer Graphics Forum, volume
          <volume>35</volume>
          , pages
          <fpage>693</fpage>
          -
          <lpage>716</lpage>
          . Wiley, DOI: 10.1111/cgf.12935.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Bostock</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogievetsky</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Heer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>D3 Data-Driven Documents</article-title>
          .
          <source>IEEE transactions on visualization and computer graphics</source>
          ,
          <volume>17</volume>
          (
          <issue>12</issue>
          ):
          <fpage>2301</fpage>
          -
          <lpage>2309</lpage>
          , DOI: 10.1109/TVCG.
          <year>2011</year>
          .
          <volume>185</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Cantara</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>METS: The Metadata Encoding</article-title>
          and
          <string-name>
            <given-names>Transmission</given-names>
            <surname>Standard</surname>
          </string-name>
          .
          <source>Cataloging &amp; classification quarterly</source>
          ,
          <volume>40</volume>
          (
          <issue>3-4</issue>
          ):
          <fpage>237</fpage>
          -
          <lpage>253</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coleman</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christensen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Heer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Visualizing the Republic of Letters</article-title>
          . https://web.stanford.edu/group/ toolingup/rplviz/papers/Vis_RofL_
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.-I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dörk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Dade-Robertson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Exploring the Promises and Potentials of Visual Archive Interfaces</article-title>
          .
          <source>In iConference 2014 Proceedings</source>
          , pages
          <fpage>735</fpage>
          -
          <lpage>741</lpage>
          . DOI:
          <volume>10</volume>
          .9776/14348.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Drucker</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Humanities Approaches to Graphical Display</article-title>
          .
          <source>DHQ: Digital Humanities Quarterly</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Efer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Graphdatenbanken für die textorientierten e-Humanities</article-title>
          .
          <source>PhD thesis</source>
          , Universität Leipzig, https://nbn-resolving.org/urn:nbn:de:bsz:
          <fpage>15</fpage>
          -qucosa-219122.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Fekete</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Plaisant</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Interactive Information Visualization of a Million Items</article-title>
          .
          <source>In IEEE Symposium on Information Visualization</source>
          ,
          <string-name>
            <surname>INFOVIS</surname>
          </string-name>
          <year>2002</year>
          ., pages
          <fpage>117</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Gibson</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faith</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vickers</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>A Survey of Two-Dimensional Graph Layout Techniques for Information Visualisation</article-title>
          .
          <source>Information Visualization</source>
          ,
          <volume>12</volume>
          (
          <issue>3-4</issue>
          ):
          <fpage>324</fpage>
          -
          <lpage>357</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Glaser</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Strauss</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>1967</year>
          ).
          <source>Weidenfield &amp; Nicolson</source>
          , London.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Henry</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Fekete</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-D.</surname>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Matrixexplorer: a Dual-Representation System to Explore Social Networks</article-title>
          .
          <source>IEEE Transactions on Visualization and Computer Graphics</source>
          ,
          <volume>12</volume>
          (
          <issue>5</issue>
          ):
          <fpage>677</fpage>
          -
          <lpage>684</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Hinrichs</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forlini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Moynihan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <source>In Defense of Sandcastles: Research Thinking through Visualization in Digital Humanities. Digital Scholarship in the Humanities</source>
          ,
          <volume>34</volume>
          (
          <issue>1</issue>
          ):
          <fpage>i80</fpage>
          -
          <lpage>i99</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Isenberg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuk</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Collins,
          <string-name>
            <given-names>C.</given-names>
            , and
            <surname>Carpendale</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Grounded Evaluation of Information Visualizations</article-title>
          .
          <source>In Proceedings of the 2008 Workshop on BEyond Time and Errors: Novel EvaLuation Methods for Information Visualization</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . DOI:
          <volume>10</volume>
          .1145/1377966.1377974.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Jacomy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venturini</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heymann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bastian</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>ForceAtlas2, a Continuous Graph Layout Algorithm for Handy Network Visualization Designed for the Gephi Software</article-title>
          .
          <source>PloS one</source>
          ,
          <volume>9</volume>
          (
          <issue>6</issue>
          ):e98679, DOI: 10.1371/journal.pone.
          <volume>0098679</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Jansen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Wald</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Netzwerktheorien</article-title>
          . In Benz,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Lütz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Schimank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            , and
            <surname>Simonis</surname>
          </string-name>
          , G., editors,
          <source>Handbuch Governance: Theoretische Grundlagen und empirische Anwendungsfelder</source>
          , pages
          <fpage>188</fpage>
          -
          <lpage>199</lpage>
          . VS Verlag für Sozialwissenschaften, Wiesbaden, DOI: 10.1007/978-3-
          <fpage>531</fpage>
          -90407-8_
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Kerracher</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kennedy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Chalmers</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>A Task Taxonomy for Temporal Graph Visualisation</article-title>
          .
          <source>IEEE transactions on visualization and computer graphics</source>
          ,
          <volume>21</volume>
          (
          <issue>10</issue>
          ):
          <fpage>1160</fpage>
          -
          <lpage>1172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Kruk</surname>
            ,
            <given-names>S. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Synak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zimmermann</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>MarcOnt - Integration Ontology for Bibliographic Description Formats</article-title>
          .
          <source>In International Conference on Dublin Core and Metadata Applications</source>
          , pages
          <fpage>231</fpage>
          -
          <lpage>234</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plaisant</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parr</surname>
            ,
            <given-names>C. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fekete</surname>
          </string-name>
          , J.
          <string-name>
            <surname>-D.</surname>
          </string-name>
          , et al. (
          <year>2006</year>
          ).
          <article-title>Task Taxonomy for Graph Visualization</article-title>
          .
          <source>In Proceedings of the 2006 Workshop on BEyond Time and Errors: Novel Evaluation Methods for Information Visualization</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . DOI:
          <volume>10</volume>
          .1145/1168149.1168168.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Matschinegg</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Nicka</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <source>REALonline Enhanced</source>
          .
          <article-title>Die neuen Funktionalitäten und Features der Forschungsbilddatenbank des IMAREAL</article-title>
          .
          <source>MEMO</source>
          ,
          <volume>2</volume>
          :
          <fpage>10</fpage>
          -
          <lpage>32</lpage>
          , DOI: 10.25536/20180202.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>McInnes</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Healy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Melville</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction</article-title>
          . arXiv,
          <year>1802</year>
          .
          <volume>03426</volume>
          , https://arxiv.org/abs/
          <year>1802</year>
          .03426.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Molina</given-names>
            <surname>León</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            and
            <surname>Breiter</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Co-creating Visualizations: A First Evaluation with Social Science Researchers</article-title>
          . Computer Graphics Forum,
          <volume>39</volume>
          (
          <issue>3</issue>
          ):
          <fpage>291</fpage>
          -
          <lpage>302</lpage>
          , DOI: 10.1111/cgf.13981.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Munzner</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>A Nested Model for Visualization Design and Validation</article-title>
          .
          <source>IEEE Transactions on Visualization and Computer Graphics</source>
          ,
          <volume>15</volume>
          (
          <issue>6</issue>
          ):
          <fpage>921</fpage>
          -
          <lpage>928</lpage>
          , DOI: 10.1109/TVCG.
          <year>2009</year>
          .
          <volume>111</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Micheel</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melenhorst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wieneke</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , et al. (
          <year>2014</year>
          ).
          <article-title>HistoGraph - A Visualization Tool for Collaborative Analysis of Networks from Historical Social Multimedia Collections</article-title>
          .
          <source>In 18th International Conference on Information Visualisation</source>
          , pages
          <fpage>241</fpage>
          -
          <lpage>250</lpage>
          . DOI:
          <volume>10</volume>
          .1109/IV.
          <year>2014</year>
          .
          <volume>47</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Pienta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abello</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kahng</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Chau</surname>
            ,
            <given-names>D. H.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Scalable Graph Exploration and Visualization: Sensemaking Challenges and Opportunities</article-title>
          .
          <source>In 2015 International Conference on Big Data and Smart Computing (BIGCOMP)</source>
          , pages
          <fpage>271</fpage>
          -
          <lpage>278</lpage>
          . DOI:
          <volume>10</volume>
          .1109/35021BIGCOMP.
          <year>2015</year>
          .
          <volume>7072812</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Pienta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hohman</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamersoy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Endert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al. (
          <year>2017</year>
          ).
          <article-title>Visual Graph Query Construction and Refinement</article-title>
          .
          <source>In SIGMOD '17: Proceedings of the 2017 ACM International Conference on Management of Data</source>
          , pages
          <fpage>1587</fpage>
          -
          <lpage>1590</lpage>
          . DOI:
          <volume>10</volume>
          .1145/3035918.3056418.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Purschwitz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Netzwerke des Wissens - Thematische und personelle Relationen innerhalb der halleschen Zeitungen und Zeitschriften der Aufklärungsepoche (1688-1818)</article-title>
          .
          <source>Journal of Historical Network Research</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>109</fpage>
          -
          <lpage>142</lpage>
          , DOI: 10.25517/jhnr.v2i1.
          <fpage>47</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Rehm</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elsholz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hegele</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al. (
          <year>2020</year>
          ).
          <article-title>European Language Grid: An Overview</article-title>
          . In Calzolari, N.,
          <string-name>
            <surname>Béchet</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blache</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cieri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , et al., editors,
          <source>Proceedings of the 12th Language Resources and Evaluation Conference (LREC</source>
          <year>2020</year>
          ), pages
          <fpage>3359</fpage>
          -
          <lpage>3373</lpage>
          . https://www.aclweb.org/ anthology/2020.lrec-
          <volume>1</volume>
          .413/.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Shneiderman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Extreme Visualization: Squeezing a Billion Records into a Million Pixels</article-title>
          .
          <source>In SIGMOD '08: Proceedings of the 2008 ACM SIGMOD international conference on Management of data</source>
          , pages
          <fpage>3</fpage>
          -
          <lpage>12</lpage>
          . DOI:
          <volume>10</volume>
          .1145/1376616.1376618.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Thudt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinrichs</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Carpendale</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>The Bohemian Bookshelf: Supporting Serendipitous Book Discoveries through Information Visualization</article-title>
          .
          <source>In CHI '12: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source>
          , pages
          <fpage>1461</fpage>
          -
          <lpage>1470</lpage>
          . DOI:
          <volume>10</volume>
          .1145/2207676.2208607.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>van Ham</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Perer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2009</year>
          ). ”Search, Show Context, Expand on Demand”
          <article-title>: Supporting Large Graph Exploration with Degree-of-Interest</article-title>
          .
          <source>IEEE Transactions on Visualization and Computer Graphics</source>
          ,
          <volume>15</volume>
          (
          <issue>6</issue>
          ):
          <fpage>953</fpage>
          -
          <lpage>960</lpage>
          , DOI: 10.1109/TVCG.
          <year>2009</year>
          .
          <volume>108</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>von Landesberger</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuijper</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kohlhammer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al. (
          <year>2011</year>
          ).
          <article-title>Visual Analysis of Large Graphs: State-of-the-Art and Future Research Challenges</article-title>
          . Computer Graphics Forum,
          <volume>30</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1719</fpage>
          -
          <lpage>1749</lpage>
          , DOI: 10.1111/j.1467-
          <fpage>8659</fpage>
          .
          <year>2011</year>
          .
          <year>01898</year>
          .x.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Vorndran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Hervorholen, was in unseren Daten steckt! Mehrwerte durch Analysen großer Bibliotheksdatenbestände. o-bib</article-title>
          .
          <source>Das ofene Bibliotheksjournal</source>
          ,
          <volume>5</volume>
          (
          <issue>4</issue>
          ):
          <fpage>166</fpage>
          -
          <lpage>180</lpage>
          , DOI: 10.5282/o-bib/2018H4S166-
          <fpage>180</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Warren</surname>
            ,
            <given-names>C. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shore</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Otis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Finegold</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Shalizi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Six Degrees of Francis Bacon: A Statistical Method for Reconstructing Large Historical Social Networks</article-title>
          .
          <source>DHQ: Digital Humanities Quarterly</source>
          ,
          <volume>10</volume>
          (
          <issue>3</issue>
          ), DOI: 10.17613/M6B020.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Wintergrün</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Netzwerkanalysen und semantische Datenmodellierung als heuristische Instrumente für die historische Forschung</article-title>
          .
          <source>PhD thesis</source>
          ,
          <string-name>
            <surname>Friedrich-Alexander-Universität</surname>
          </string-name>
          Erlangen-Nürnberg, https: //nbn-resolving.org/urn:nbn:de:bvb:
          <fpage>29</fpage>
          -
          <lpage>opus4</lpage>
          -111899.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>