<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building a Language Technology Infrastructure for Digital Humanities: Challenges, Opportunities and Progress</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dana Danne´lls</string-name>
          <email>dana.dannells@svenska.gu.se</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Brode´n</string-name>
          <email>daniel.broden@lir.gu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of History of Ideas</institution>
          ,
          <addr-line>Literature and Religion</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Spra ̊kbanken Text/Department of Swedish University of Gothenburg</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Swe-Clarin/Centre for Digital Humanities University of Gothenburg</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>75</fpage>
      <lpage>83</lpage>
      <abstract>
        <p>Spra˚kbanken Text, a research unit at the University of Gothenburg, forms part of the National Language Bank of Sweden and is the main coordinating node of Swe-Clarin, the Swedish national CLARIN node. During the past years, Spra˚kbanken Text has been actively engaged in a number of humanities and social sciences related research projects. This engagement has primarily concerned the development of new resources, methods and tools to accurately process large amounts of digitized material, in addition to interfaces for visualizing the materials, making them easily accessible for further analysis. The activities within Swe-Clarin have been essential for the progress and the success of this work. In this paper we present what was required from Spra˚kbanken Text in order to meet the expectations of researchers from the humanities and social sciences. We discuss some of the challenges this work involves and describe the opportunities this field brings with it and how these opportunities could help to progress the work of Spra˚kbanken Text toward building a language technology infrastructure that supports interdisciplinary research.</p>
      </abstract>
      <kwd-group>
        <kwd>Digital Humanities</kwd>
        <kwd>Language Technology</kwd>
        <kwd>National infrastructure</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>The humanities research problem</title>
      <p>Researchers coming from digital humanities and its adjacent fields are in need of
sophisticated tools, interfaces and materials to help them answer their current research
questions and form new ones.</p>
      <p>As a prominent institution in the field of language technology (LT), Spra˚kbanken
Text, a research unit at the University of Gothenburg, is on a regular basis approached
by researchers from the field of humanities but also the social sciences with various
methodological questions and feature requests. Consequently, over the years,
Spra˚kbanken Text has nurtured a growing interest in projects combining a language technology
approach with interpretative methods associated with other fields. While such
collaborative efforts have contributed to positioning Spra˚kbanken Text centrally in the Nordic
digital humanities community, they have also provided new challenges.</p>
      <p>A key question for Spra˚kbanken Text today is how to continue developing language
technology tools and services and at the same time answer to the needs and wishes
of the wider research community. As critical commentators have stressed, while
research infrastructure projects are vital for the introduction of computational methods
for humanities research, there needs to be a strong focus on intensifying and allowing
”creative forms of humanities research for the 21th century, not their replacement by
science as a hegemonic form of knowledge creation” [2, page 84].</p>
      <p>In this paper, we address the specific question of what is required from a language
technology infrastructure, such as Spra˚kbanken Text, in order to truly meet the
expectations coming from the humanities and social sciences (HSS) researchers. More
specifically, what actual steps and activities should be considered and above all
prioritized in this context? We also address the question of how these measures contribute
to helping language technology to progress with the work towards building a national
infrastructure that supports truly interdisciplinary research.</p>
      <p>We begin by presenting ongoing efforts at Spra˚kbanken Text, where a major focus
has been on developing and enhancing the lexical and corpora resource repositories
and improving the tools to access them.4 The enhancements and improvements are
partly the results of close collaboration and dialogue with, among others, historians,
librarians, rhetoricians, and literary scholars. Because the lexical and corpora resources
of Spra˚kbanken Text have been developed in interdisciplinary collaboration with
researchers from different fields who possess different perspectives and knowledge about
which pieces of information are important to encode, it has proved to be a challenging
task to apply computational methods that are able to access the resources uniformly, as
well as to develop tools that allow users without substantial programming skills to use
the resources and to present the analysis results.</p>
      <p>After this, we turn to the user involvement activities of Swe-Clarin, the Swedish
CLARIN node, of which Spra˚kbanken Text is the main coordinating partner, focusing
in particular on interdisciplinary pilot projects and user workshops. In order to spread
awareness of the research potentials of LT tools, Swe-Clarin and Spra˚kbanken Text has
initiated a series of one-off projects, in which we have collaborated with researchers
from various fields of the humanities. Besides the research results and their
dissemination, one vital outcome has been further insights into the needs and challenges posed by
scholars from other disciplinary domains.
1.1</p>
      <p>Spra˚kbanken Text – a Growing Swedish Infrastructure
Spra˚kbanken Text carries research on language technology for historical and modern
Swedish, and develops a research infrastructure to support research in linguistics and
other disciplines in the humanities, social sciences, and medicine.</p>
      <p>
        Spra˚kbanken Text was established by a governmental decree as a national centre
of computational lexicography (then named Logoteket) in 1975. Through the centre,
4 All the resources and tools at Spra˚kbanken Text are CC-BY licensed.
76/143
computational linguists and corpora users, in and outside of Sweden, have been able to
access linguistic and statistical data about a wide range of texts in Swedish for more
than four decades [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. During the 2000s, the work at Spra˚kbanken Text rapidly
expanded towards the development of high quality language technology tools. The first
version of the widely used corpus tool Korp [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] was launched in 2011. Today, Korp is
used in Sweden as well as in several other countries and has been modified to support
other languages, among others by the Centre for Language Technology in Copenhagen
and by the Language Bank of Finland, Kielipankki. Other tools, also named after birds
and other animals, have followed, including the annotation tool Sparv [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and the
lexical platform Karp [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Furthermore, since 2014 Spra˚kbanken Text is officially the main
coordinating node of Swe-Clarin, the Swedish national node in the European Research
Infrastructure Consortium CLARIN (Common Language Resources and Technology
Infrastructure). As a part of, and a so-called B-centre, the CLARIN ERIC consortium,
Spra˚kbanken Text offers primarily HSS researchers access to a range of tools and
corpora as well as services and knowledge on a sustainable basis.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The solution</title>
      <p>
        The work at Spra˚kbanken Text is a collaborative effort between researchers, so called
experts, and system engineers. The daily work primarily comprises: (a) collection and
development of digital resources; (b) development of tools and metadata standards to
annotate and classify the material and link it to other resources; (c) implementation and
maintenance of interfaces to visualize and navigate between the resources; and (d)
analysis of the content of the resources by applying state-of-the-art methods. All of these
are essential tasks for meeting the expectations coming from HSS researchers [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. In
the following, we draw on some them with emphasis on how the work at Spra˚kbanken
Text is progressing.
      </p>
      <p>
        Resources As the result of mass digitization of historical and literary data, a wide
range of new electronic resources have been added to Spra˚kbanken Text’s repository.
One recently added resource is the Kubhist corpus of Swedish historical newspapers,
spanning from 1645 until 1926, and containing around 5.5 billion tokens [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It is a
rich resource for conducting diachronic studies and for examining various research
questions concerning language and culture. Even though this new Kubhist has been
digitized by the National Library of Sweden by applying state-of-the-art OCR
technology, it contains many OCR-errors which pose new challenges to the annotation tools of
Spra˚kbanken Text. While there is an ongoing initiative to improve the OCR processes
of Swedish newspapers [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], accessing the data in meaningful ways still requires specific
methods and competencies to improve the performance of the analysis [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Meanwhile,
regardless of the OCR-errors, the material could be explored to answer important
research questions about language properties such as: What are the meanings and
associations of different concepts at different points in time? What spelling variation is most
prominent during a certain time period?
      </p>
      <p>Two other complex and large resources that have been added to the resource
repository of Spra˚kbanken Text are the Twitter and Flashback corpora. Social media data
is extremely diverse and the sensitive information it contains makes it compelling to
77/143
explore, for example, for sociologists or media scholars who are interested in finding
value-laden, harsh words directed towards journalists and other public figures. Typical
research questions this material could potentially provide answers to are: How have
the statements made by a certain person affected his/her public persona? What kind of
attention has a certain public figure attracted and through what social media activities?</p>
      <p>
        A valuable lexical resource, the Swedish sentiment lexicon [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], was developed
and added to the repository of lexical resources of Spra˚kbanken Text. The resource
is forming the first step towards the creation of a full-fledged sentiment analysis for
Swedish. It has been proven in particular useful for analyzing opinions on social media.
      </p>
      <p>
        Spra˚kbanken Text has ongoing research collaborations with the Department of
Historical Studies, the Department of Languages and Literature and the Centre for
Digital Humanities at the University of Gothenburg. In relation to these collaborations,
Spra˚kbanken Text develops and maintains several databases, including the
Biographical Dictionary of Swedish Women, a historical database containing biographies of
women who have made significant contributions to Swedish society and culture,5 and
the NordiCon database, containing medieval Nordic personal names attested in
Continental sources [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].6 Spra˚kbanken Text provides easy access and editing possibilities
of these databases though Karp, Spra˚kbanken Text’s lexical infrastructure [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Drawing
on the data from these databases one could ask research questions such as: How has the
professional life of women in Swedish societies progressed the last couple of hundred
years? What are the naming preferences among different religious communities? Are
there any naming traditions that characterizes certain time periods?
      </p>
      <p>
        Annotations Spra˚kbanken Text’s annotation pipeline is available through the tool
Sparv [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and over the years the pipeline has been increased with several annotation
layers. Sentiment annotation is an example of an annotation layer that has been added to the
pipeline [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], as one outcome of the Swe-Clarin activities. Furthermore, Sparv has been
empowered with a user-friendly interface with possibilities to upload a digitized
material in various formats and to automatically enhance it with different annotations such as
sentiment, name entity, part-of-speech. The annotation results are displayed through the
interface and can also be downloaded. Sparv’s user-interface has been proven valuable
for humanities scholars, and by following their requests it is constantly evolving.
      </p>
      <p>Analysis Korp is Spra˚kbanken Text corpus infrastructure tool. The tool comes with
a graphical web user interface and functionalities for importing, annotating and
exporting corpus resources. Users can explore the resources they are interested in by, for
example, compiling statistics over the data and visualize the results with graphs,
geographical maps or ”word pictures”.</p>
      <p>Word picture is a functionality in Korp for visualizing word correlations in a text.
Word picture has been widely used to explore various research questions related to
conceptual transformations as well as to specific issues concerning equality or
discrimination of old people or people with disabilities. Some examples are: When did certain
words come into use, and when did they disappear? What words are used to describe
men and women today compared to the early 20th century?
5 https://skbl.se/en
6 https://spraakbanken.gu.se/karp/tng/?mode=nordicon&amp;lexicon=
nordicon
78/143</p>
      <p>
        Word picture has also been used to help scholars of political science to identify what
different topics are prominently discussed by different political parties [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], answering
research questions such as: What are the most frequent questions discussed under a
certain election year? How does the public interest in certain topics change over time?
      </p>
      <p>
        Visualization of geographical locations is another example of a functionality that
has been added to generate maps marked with the locations of the place names that are
are mentioned in the text. It can be exploited for answering research questions such as:
Which geographical locations are mentioned in the work of a particular author and in
what contexts? What do we know about the the etymology and history of a particular
place name, and how has its geographical associations changed over time? What place
names no longer exists? Such questions have received a great interest over the past
years, especially from literary scholars. However, because of annotation errors, spelling
variation, and lack of normalization, the automatic analysis unfortunately still fails to
recognize many significant place names [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. As a result, it is not always possible to
explore the research questions to a satisfactory extent [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. What adds to this
insufficiency is the lack of information about the accuracy of the results, and a functionality
that allows users to inspect the material through close reading.
      </p>
      <p>Recently, Korp was enhanced with a functionality that enables researchers to
conduct manual analysis of video transcriptions. This was developed as part of a
collaboration with a group of researchers at the Institute of Language and Folklore in Sweden,
who are working on a project entitled Interaction and Variation in Pluricentric
Languages (IVIP). The new functionality allows, for example, anthropologists or
sociologists to closely study how interpersonal relationships are expressed in institutional
conversations in Sweden and Finland in the domains of service, education and healthcare.
Possibilities for close readings of the text and for analyzing it directly on the screen is
a functionality that has proved to be highly appreciated.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The collaboration experience</title>
      <p>
        As mentioned above, Spra˚kbanken Text is a also driving force in the user involvement
activities of Swe-Clarin, that aim at spreading awareness of what research possibilities
language technology tools can offer HSS scholars who use text and speech as primary
research data. Digital humanities projects are often conducted with either strong data
science or humanities bias [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Thus, one of the central means for spreading
awareness of what research possibilities Swe-Clarin’s digital tools can offer and at the same
time meet the requirements of HSS researchers is collaborative pilot projects.
SweClarin has initiated a series of interdisciplinary one-off projects, in which scholars and
technicians from Spra˚kbanken Text collaborate with HSS researchers from various
disciplinary fields. The projects have concerned, among other things, rhetoric,
secondlanguage acquisition and political science [
        <xref ref-type="bibr" rid="ref13 ref19">13, 19</xref>
        ].
      </p>
      <p>
        A recent historical study [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] used text mining of the large newspaper corpus
Kubhist to study the emergence of terrorism in the Swedish newspaper discourse from late
18th to early 20th century. The aim was partly to evaluate prior research claims
regarding the historical connotations and ideological contexts of terrorism before World War I.
      </p>
      <p>The study could confirm what we already knew about the meanings of terrorism from
79/143
the historical record, but also go beyond common historical wisdom in pointing to the
diversity of the meaning of terrorism in the period. For example, the word pictures of
”terrorism” and ”terrorist” had more prominent attributions of so-called state terrorism
than expected. However, it should be noted that even with the large quantity of text in
the Kubhist corpus, many of the queries returned only a few hundred hits, which does
not allow for solid generalizations.</p>
      <p>
        Besides the specific research results and their dissemination, one vital outcome of
the user involvement activities of Swe-Clarin have been concrete insights into the needs
of HSS researchers and into the further development of Spra˚kbanken Text’s resources to
meet such demands. Apropos the Kubhist corpus, the researchers involved are looking
forward to continuing the work on historical discourses on terrorism in Sweden by
looking at an expanded version of Kubhist. The new data-set is five times as large and
have better OCR quality [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a factor which affects search accuracy [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The updated
Kubhist corpus also allows for tests of the usefulness of diachronic word embeddings
for studying the change in meanings of terms related to terrorism and other concepts
over time. Furthermore, the pilot project on the conceptual history of terrorism has
pointed to the need of supporting comparative analysis with versions of Korp installed
in other countries. As the lead researcher has an interest in both Swedish and Finnish
history, the project’s focus has now been extended to include a comparative analysis
of a Finnish newspaper corpus (in the Swedish language). This requires content search
involving both the Swedish Korp tool and the version used by the Language Bank of
Finland, similar to the more general federated content search that Spra˚kbanken Text
already provides to CLARIN.
      </p>
      <p>
        Another important user involvement activity of the CLARIN endeavor is the
organization of user workshops. Although Spra˚kbanken Text and other members of the
Swe-Clarin consortium provide various digital tools and materials for humanities
scholars to use in their research, the resources are hardy used by all who would benefit from
them [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The Swe-Clarin workshops (nicknamed “Swe-Clarin on tour”), which are
held at universities and memory institutions around Sweden, give hands-on training in
the use of some of the key tools of Spra˚kbanken Text, such as Korp and Sparv, and in
the formulation of productive research questions.
      </p>
      <p>
        However, the aim of the workshops is not only to promote a wider utilization of
Swe-Clarin’s resources, but also to get user-feedback. Consequently they are carried
out with an active interest in the research perspectives that the participants bring to the
table. Similar to the collaborative pilot projects, the workshop dialogues have given
us access to a body of perspectives and experiences that have provided insight into
current limitations of our tools, from bugs to functionality gaps. A recurring feedback
from the workshops has been the need for document-level access to the text materials
provided by Spra˚kbanken Text. While researchers in language technology have limited
interest in, for example, cultural and historical content aspects of our corpora, such
aspects are elemental to many researchers in the humanities. Thus, Spra˚kbanken Text
has initiated extensive work to meet the need for document-level access to texts, the
most prominent example being the further development of the Korp tool [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Looked at
from this perspective, as much as the workshops are about providing hands-on training
in the use of our tools and in the formulation of productive research questions, they are
80/143
also about gathering feedback that helps us to develop the tools and data-sets for the
needs of HSS scholars.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and recommendations to educators</title>
      <p>The various collaborative projects tied to the activities of Spra˚kbanken Text and
SweClarin, in addition to the feedback from user workshops, has given us access to a body
of perspectives and experiences that have provided valuable insights into current
limitations of our tools and resources and contributed to their further development. Above
all, they made us aware of the importance of integrating interdisciplinary perspectives
and cooperations at different levels of digital resources and tools development.</p>
      <p>As the results of extensive Swe-Clarin activities, Spra˚kbanken Text has become
aware of a range of desirable functionalities through input from political science
scholars, historians, sociologists, rhetoricians, and ethnologists, who appreciate our
infrastructure but have needs and research questions that can not currently be answered
because of various technical and computational limitations. While some of the
limitations were outlined in Section 2, we would like to round-off with highlighting some
others that are currently lacking but will hopefully be added to the infrastructure of
Spra˚kbanken Text.</p>
      <p>Real-time data analysis Social data contains huge amounts of up-to-date
information and should therefore be analyzed the moment the data becomes available.
Spra˚kbanken Text does not yet have any support for uploading the data and for performing real
time analysis. Moreover, currently, real-time calculations over a large amount of
material take too long to be included in an interactive interface, making it difficult for
investigating synchronic and diachronic questions which requires large data-sets.</p>
      <p>Semantic similarity analysis We have so far been able to significantly improve
word picture functionalities thanks to feedback and requests from researcher and users.
Still, identifying words with semantic similarity and comparing their usages in positive
and negative sense is one example of a user request that has been suggested to enhance
the usability of word picture. Another suggested improvement of the word picture
function is to generalize it to allow for comparison between the different annotations.</p>
      <p>Upload ones own material and access it through the interfaces Researchers
outside Spra˚kbanken Text are still unable to upload their own material and explore the
annotation results through our interfaces. Materials that have subsequently been
analyzed should also be made available for download in a number of different formats for
import in other tools than those of Spra˚kbanken Text. This would allow researchers to
further explore the materials, but also to continue working with their own research data,
for example, by running analysis against existing collections at Spra˚kbanken Text. In
addition, while functionality for automatically identifying and annotating lexical
entities in a text is available in the annotation pipeline of the tools of Spra˚kbanken Text,
it could be improved to allow for searches across internal and external resources. The
merits of offering such a platform is two-sided; not only will HSS researchers benefit
from it, but it will also be beneficial for Spra˚kbanken Text by increasing our resource
repository with valuable data-sets. Because there are various formats for how different
materials are stored, we are interested in what formats different researchers prefer.
81/143</p>
      <p>Normalization of articles Normalization is a necessary next step for drawing
objective conclusions about the analysis results, and for answering questions such as how
many authors are represented in the data or how many texts have been written by a
specific author. However, this is a user request that requires substantial technical efforts,
but becoming aware of the types of normalization that researchers are primarily
interested in might help the developers at Spra˚kbanken Text with finding plausible technical
solutions.</p>
      <p>As pointed out earlier in this paper, the daily work at Spra˚kbanken Text is carried
out by a group of experts and research engineers who are collaborating actively to take
in user requests and generalize them to find suitable solutions and appropriate
methods that could be made available through our interfaces. Thus, knowing the needs and
wishes of users is an important step towards making the infrastructure accessible to a
larger group of HSS researchers. Conducting pilot studies with researchers about their
specific research interests is also a step toward spreading wider awareness in the HSS
community about the potentials of language technology in research as well as further
identifying user needs, and in extension, toward strengthening the capabilities of the
infrastructure of Spra˚kbanken Text and the National Language Bank of Sweden for all
of their users.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>The research presented here is supported by Spra˚kbanken Text and Swe-Clarin, a Swedish
consortium in Common Language Resources and Technology Infrastructure (CLARIN)
Swedish CLARIN (grant agreement 821-2013-2003).
82/143</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Adesam</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , Danne´lls,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Tahmasebi</surname>
          </string-name>
          , N.:
          <article-title>Exploring the Quality of the Digital Historical Newspaper Archive KubHist</article-title>
          .
          <source>In: Proceedings of the 4th Conference of The Association Digital Humanities in the Nordic Countries (DHN)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Berry</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fagerjord</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Digital Humanities: Knowledge and Critique in a Digital Age</article-title>
          . Polity Press, London (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Danne´lls,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Olsson</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.J.:</surname>
          </string-name>
          <article-title>Geographic visualization of place names in Swedish literary texts</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          <volume>29</volume>
          (
          <issue>3</issue>
          ),
          <fpage>400</fpage>
          -
          <lpage>404</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forsberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hammarstedt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Rose´n,
          <string-name>
            <given-names>D.</given-names>
            , Scha¨fer, R.,
            <surname>Schumacher</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Sparv: Spra˚kbanken's corpus annotation pipeline infrastructure</article-title>
          .
          <source>In: SLTC</source>
          <year>2016</year>
          .
          <article-title>The Sixth Swedish Language Technology Conference</article-title>
          , Umea˚ University,
          <fpage>17</fpage>
          -
          <lpage>18</lpage>
          November,
          <year>2016</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forsberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olsson</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          , Uppstro¨m, J.:
          <article-title>The open lexical infrastructure of Spra¨kbanken</article-title>
          .
          <source>In: Proceedings of LREC 2012</source>
          . pp.
          <fpage>3598</fpage>
          -
          <lpage>3602</lpage>
          . ELRA, Istanbul (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forsberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roxendal</surname>
          </string-name>
          , J.:
          <article-title>Korp - the corpus infrastructure of Spra¨kbanken</article-title>
          .
          <source>In: Proceedings of LREC 2012</source>
          . pp.
          <fpage>474</fpage>
          -
          <lpage>478</lpage>
          . ELRA, Istanbul (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tahmasebi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volodina</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ekman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viklund</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Megyesi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Na¨sman, J., Palme´r, A., Wire´n,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , Bjo¨rkenstam,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Grigonyte</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          , Gustafson Capkova´,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , Kosin´ski, T.:
          <article-title>Swe-clarin: Language resources and technology for digital humanities</article-title>
          .
          <source>In: Digital Humanities</source>
          <year>2016</year>
          .
          <article-title>Extended Papers of the International Symposium on Digital Humanities Va¨xjo¨, Sweden</article-title>
          . Edited by Koraljka
          <string-name>
            <surname>Golub</surname>
            ,
            <given-names>Marcelo</given-names>
          </string-name>
          <string-name>
            <surname>Milra</surname>
          </string-name>
          . Vol-2021. M. Jeusfeld c/o Redaktion Sun SITE,
          <string-name>
            <surname>Informatik</surname>
            <given-names>V</given-names>
          </string-name>
          , RWTH Aachen.,
          <string-name>
            <surname>Aachen</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Danne´lls,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Johansson</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          , Bjo¨rk, L.:
          <article-title>Evaluation and refinement of an enhanced ocr process for mass digitisation</article-title>
          .
          <source>In: Proceedings of 4th Conference of the Association Digital Humanities in the Nordic Countries (DHN)</source>
          .
          <source>CEUR</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Danne´lls,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Persson</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Supervised post OCR correction of historical Swedish texts: What role does the OCR system play</article-title>
          ?
          <source>In: Proceedings of DHN</source>
          <year>2020</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Engwall</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hedmo</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Persson</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Corpus linguistics in Sweden: Pioneers and their contexts. Kungli vitterhets historie och antikvitets akademien</article-title>
          , Stockholm (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Fridlund</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olsson</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          , Brode´n,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Borin</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Trawling for terrorists: A big data analysis of conceptual meanings and contexts in Swedish newspapers</article-title>
          ,
          <fpage>1780</fpage>
          -
          <lpage>1926</lpage>
          .
          <source>In: Proceedings of HistoInformatics 2019</source>
          . pp.
          <fpage>30</fpage>
          -
          <lpage>39</lpage>
          . CEUR-ws.org,
          <string-name>
            <surname>Aachen</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Jarlbrink</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snickars</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Cultural heritage as digital noise: Nineteenth century newspapers in the digital archive</article-title>
          .
          <source>Journal of Documentation</source>
          <volume>73</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1228</fpage>
          -
          <lpage>1243</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Karsvall</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>SDHK meets NER: Linking place names with medieval charters and historical maps</article-title>
          .
          <source>In: Proceedings of DHN 2018</source>
          . pp.
          <fpage>38</fpage>
          -
          <lpage>50</lpage>
          . CEUR-ws.org,
          <string-name>
            <surname>Aachen</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Rouces</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tahmasebi</surname>
          </string-name>
          , N.:
          <article-title>Political stance analysis using swedish parliamentary data</article-title>
          .
          <source>In: Workshop Proceedings</source>
          (Vol.
          <volume>2364</volume>
          ).
          <source>Digital Humanities in the Nordic Countries 4th Conference. CEUR Workshop Proceedings</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Rouces</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tahmasebi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eide</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          :
          <article-title>Defining a gold standard for a swedish sentiment lexicon: Towards higher-yield text mining in the digital humanities</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          vol.
          <source>2084. Proceedings of the Digital Humanities in the Nordic Countries 3rd Conference Helsinki, Finland, March 7-9</source>
          ,
          <year>2018</year>
          . Edited by Eetu Ma¨kela¨ Mikko Tolonen Jouni Tuominen. University of Helsinki, Faculty of Arts, Helsinki (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rouces</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tahmasebi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eide</surname>
            ,
            <given-names>S.R.:</given-names>
          </string-name>
          <article-title>SenSALDO: Creating a Sentiment Lexicon for Swedish</article-title>
          .
          <source>In: LREC</source>
          <year>2018</year>
          ,
          <article-title>Eleventh International Conference on Language Resources and Evaluation</article-title>
          .
          <source>ELRA</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Tahmasebi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , Brode´n,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Malm</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>A convergence of methodologies: Notes on data-intensive humanities research</article-title>
          .
          <source>In: Proceedings of DHN 2019</source>
          . pp.
          <fpage>437</fpage>
          -
          <lpage>449</lpage>
          . CEURws.org,
          <string-name>
            <surname>Aachen</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Tangherlini</surname>
            ,
            <given-names>T.R.:</given-names>
          </string-name>
          <article-title>The folklore macroscope. Challenges for a computational folkloristics</article-title>
          .
          <source>Western Folklore</source>
          <volume>72</volume>
          (
          <issue>1</issue>
          ),
          <fpage>7</fpage>
          -
          <lpage>27</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Viklund</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>How can big data help us study rhetorical history?</article-title>
          <source>In: Linko¨ping Electronic Conference Proceedings, No. 123</source>
          .
          <string-name>
            <surname>Edited by Koenraad De Smedt</surname>
          </string-name>
          .
          <source>Selected Papers from the CLARIN Annual Conference 2015. October 14-16</source>
          ,
          <year>2015</year>
          , Wroclaw, Poland. vol.
          <volume>123</volume>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>93</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. Waldispu¨hl,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Dannells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Borin</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Material philology meets digital onomastic lexicography: The NordiCon database of medieval nordic personal names in continental sources</article-title>
          .
          <source>In: Proceedings of The 12th Language Resources and Evaluation Conference</source>
          . pp.
          <fpage>860</fpage>
          -
          <lpage>867</lpage>
          . European Language Resources Association, Marseille, France (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Wissik</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Resch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Researcher Hands-On Training in the Digital Humanities: The ACDH Tool Gallery as an Austrian Case Study</article-title>
          . In: Borin,
          <string-name>
            <surname>L</surname>
          </string-name>
          . (ed.)
          <source>Selected papers from the CLARIN Annual Conference</source>
          <year>2016</year>
          ,
          <article-title>Aix-en-</article-title>
          <string-name>
            <surname>Provence</surname>
          </string-name>
          ,
          <source>CLARIN Common Language Resources and Technology Infrastructure</source>
          . pp.
          <fpage>131</fpage>
          -
          <lpage>138</lpage>
          . Linko¨ping Electronic Conference Proceedings, Linko¨ping University Electronic Press, Linko¨ping,
          <string-name>
            <surname>Sweden</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>