<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extracting literature references in German Speaking Geography - the GEOcite project</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bastian Birkeneder</string-name>
          <email>Bastian.Birkeneder@uni-passau.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Aufenvenne</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Haase</string-name>
          <email>Christian.Haase@uni-passau.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Mayr</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Malte Steinbrink</string-name>
          <email>Malte.Steinbrink@uni-passau.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Chair of Human Geography, University of Passau</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>GESIS - Leibniz Institute for the Social Sciences</institution>
          ,
          <addr-line>Cologne</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>The paper outlines the motivation and build-up of the DFG-funded GEOcite project at University of Passau. The project works on a domain-specific approach to automatically extract, segment, match and visualize literature references in the German speaking geography domain with the objective to provide a novel basis for a scientometric monitoring instrument for the community. In this paper, we describe the GEOcite corpus, its construction and elaborate on a preliminary evaluation of diferent approaches to extract and segment references from the digitized part of the corpus. We further evaluate the EXCITE segmentation model [1] on diferent datasets of German research papers. The results of our evaluation show small improvements with domain-specific and increased training data.</p>
      </abstract>
      <kwd-group>
        <kwd>Reference extraction</kwd>
        <kwd>Geography papers</kwd>
        <kwd>Network analytics</kwd>
        <kwd>Scientometric monitoring</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        (M. Steinbrink)
citation patterns, the disciplinary structure of German speaking geography is investigated. In
the first phase of the project (2014-2018), the citation relationships of all geographers who held
a professorship at a German, Austrian or Swiss university in 2012 were collected. The citation
data was retrieved from journal papers published by these actors in the decade from 2003 to
2012. The data collection was carried out partly automated using Scopus. Reference data from
geographical journals not listed in Scopus were manually extracted. The results of the first
project phase show that the discipline has clearly split into diferent subdisciplinary clusters
[
        <xref ref-type="bibr" rid="ref2 ref5 ref6">5, 6, 2</xref>
        ]. However, the subdisciplines are still more or less linked by citations. Though, the
temporal dimension of the structuring process could not be taken into account. So it is not clear
whether the current situation is the result of growing together or drifting apart. Therefore, in
the second project phase (since 2019), the data basis was comprehensively expanded in order to
enable longitudinal analyses focusing on disciplinary dynamics over time. Our aim is to include
all journal publications of geography professors from German speaking countries from 1949
until today. For this purpose the scientometric monitoring tool GEOcite was developed. In the
following, the structure and functionality of GEOcite will be explained.
      </p>
      <p>
        This paper matches with a couple of focus topics at the ULITE workshop1: GEOcite completely
builds on open source software and has the objective to produce ”Open infrastructures and
services for reference mining”; in addition, GEOcite is an application of an established software
framework for reference extraction and matching EXCITE [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] in the Geography domain. Thirdly,
GEOcite matches with the topic ”Search, exploration and mining of the reference graph” in the
way that the retrieved data will ultimately be used for network analysis aiming at a deeper
understanding of historical changes and paradigmatic shifts within geography.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. GEOcite: technical background</title>
      <p>As a software solution, GEO cite aims to locate, collect and process extensive historical and
recent citation data (from 1949 to the present). Digital and analogue archives are used for this.
Extensive digitization work is being carried out, which forms the basis for an automated citation
data extraction. For this task the EXCITE tools are used (see Figure 1). Our aim is to create a
database for the analysis of current structures of disciplinary knowledge networks and their
historical genesis and development. GEOcite enables us to create the conditions for bibliometric
network analysis to better understand the disciplinary dynamics. The data obtained are made
available to the scientific community and is thus permanently available for empirical research
and historical discipline observation.</p>
      <p>Figure 1 gives an overview of the structure and the data flows of the GEO cite tool. At the
center is the GEOcite database [M1]. This links three datasets [M1a, b, c] necessary for the
planned bibliometric network analyses.</p>
      <p>
        In GEOcite, as in the first phase of the project, the actors considered are also the geographic
professors in German speaking countries. The GEOprof -Database [actor data, M1a] contains a
list of all geography professors since 1949 as well as additional biographical attribute data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The second dataset [M1b] is a comprehensive compilation of the bibliographic information of
journal papers published by those geography professors listed in the GEOprof -Database (citing
articles). To cover both historical and current publication activity, the database feeds from two
diferent sources: In addition to texts listed in Scopus, the data is taken from analog journals
that were digitized by the Göttingen Digitization Center (GDZ). Due to copyright legislation
only the title pages and the bibliographies of a paper were scanned. The copies were provided
as single files in TIFF format and needed to be processed in several steps: First, automated
text recognition (OCR) was performed. Errors detected during text recognition were corrected
manually. The image files were then converted into PDF format. After that the documents were
merged to combine the title pages of an article and the associated bibliography into one file.
The open source programs Cermine [9] and Grobid [10] were used to extract bibliographic data
from the title pages, specifically authors and title of an article.</p>
      <p>
        The third component depicted [M1c], represents a list of all works cited in the articles. In
addition to the complete bibliographic information of the references, the dataset also contains
the link to the actor data (GEOprof database) [M1a] and the source texts [M1b] in which they
were cited. While this bibliographic information can be queried directly in Scopus, extracting
the information of the cited works from the digitized corpus is more challenging. This is
precisely the application field of the EXCITE project, which has been running since 2016 and
is funded by the DFG [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. EXCITE provides a tool for extracting literature references from
PDF files. For this purpose, the reference strings in PDF documents are automatically detected,
extracted and segmented. EXCITE was developed specifically for the extraction of citation
data from social science texts and has been trained on mainly recent German-language papers.
Relevant extracted bibliographic information is used to find matches between our actor database
and processed scientific articles. The resulting links between citing actor (matched author of
an article) and cited actor (matched author in a reference) is then used to create our citation
network.
      </p>
      <p>
        GEOcite reuses the following EXCITE tools2 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]:
1. EXannotator3 to build a dataset for Exparser model training.
      </p>
      <p>
        2. Exparser [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to process and extract the references from the PDF corpus.
      </p>
      <sec id="sec-2-1">
        <title>2.1. GEOcite Data</title>
        <p>In the following Table 1, we outline the GEOcite corpus consisting of active male and female
professors of Geography in Germany and other German speaking countries. In addition, we list
the amount of considered relevant papers in Scopus and our digitized article corpus. We divide
our data into bins of 20 years (1949–1968; 1969–1988; 1989–2008; 2009-2022).
1 Included are those professors who were actively holding a professorship in</p>
        <p>the respective time interval.
2 The Scopus corpus includes only scientific papers written by relevant actors.
3 23 articles from Scopus do not include a publication date.
4 The digitalized corpus also contains articles from authors, which are not
part of our research group.
total
1,180
29,4843
16,052</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. GEOcite tools</title>
        <p>
          As the first output of the GEO cite project, a comprehensive list of the geographic professors in
Germany, Austria and Switzerland since 1949 is available. The GEOprof dataset is available to
the research community as a static download (CSV format) at the geoscience data publisher [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
There you will also find further information about the methods used to collect the data on the
professorship. In addition, an interactive map to explore the dataset is available on the project’s
website4 (see Figure 2).
        </p>
        <p>We further created a geography-specific dataset of annotated and segmented references,
extracted from scientific articles. At the end of the project, a web-based platform for bibliometric
network analysis of the collected geogrphic citation data is planned. All datasets and tools are
or will be available for reuse5.</p>
        <p>2https://github.com/exciteproject/
3https://github.com/exciteproject/EXannotator
4https://geographische-netzwerkstatt.uni-passau.de/de/geoprof/
5https://github.com/GeoCite</p>
        <p>In the following section, we describe a preliminary evaluation and discussion of an analysis
of reference segmentation in our Geocite corpus, as well as the next steps to be taken in the
project in order to optimize the segmentation process.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Evaluation of the reference segmentation</title>
      <sec id="sec-3-1">
        <title>3.1. Set-up and Training</title>
        <p>
          In order to create a comprehensive citation network between our actors, it is essential to extract
important bibliographic data with preferably low error rates. While the default Exparser models
provide suficient results, the used toolchain allows its user to train a model on a custom dataset.
To maximize the results of the segmentation of our extracted references, we trained three
models on diferent datasets to examine the efects of more domain specific articles and more
articles in general. To increase our training data for the machine learning (ML) models used by
Exparser, we extracted references from 170 German geography research papers. These articles
were randomly chosen from our digital corpus and were published between 1952 and 2019. The
EXCITE dataset7 contains 125 German articles. We further annotated and segmented these
references according to the specified EXCITE requirements [ 11]. As test set we combined 10%
of our dataset and 10% of the EXCITE German Goldstandard. Training parameters were set
identical as reported by Hosseini et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We trained one model with our data (GEOcite model)
and the EXCITE Goldstandard (EXCITE model) respectively, as well as one model with both
7https://github.com/exciteproject/EXgoldstandard
        </p>
        <p>Label
publisher
last page
surname
article-title</p>
        <p>url
volume
source
given-names</p>
        <p>editor
first page</p>
        <p>year
identifier
issue
other
training sets combined (Combined model).</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Results</title>
        <sec id="sec-3-2-1">
          <title>The results of all three models are shown in Table 2.</title>
          <p>F1 GEOcite model</p>
          <p>F1 EXCITE model</p>
          <p>F1 Combined model</p>
          <p>Our evaluation shows that a domain-specific dataset does not necessarily improve the output
of the Exparser segmentation model. In particular we can observe that important tags for our
work, like surname, given-names, and article-title achieve lower F1-scores that the EXCITE
model. For several tags, we notice slightly improved results with the combined model. Our
results indicate that more general data and training data from our target domain can improve
the Exparser segmentation model. Considering the predominating use of English language in
the scientific community, it might be no surprise that the majority of datasets in this domain
were collected from English publications. This unfortunately limits the usage of large scale
datasets like PMC Open Access [12] or DocBank [13].</p>
          <p>One subject of our future research is the utilization of more sophisticated ML models. In recent
years a paradigm shift for ML can be observed [ 14]. Models like BERT [15] or GPT-3 [16]
trained on broad data at scale and used as foundation models show exceeding results in NLP or
Computer Vision tasks. We experiment with multilingual models (e.g. XLM-R [17]) for text
features and instance segmentation models (e.g. Mask R-CNN [18]) for structural features. The
underlying idea is to use these available large scale datasets for language independent models
and circumvent the sparsity of data in diferent languages.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Outlook</title>
      <p>With the completion of the project, we will release a dataset of our citation network, as well as
all extracted citations from our corpus. Additionally, we will provide a REST API for all members
of the scientific community to query data from diferent actors, corresponding citations, and
various other attributes.</p>
      <p>Similar to our GEOprof dataset, we will provide an interactive website where our citation
network and research results are visualized. Furthermore our software platform GEO cite will
be entirely Open Source.</p>
      <p>Initial empirical analyses based on the GEOcite corpus are also already planned. For example,
there will be further investigations on the question of the unity of geography (see above). In
addition, work is planned on paradigm genesis and evolution in German speaking geography
as well as specific bibliometric studies on the disadvantage of female geographers in the sense
of the so called Matilda efect [ 19, 20] in the course of the discipline’s history. As another
example, self-citation behavior [21] of this special community covered in the Geocite corpus
can be analysed over the covered period.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work was funded by DFG under grant 249237273, Die Säulen der Einheit und die
Brücken im Fach: Geographische Forschung zwischen Rhetorik und Praxis (GEOcite)
project, https://geographische-netzwerkstatt.uni-passau.de/geocite/.
der geographischen ProfessorInnenschaft im deutschsprachigen Raum ab 1949, 2021.
doi:1 0 . 5 8 8 0 / F I D G E O . 2 0 2 1 . 0 1 8 .
[9] D. Tkaczyk, P. Szostek, M. Fedoryszak, P. J. Dendek, Ł. Bolikowski, CERMINE: automatic
extraction of structured metadata from scientific literature, Int. J. Doc. Anal. Recognit. 18
(2015) 317–335.
[10] Grobid, https : / / github.com / kermitt2 / grobid, 2008–2022.</p>
      <p>a r X i v : 1 : d i r : d a b 8 6 b 2 9 6 e 3 c 3 2 1 6 e 2 2 4 1 9 6 8 f 0 d 6 3 b 6 8 e 8 2 0 9 d 3 c .
[11] Excite documentation, https://exparser.readthedocs.io/en/latest/ReferenceParsing/, 2019.</p>
      <p>[Online; accessed 1-May-2022].
[12] Pmc open access subset [internet]. bethesda (md): National library of medicine, https:
//www.ncbi.nlm.nih.gov/pmc/tools/openftlist/, 2003. [Online; accessed 1-May-2022].
[13] M. Li, Y. Xu, L. Cui, S. Huang, F. Wei, Z. Li, M. Zhou, Docbank: A benchmark dataset for
document layout analysis, 2020. a r X i v : 2 0 0 6 . 0 1 0 3 8 .
[14] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein,
J. Bohg, A. Bosselut, E. Brunskill, et al., On the opportunities and risks of foundation
models, arXiv preprint arXiv:2108.07258 (2021).
[15] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional
transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018).
[16] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan,
P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners, arXiv
preprint arXiv:2005.14165 (2020).
[17] A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave,
M. Ott, L. Zettlemoyer, V. Stoyanov, Unsupervised cross-lingual representation learning at
scale (2019). a r X i v : 1 9 1 1 . 0 2 1 1 6 .
[18] K. He, G. Gkioxari, P. Dollár, R. Girshick, Mask r-cnn, in: Proceedings of the IEEE
international conference on computer vision, 2017, pp. 2961–2969.
[19] M. W. Rossiter, The matthew matilda efect in science, Social Studies of Science 23 (1993)
325 – 341. doi:1 0 . 1 1 7 7 / 0 3 0 6 3 1 2 9 3 0 2 3 0 0 2 0 0 4 .
[20] P. Aufenvenne, C. Haase, F. Meixner, M. Steinbrink, Participation and communication
behaviour at academic conferences – an empirical gender study at the german congress of
geography 2019, Geoforum 126 (2021) 192–204. URL: https://www.sciencedirect.com/science/
article/pii/S0016718521001986. doi:h t t p s : / / d o i . o r g / 1 0 . 1 0 1 6 / j . g e o f o r u m . 2 0 2 1 . 0 7 . 0 0 2 .
[21] A. Kacem, J. W. Flatt, P. Mayr, Tracking self-citations in academic publishing,
Scientometrics 123 (2020) 1157–1165. doi:1 0 . 1 0 0 7 / s 1 1 1 9 2 - 0 2 0 - 0 3 4 1 3 - 9 .</p>
    </sec>
    <sec id="sec-6">
      <title>A. Online Resources</title>
      <sec id="sec-6-1">
        <title>The sources for GEOcite project will be available via</title>
        <p>• GitHub.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Boukhers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ambhore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          ,
          <article-title>An end-to-end approach for extracting and segmenting high-variance references from pdf documents</article-title>
          ,
          <source>in: Proceedings of the ACM/IEEE Joint Conference on Digital Libraries</source>
          <year>2019</year>
          ,
          <year>2019</year>
          , pp.
          <fpage>186</fpage>
          -
          <lpage>195</lpage>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>0</volume>
          <fpage>9</fpage>
          <string-name>
            <surname>/ J C D L</surname>
          </string-name>
          .
          <volume>2 0 1 9 . 0 0 0 3 5 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Aufenvenne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Steinbrink</surname>
          </string-name>
          , Brüche und Brücken:
          <article-title>Netzwerk- und zitationsanalytische Beobachtungen zur Einheit der Geographie</article-title>
          ,
          <source>Geographie und Landeskunde</source>
          <volume>88</volume>
          (
          <year>2014</year>
          )
          <fpage>257</fpage>
          -
          <lpage>292</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Kesteloot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bagnoli</surname>
          </string-name>
          ,
          <article-title>Human and physical geography: Can we learn something from the history of their relations?</article-title>
          ,
          <source>BELGEO</source>
          (
          <year>2021</year>
          ).
          <source>doi:1 0 . 4 0</source>
          <volume>0 0</volume>
          / b e l g e
          <source>o . 5 2</source>
          <volume>6 2 7 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Demeritt</surname>
          </string-name>
          ,
          <article-title>Dictionaries, disciplines and the future of geography</article-title>
          ,
          <source>Geoforum</source>
          <volume>39</volume>
          (
          <year>2008</year>
          )
          <fpage>1811</fpage>
          -
          <lpage>1813</lpage>
          .
          <source>doi:1 0 . 1 0</source>
          <volume>1 6</volume>
          / j . g
          <source>e o f o r u m . 2 0 0 8 . 0 9 . 0 0 8 .</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Steinbrink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Aufenvenne</surname>
          </string-name>
          , Integrative Geographiedidaktik?
          <source>Versuch einer Positionsbestimmung der Fachdidaktik innerhalb der deutschsprachigen Geographie</source>
          <volume>142</volume>
          /143 (
          <year>2016</year>
          ). URL: http://hw.oeaw.ac.at/?arp=
          <fpage>7887</fpage>
          -
          <lpage>0inhalt</lpage>
          /
          <fpage>gwu142</fpage>
          -143_03_
          <string-name>
            <surname>Steinbrink-Aufenvenne</surname>
          </string-name>
          .pdf.
          <source>doi:1 0 . 1 5</source>
          <volume>5 3</volume>
          / g w - u
          <source>n t e r r i c h t 1 4 2 / 1 4 3 s 5 .</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Steinbrink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Aufenvenne</surname>
          </string-name>
          ,
          <article-title>On othering and mainstreamisation of new cultural geography. some scientometric observations</article-title>
          ,
          <source>Mitteilungen der Osterreichischen Geographischen Gesellschaft</source>
          <volume>159</volume>
          (
          <year>2017</year>
          )
          <fpage>83</fpage>
          -
          <lpage>104</lpage>
          . doi:
          <article-title>1 0 . 1 5 5 3 / m o e g g 1 5 9 s 8 3</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hosseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ghavimi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Boukhers</surname>
          </string-name>
          , P. Mayr, EXCITE
          <article-title>- A toolchain to extract, match and publish open literature references</article-title>
          ,
          <source>in: Proceedings of the ACM/IEEE Joint Conference on Digital Libraries</source>
          <year>2019</year>
          , ACM,
          <year>2019</year>
          , pp.
          <fpage>432</fpage>
          -
          <lpage>433</lpage>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>0</volume>
          <fpage>9</fpage>
          <string-name>
            <surname>/ J C D L</surname>
          </string-name>
          .
          <volume>2 0 1 9 . 0 0 1 0 5 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Steinbrink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Aufenvenne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Köhler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Birkeneder</surname>
          </string-name>
          , GEOprof-Database: Datenbank
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>