<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards a Scalable Geoparsing Approach for the Web</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sheikh Mastura Farzana</string-name>
          <email>Sheikh.Farzana@dlr.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tobias Hecking</string-name>
          <email>Tobias.Hecking@dlr.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Geoparsing, Geographic Information Extraction, Geotagging, Geocoding</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>German Aerospace Center (DLR), Institute for Software Technology</institution>
          ,
          <addr-line>Linder Höhe, 51147 Cologne</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The ongoing surge in web data generation and storage, coupled with embedded geographic information, holds immense potential for enhancing search applications across diverse domains. However, extracting geographic information for further enhancement of web search remains inadequately explored. This paper addresses a critical gap in the realm of geographic information extraction from web data, emphasizing the absence of unified pipelines for processing such information. In response to this void, we present a pipeline specifically tailored for web data. Furthermore, our contribution extends beyond the development of the pipeline itself to include a comparative analysis of various gazetteer-based geotagging methods in terms of accuracy and scalability along with a sizable corpora of location annotated web documents.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Accurate geographic information extraction from web resources is a building block of a
geoenriched search index, which enables a wide range of location aware search applications. One
way of achieving this is to apply geoparsing on web data. Geoparsing is a standard way of
extracting locations from text (geotagging) and mapping them to their respective coordinates
(geocoding) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
websites [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        However, apart from the challenges of the geoparsing process itself [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], scaling up geoparsers
to the web comes with additional issues regarding scalability and eficient preprocessing of
      </p>
      <p>Currently there is a lack of large scale location annotated web data to test the scalability of
available geoparsers. Existing public datasets are heavily domain specific and contain particular
content types (eg: social media data), making them unsuitable for evaluating web geoparsers.
End-to-end geoparsing pipelines that can eficiently extract geo-coordinates from websites in a
single pass are still under development.</p>
      <p>This paper outlines a pipeline that can streamline processing web data for geographic
information extraction. Additionally, presents a large corpora of location annotated web contents.
CEUR
Workshop
Proceedings</p>
      <p>© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
Furthermore, the paper delves into the advantages and limitations of various gazetteer-based
geoparsers, providing insights for future advancements in the field.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related Work</title>
      <p>
        There are many of-the shelf geoparsers available with diferent capabilities. However, these
parsers were not intended for web scale geoparsing and display individual limitations. The
Edinburgh Geoparser, designed for standalone use, poses challenges in seamless integration
with diverse data processing frameworks. Its computational eficiency per document remains
ambiguous 1. CLAVIN, an open-source geotagging and geoparsing tool, excels in context-based
resolution but demands specialized expertise in Java which may be dificult to integrate in
an existing web data processing pipeline 2. Geoparsepy, utilizing OpenStreetMap, requires a
dedicated PostgreSQL database integration and returns coordinate polygons instead of individual
coordinates which is helpful when projecting locations on a map but less so when precise
coordinates are expected 3. The newly available spatial clustering-based voting approach that
combines diferent parsing approaches shows remarkable results for location disambiguation,
however has an incredibly high time consumption rate unsuitable for web scale geoparsing [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        here are diferent types of geoparsers than can be explored to find the appropriate solutions.
Hu et al. explains in his paper several types of geoparsers [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Our previous experiments have
shown that learning based geoparsers are significantly slower than simpler parsers.[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] Moreover,
since crawling and indexing web data requires complex framework of it’s own, a geoparsing
component must be capable of seamlessly integrating into these frameworks which is dificult
to achieve with existing geoparsers as they were not intended for such use cases.
      </p>
      <p>Our studies have shown that among the many approaches of geoparsing, although learning
based and hybrid solutions perform better, gazetteer based solutions are the easiest to integrate
into existing pipelines, have acceptable performance, can be scalable and extend to multiple
languages. Therefore, in this paper we have implemented several possible gazetteer based
geographic information extractors, all of them are capable of integrating into existing web data
processing frameworks. Their performance has been analyzed on diferent metrics on a large
location annotated web data corpus created by us as well as on another pre-existing location
annotated corpora.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Geographic information extraction pipeline</title>
      <p>Web content analysis serves as a valuable tool for extracting various types of information,
ranging from content category, language, various meta-information etc. The extraction of
these components relies on diferent modules, many of which are pre-existing. As a result,
any module for extracting geographic information must be capable of merging with the rest of
the analysis pipeline without added complexity. 1 shows diferent parts of our proposed web
content analysis pipeline.
1https://www.ltg.ed.ac.uk/software/geoparser/
2https://github.com/bigconnect/clavin
3https://github.com/stuartemiddleton/geoparsepy</p>
      <p>• WARC Repository: Web crawlers4 are used to crawl the web and the crawled data is
stored as WARC (Web ARChive) files, a specialized format for archiving web content 5.
Crawling can be seen as a process that accesses webpages and downloads their content.
Therefore, WARC files contain metadata such as the URL, timestamp, and content type,
along with the actual content.
• WARC Preprocessing: Each WARC file can contain information of several webpages.</p>
      <p>These files are required to be processed in order to retrieve individual data of each web
page, mainly plain text and various forms of metadata including microformats, language,
HTML headers etc. Robust libraries, such as Resiliparse,6 can be employed for eficient
processing.
• Geoparsing Component: It is the central element of the web data processing pipeline
for our case, data extracted in the previous step are fed into specialized modules for
location information extraction. This component comprises of two key processes:
– Location Extraction from Microformat: Specific locations and relevant
information related to a web article are often embedded within HTML code using
microformats. This special formatting, supported by conventions within HTML, is
easily extractable using existing libraries such as extruct7. We focus on popular
microformats like microdata 8 and json-ld 9 to extract locations without the need
for additional disambiguation as the location extracted from here are added by the
webpage publishers and definitely link to the content of the webpage.
4https://stormcrawler.net/
5https://iipc.github.io/warc-specifications/
6https://resiliparse.chatnoir.eu/en/stable/
7https://pypi.org/project/extruct/
8https://developer.mozilla.org/en-US/docs/Web/HTML/Microdata
9https://schema.org/docs/schemas.html
– Geoparsing Plain Text: The second approach involves running the extracted plain
text through first a geotagging module, in our case, a gazetteer based model. From
the geotagging step we extract possible place names from text and associate them
with both locations and geographic coordinates using a gazetteer (4.1). Section 4.3
describes at length the geoparsing models we have employed in this paper.
• Metadata Table: Extracted geographic information, including locations and coordinates,
as well as other information relevant to the particular web resource (not relevant for this
paper) is incorporated into a large table. One possible way of storing such large tables
is using Apache Parquet file format 10. These files play a crucial role in enriching web
indexes and enhancing web search results.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Geoparsing Module</title>
      <p>In order to correctly identify locations, inference of each word in a text is important, followed
by precise disambiguation of words that can mean both location or something else, eg: Paris,
can be a person’s name or a location in France. This step is followed by mapping each extracted
location to it’s geographic coordinates.</p>
      <p>For the sake of simplicity and reduction of processing time of huge web corpora, in this
work we do not apply any location disambiguation. Instead, each identified location is already
annotated with all coordinates along with country codes. Disambiguation can then be performed
on demand in a post-processing step. However, since place name disambiguation and one-to-one
mapping of places to specific coordinates is an essential step of geoparsing, henceforth we will
address our proposed geoparsers as geographic information extractors instead.</p>
      <p>For testing and refining the extraction approaches, we have created a dataset of annotated
web data. The subsequent sections provide detailed explanations of the location gazetteer,
the process of location extraction from microformats, various gazetteer-based geoparsers, and
performance evaluation.</p>
      <sec id="sec-5-1">
        <title>4.1. Location Gazetteer:</title>
        <p>A pre-existing database of locations is considered as a location gazetteer.1112 We have analyzed
GeoNames and OpenStreetMap(OSM) data, both publicly available. However, OSM database
structure is unsuitable for our task as it returns polygon for area instead of single coordinate
per location. On the other hand, GeoNames is comprehensive, easy to transform into diferent
formats and contains necessary coordinate information. Hence, we have chosen GeoNames
data for crafting our location gazetteer; we use the ‘allCountries.txt’ file provided for free public
use by GeoNames.
10https://parquet.apache.org/
11https://www.geonames.org/
12https://www.openstreetmap.org/</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Location Extraction from Microformats</title>
        <p>Microformats are formatting conventions intricately integrated into HTML code of web-pages,
commonly applied for structuring data such as contact information, events, and geographic
details, enhancing accessibility and understanding of web content. In this paper, we have only
included JsonLD microformat and we plan to integrate microdata format in future. JsonLD is
embedded as dictionary which can contain many nested dictionaries. During extraction process,
we look for entities which has the key ‘address’. If found we process it to gather all the locations
included inside the tag and search for a match in GeoNames gazetteer 4.1.</p>
      </sec>
      <sec id="sec-5-3">
        <title>4.3. Gazetteer based Place Name Extraction Models</title>
        <p>The gazetteer matching mechanism entails cross-referencing location information in text data
with an established location database, in our case the gazetteer mentioned in 4.1. While this
method is generally successful in precisely identifying location names, its eficacy highly depends
on the quality of the gazetteer and the parsing approaches used. Based on the limitations of
existing gazetteer based geoparsers (eg: Geoparsepy) we chose to instead implement diferent
types of place name extractors namely string matching, GateNLP 13, Grammar based approach
and Named Entity Recognition (NER).</p>
        <sec id="sec-5-3-1">
          <title>4.3.1. String matching</title>
          <p>The easiest solution for looking for a location in a document is to match each word in that
document against a location gazetteer. Due the simplicity of the approach, it can be incredibly
fast. In our implementation, we first match pairs of consecutive words with the gazetteer to be
able to identify locations such as San Francisco, if found we remove them from the main text
and proceed to look for a match with single worded locations. It is necessary to mention that
we implemented this solution with first four consecutive words look-up (eg: San José del Cabo)
followed by three consecutive words (eg: Andorra La Vella) and then word pair and single
worded cities. However, this solution decreases the recall and precision values with an increase
in time consumption.
4.3.2. GateNLP
GateNLP is an open-source framework for natural language processing and text mining. We
take advantage of this library to extract locations from plain text. GateNLP accepts a specially
formatted gazette and has a look-up function that can match exact parts of strings to the
provided gazette. In our case the gazette is a modified version of the location gazetteer 4.1.
GateNLP is able to retrieve locations that comprises of multiple words. Looking up each location
from gazetteer in the text without this library would be very time consuming and redundant.
13https://gatenlp.github.io/python-gatenlp/
4.3.4. NER</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Evaluation</title>
      <sec id="sec-6-1">
        <title>5.1. Dataset</title>
        <sec id="sec-6-1-1">
          <title>4.3.3. Grammar based approach</title>
          <p>We use language specific grammar rules to identify locations in a text. Often proper nouns in a
text are locations. We use NLTK14 library to construct a tagged tree from all words of a text
and retrieve the ones tagged as proper noun (NNP). It is possible to add diferent grammars
from diferent languages for similar extractions. Here is the grammar equation we have used
for English.</p>
          <p>NP : {&lt; NNP &gt; +}
The extracted proper nouns are then matched with the gazetteer to identify locations,
corresponding coordinates and country codes.</p>
          <p>Named entity recognition is a popular method of identifying the entity types of a text. We use
SpaCy15 to identify location (LOC, GPE) and match them with the gazetteer. A multilingual
SpaCy model (xx-ent-wiki-sm) is used to be able to process documents is several languages.
In order to create a large location annotated corpora we required web data. Instead of crawling
random web pages, we applied a reverse parsing technique. We first created a set of queries
using the ‘cities500’ table of the GeoNames database. Further filtering the data by removing all
locations that are outside Germany in order to keep the number of queries to a manageable
amount. For each location of the table we formulate the queries in the format city, state, country
(e.g., Bonn, North-Rhine-Westphalia, Germany), resulting in approximately 11, 500 queries. The
assumption here is that, inclusion of state and country information along with city names
will provide certain levels of disambiguation to the search engine. For example, to provide a
distinction between Frankfurt, Hesse, Germany and Frankfurt, Brandenburg, Germany etc.</p>
          <p>
            Next, we executed the queries using the ChatNoir search engine [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ], on the
CommonCrawl201716 and ClueWeb202217 data indices downloading maximum of top 10 search
results. From each resulting URI we the downloaded plain text using the Resiliparse library and
corresponding microformat data using the Extruct library.
          </p>
          <p>Afterwards, we annotated the plain texts with only the queried locations it is associated with.
This annotation process resulted in a test data corpus comprising 21, 834 web documents and a
total of 60, 070 annotated locations, average length of the texts are 26, 482 characters.</p>
          <p>
            Furthermore, we have used the ’LGL corpora’ as referenced by Gritta et al. in his paper
[
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. The dataset was created in 2010 by Lieberman et al.[
            <xref ref-type="bibr" rid="ref7">7</xref>
            ] This dataset comprises 557 news
articles, each meticulously annotated with location information, including geo-coordinates,
14https://www.nltk.org/
15https://spacy.io/
16https://commoncrawl.org/search?query=2017
17https://lemurproject.org/clueweb22.php/
by experts. Evaluating the performance of the geoparsers on this dataset alongside our larger
corpora ensures the consistency of our results across both datasets.
          </p>
          <p>The LGL corpora does not include microformat information, subsequently we did not attempt
microformat location extraction during analysis on this dataset.</p>
          <p>It is to be mentioned that both dataset contain only English text.</p>
          <p>Considering the substantial size of our web corpora, we sampled 10, 000 instances to eficiently
analyze time consumption. We had a total of 4140 JsonLD microformat in our 10k sample. 1580
instances of the web corpora returned locations extracted from microformat. For cases where
no location is found from microformat, we proceed to apply diferent geoparsing approaches,
the configurations of which are explained in Section 4.3, covering a spectrum from naive to
sophisticated implementations.</p>
        </sec>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Evaluation Metrics</title>
        <p>In our case, the most important evaluation metrics are recall rate and time consumption.
Recall shows us the percentage of locations that are successfully retrieved compared to actual
annotations. Since our location gazetteer includes all countries but the Web Corpora 10k
exclusively features German city annotations, during performance analysis, if we apply Precision
and F1-score metrics, although important indicators of performance, the results will be very
inconsistent and incomparable, hence we refrain from applying these metrics on the Web
Corpora 10k dataset.</p>
        <p>However, precision, recall, f1-score, and runtime, all 4 metrics are applied on the LGL corpora
for analyzing performance. We have computed precision, recall, and F1-score (denoted as  , 
and  ) over total number of annotated locations across all documents against total number of
retrieved locations across all documents. Runtime, measured in seconds, it is the total amount
of time that was required to process all texts from each dataset, it excludes the time needed to
load datasets or models that have an one-time requirement for the entire process. Our tasks
were executed on an Intel(R) Xeon(R) Platinum 8380 CPU @ 2.30GHz machine.</p>
      </sec>
      <sec id="sec-6-3">
        <title>5.3. Results</title>
        <p />
        <p>LGL Corpora</p>
        <p>Web Corpora 10k</p>
        <p />
        <p>Table 1 shows very interesting results for each approach. Here, sr denotes instances where
stop words were removed. It was an attempt on our part to reduce the amount of text to be
analyzed hoping it will result in reduced time consumption. However, as we can determine
from the table that inclusion of a stop word removal function increases the time consumption
substantially but fails to improve any of the metrics significantly.</p>
        <p>Our fastest model for both dataset is expectedly string matching, it has a runtime requirement
far less than any of the other models. Regardless, this naive approach of extracting locations
from text fails to have acceptable precision values resulting in low F1 Scores as well (LGL
Corpora). We can observe similar performance from GateNLP, it has incredibly high recall
compared to precision in retrieving locations. The precision score is low due to the fact that the
parser finds match in the gazetteer for a lot of regular words such as work, home, road etc.</p>
        <p>Noticing this behavior, we delved into analyzing if our location gazetteer is up to the mark.
What we found is that, there are many locations across the world that are also regular words
such as work, tuesday, move, home, road and many more. We have searched for such odd
locations on a publicly available map and was surprised to see these locations exist bringing
us to the conclusion that the location gazetteer is quite extensive. At the same time, very
naive approaches of extracting geographic locations that do not use any other information such
as sentence structure or context are bound to underperform. Additionally, reflecting on the
necessity of inference and disambiguation techniques while extracting place names from plain
text.</p>
        <p>For the case of Web Corpora 10k, the runtime and recall values of diferent approaches show
comparable values to the LGL dataset, indicating consistent behavior. We see that GateNLP has
a recall of 0.98, however looking at the precision and f1-score of this parser on the LGL dataset
we can conclude that GateNLP extracts high numbers of non-location words as locations.</p>
        <p>We had expected the grammar rule based approach to be faster than it is in reality, this
approach also sufers from lack of additional context information on extracted proper nouns,
as many proper nouns can be place names as well as other things. The NER has much more
acceptable recall values (LGL Corpora) due to the fact that SpaCy models are trained to be able
to distinguish between diferent types of entities, providing automatic disambiguation between
place names and other entities. It can be observed NER approach show acceptable recall values
with not very low precision hence The NER based approach is overall the most hopeful one
amongst the four models we have tried based on the metrics.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Discussion</title>
      <p>Importance of geographic information extraction is undeniable for information linking and
the challenge of geoparsing web content is far from resolved. In this paper we have shown
valuable insights into the scalability and capabilities of gazetteer based parsing solution for web
content with one solution having acceptable performance. However, our current approach has
limitations in-terms of place name disambiguation and one-to-one mapping of places to specific
coordinates. Regardless, more exploratory work is needed in this area in order to achieve an
well performing end-to-end web scale geoparser that has similar accuracy as existing geoparsers
for small texts as well as high robustness when integrated into web data processing pipelines.
This work has received funding from the European Union’s Horizon Europe research and
innovation programme under grant agreement No 101070014 (OpenWebSearch.EU,
https://doi.org/10.3030/101070014).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kersten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Klan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Wiegmann, Gazpne2: A general place name extractor for microblogs fusing gazetteers and pretrained transformer models</article-title>
          ,
          <source>IEEE Internet of Things Journal</source>
          <volume>9</volume>
          (
          <year>2022</year>
          )
          <fpage>16259</fpage>
          -
          <lpage>16271</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gritta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Pilehvar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Collier</surname>
          </string-name>
          , Which Melbourne?
          <article-title>Augmenting Geocoding with Maps, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics</article-title>
          , Melbourne, Australia,
          <year>2018</year>
          , pp.
          <fpage>1285</fpage>
          -
          <lpage>1296</lpage>
          . URL: http://aclweb.org/anthology/P18-1119. doi:
          <volume>10</volume>
          .18653/ v1/
          <fpage>P18</fpage>
          - 1119.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Farzana</surname>
          </string-name>
          , T. Hecking,
          <article-title>Geoparsing at web-scale-challenges and opportunities (</article-title>
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kersten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Klan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <article-title>How can voting mechanisms improve the robustness and generalizability of toponym disambiguation?</article-title>
          ,
          <source>International Journal of Applied Earth Observation and Geoinformation</source>
          <volume>117</volume>
          (
          <year>2023</year>
          )
          <fpage>103191</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          , Elastic ChatNoir:
          <article-title>Search Engine for the ClueWeb and the Common Crawl</article-title>
          , in: L.
          <string-name>
            <surname>Azzopardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Pasi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          Piwowarski (Eds.),
          <source>Advances in Information Retrieval. 40th European Conference on IR Research (ECIR</source>
          <year>2018</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gritta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Pilehvar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Collier</surname>
          </string-name>
          ,
          <article-title>A pragmatic guide to geoparsing evaluation, arXiv preprint</article-title>
          arXiv:
          <year>1810</year>
          .
          <volume>12368</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Lieberman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Samet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sankaranarayanan</surname>
          </string-name>
          ,
          <article-title>Geotagging with local lexicons to build indexes for textually-specified spatial data</article-title>
          ,
          <source>in: 2010 IEEE 26th international conference on data engineering (ICDE</source>
          <year>2010</year>
          ), IEEE,
          <year>2010</year>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>