<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Methods and Technologies for Integration and Processing of Geographically Distributed Quantitative Geological Information</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>© K.A. Platonov Vernadsky State Geological Museum</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moscow</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russia twinco@mail.ru</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Metadata and datasets Metadata</institution>
        </aff>
      </contrib-group>
      <fpage>250</fpage>
      <lpage>255</lpage>
      <abstract>
        <p>There are suggested methods and technologies for integration and processing of geographically distributed quantitative geological information. The methods based on DataCite approaches, DOI-system, OAI standards and protocols are applied. The developed and adapted methods and technologies were taken as principles in creation of Information system for integration and processing of quantitative geological information.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>A large geological empirical material is generated and
published annually on the Internet - quantitative
determinations of chemical and mineral composition of
rocks, ores, minerals and their aggregates in the Russian
Federation. The data are placed at the spot of their receipt
and the most intensive use in databases and information
systems, world networks for the exchange of scientific
data, as well as in scientific journals and monographs.
Geographically distributed information resources make
it difficult for geologists to obtain complete, reliable and
sufficient information for solving the scientific and
production tasks assigned to them.</p>
      <p>At the present, the organization of a single point of
access to geographically distributed quantitative
information through a single unified interface is an
urgent problem for organization of information support
and support of scientific geological research in the
Russian Federation. For a long time the integration of
quantitative information in geology was carried out at the
physical level, in particular the consolidation of data with
subject and territorial constraints.</p>
      <p>
        In 2014 the final "Declaration of Data Citation Principles"
and the mechanism for data publishing as an independent
unique product of scientific work were published [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Table 1 Sources of quantitative geological information
In 2010-2014 the developers of the DataCite project
proposed a software implementation based on the
"Declaration of Data Citation Principles". The key point
of the project is the creation of a DOI registration agency
for scientific data [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The procedure for quantitative
datasets publishing includes a DOI identifier assignment,
a unified meta description creation and registration in
the central DataCite repository
(http://www.datacite.org/). Datasets are placed in
specialized systems supporting OAI metadata exchange
protocols and sending the current information to the
central DataCite repository.
      </p>
      <p>The appearence of permanent data identification
system and a uniform global metadata scheme allows to
organize the integration of information contained in
geographically distributed sources at a logical level.</p>
      <p>
        The basic principles of information systems(IS) of
integration and data management organization are
described in the work [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Each IS should support
maintain conditions for being Findable, Accessible,
Interoperable, and Reusable data. The uniqueness of the
stored digital objects is based on the "Declaration of Data
Citation Principles".
      </p>
      <p>The main sources of tables of geological quantitative
data are listed in the table(see Table 1).</p>
      <sec id="sec-1-1">
        <title>Sources</title>
      </sec>
      <sec id="sec-1-2">
        <title>Repositories</title>
      </sec>
      <sec id="sec-1-3">
        <title>Database</title>
        <p>
          "GeoRoc" [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
FAIR Systems
"Pangaea"[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
        </p>
        <p>DataCite
Access protocols</p>
      </sec>
      <sec id="sec-1-4">
        <title>HTTP,OAI</title>
      </sec>
      <sec id="sec-1-5">
        <title>HTTP</title>
      </sec>
      <sec id="sec-1-6">
        <title>API,HTTP,OAI</title>
      </sec>
      <sec id="sec-1-7">
        <title>API,HTTP,OAI</title>
      </sec>
      <sec id="sec-1-8">
        <title>Data adding form Manually oai-pmh 2.0</title>
      </sec>
      <sec id="sec-1-9">
        <title>Manually</title>
      </sec>
      <sec id="sec-1-10">
        <title>Manually oai-pmh 2.0</title>
      </sec>
      <sec id="sec-1-11">
        <title>Metadata format</title>
        <p>oai_dc</p>
        <p>N/A
oai_dc, pan_md</p>
        <p>datacite3
oai_dc, oai_datacite</p>
      </sec>
      <sec id="sec-1-12">
        <title>Storage form Full texts of publications Table files</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2 Methods and technologies for the integration and processing of quantitative information in geology</title>
      <p>
        In the frames of the work on development of the
Information infrastructure of support and follow up of
the scientific geological research[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] , was targeted the
goal to develop methods and technologies for creation
an information system for the integration and processing
of geographically distributed quantitative geological
information.
      </p>
      <p>The following tasks were formulated:
• Development of methods and technologies for
integration quantitative data from various
technological, geographically distributed sources:
repositories, scientific journals and monographs,
information systems and databases, scientific
information exchange networks;
• Development of methods and technologies for
generation metadata of the received tables of
quantitative data in international formats;
• Implementation of a system for integration and
processing geographically distributed quantitative
geological information;
• Creation of thematic processing block of geological
information.</p>
      <sec id="sec-2-1">
        <title>2.1 The integrated method of quantitative information tables extraction from scientific publications</title>
        <p>Scientific publications have long been perceived as "Big
Data." Therefore, the tasks of this information flow
processing are solved with a high level of automation and
availability of horizontal scaling of the used algorithms.</p>
        <p>
          The structure of the scientific publication is analyzed
in [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. The process of automatic extraction of textual
information from files in PDF format is described in the
same work: article’s metadata and list of references. The
present work became a prerequisite for proposed by the
author integrated method of quantitative data tables
extraction from scientific publications.
        </p>
        <p>The PDF format, unlike other formats, does not
contain data about any structures in the document. The
file stores information about the output to the right places
of symbols, lines, curves, rectangles, raster images and
other geometric primitives. Thus, the task of finding a
table in PDF format on an article page is similar to the
tasks of recognition a table structure on a raster image.
The technology for obtaining quantitative data in the
form of a publishing electronic table consists of 5 stages:
page layout analysis, table structure detection, structure
recovery, cell function analysis, header acquisition and
table notes.</p>
        <p>The following methods were considered for the
problem solving:
• RLSA - the essence of the method is to create a binary
mask image in the vertical and horizontal direction.
Further, the masks are combined and, based on the
frequency characteristics of the black and white</p>
        <p>
          pixels, the page layout is restored. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ];
• X-Y Cut - the essence of the method is that the page
is alternately divided into blocks by a horizontal or
vertical cut. The result is a tree-like structure of the
page, where the whole page is the root, and the related
blocks of the layout are the descendants. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
• Segmentation using the maximum white rectangles
the method searches for all high maximum white
rectangles and evaluates as candidates in between the
column separators. In accordance with the found
column structure, strings, paragraphs and other
elements of the page are searched [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
• Docstrum - a method designed for segmentation of
textual information. All characters are clustered by
size (the author suggests 2 clusters). The distance
from the symbol to 4-5 of its neighbors is calculated.
A histogram of distances is constructed. The first
three peaks mean: the distance between the letters in
the word, the line spacing, the space. We combine
symbols in words, words in strings, lines in
paragraphs. The result is a layout [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
• The Voronoi diagram. The page is divided into areas,
each of which corresponds to one reference point (the
center of the symbol) and is the set of points of the
plane for which the given reference point is closer
than any other reference point [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>
          It should be noted that all algorithms are designed
for text search for the subsequent application of
algorithms for optical character recognition. In the case
of a PDF file, this is not required. Using the PDFminer
software [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], a digital version of the characters and
coordinates of the area they occupy on the page are
available. The transformation of the PDF-format into a
bitmap is done using a virtual printer.
        </p>
        <p>I proposed an integrated method for extraction
quantitative information tables from the publication text
in PDF format, which performs a complete cycle from
segmentation the page to receipt an electronic version
using a series of related methods.</p>
        <p>For each stage, the adapted versions of the above
algorithms were used. The analysis of the page layout
and table detection is carried out by a combination of
RLSA and X-Y Cut algorithms. The difference was the
use of 3 gradations (white, gray, black) when winding the
lines in a horizontal and vertical direction. A modified
version of the X-Y Cut method for each column builds a
block tree (text, table, image, graph, etc.).</p>
        <p>At the stage of restoration the structure of the table,
the segmentation method is used using the maximum
white rectangles. In my case, the maximum white and
black rectangles are searched, depending on the presence
of internal boundaries.</p>
        <p>Finally, using the Docstrum method, the values of
cells are extracted and their function is analyzed (a cell
with a value, the name of a column or rows, etc.)</p>
        <p>Regular expressions and keywords are used to find
the table header and notes: Table and Note.</p>
        <p>The test sample consisted of 136 scientific
publications. 259 objects are extracted. The error
percentage was 5%. Part of the tables cannot be
presented in a spreadsheet format due to problems with
the encoding of Russian characters in PDF files.</p>
        <p>Integration of quantitative information from
FAIRness sources and data networks is carried out in the
standard mode of metadata exchange and OAI protocol
data.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Storage of geological tables of quantitative data</title>
        <p>In the Earth sciences, in particular in geology,
quantitative tables of experimental data are accompanied
by geological, geographic, temporal and analytical
descriptions. This information can be contained in the
header of the table, a note, and in other elements of the
table, as well as in its metadata. For optimal organization
of storage and development of a topical search, the
author proposes to catalog the tables of quantitative
geological data using this information. The extracted
textual geographic characteristics are important for the
subsequent obtaining of the absolute coordinates of the
geological object. The author proposes to extract
geological descriptions using three developed thesauri:
names of geological complexes, geologic time scale,
names of mineral deposits. A textual description of the
analytical method of geological data analyzing (data on
the instrument, the method of analysis, the location and
time of the study) is in the note and is easily extracted.</p>
        <p>
          Storage of tables, retrieved from publications, is
carried out in XML-form which repeats the structure of
the table, the title and the notes to the local storage
system organized as a relational database. Each
quantitative data set is then assigned a unique identifier
(DOI) and its card is generated according to the Dublin
Core metadata format specification or its modified
version used in the DataCite project. In 2017, the
DataCite working group published a 4.1 version of the
metadata schema for publishing and citing research data
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. According to the documentation, the metadata
should contain three levels of properties: mandatory,
recommended and optional. There are enough mandatory
properties to include records in the DataCite database,
which include information about identifiers, collection
header, data about the author, publisher and the year of
publication, type or format of the resource. To improve
search and integration properties of the data sets it is
necessary to fill in the recommended additional
properties, such as subject area, keyword, members,
description, additional identifiers, geolocation (point,
rectangle or polygon), the publication or project data.
        </p>
        <p>The author developed an algorithm that compares the
information from the metadata of the publication (author,
a link to the publication), geological (subject area),
geographic (geolocation), temporary and analytical
descriptions of the table to generate metadata
automatically.</p>
        <p>
          The author proposes to store the metadata obtained
this way in the IS metadata base which will be available
to external ISs via API and OAI protocols.
The system for integration and processing
geographically distributed quantitative geological
information was developed and implemented basing on
the proposed methods. Unlike the previous
implementation of the system [
          <xref ref-type="bibr" rid="ref12 ref13">12-13</xref>
          ], in the new version
there are two blocks: the integration unit and the
geological information processing unit. The first
performs the integration of quantitative data, their
primary processing, storage, retrieval, external exchange
with other systems on standard protocols and
information management. The second block is
computationally analytical, processing and analyzing
quantitative tables of geological data. This separation is
required to solve the problems of scaling the system's
capabilities while processing and analyzing scientific
geological data, including BigData.
        </p>
        <p>The following tasks are solved:
• Organization of the mechanism for the integration of
quantitative data and their descriptions from different
types of geographically distributed sources: scientific
publications, world databases, data networks;
• Development of a system for storage geological
quantitative data and metadata;
• Creation of an algorithm for generation metadata in
the DataCite format;
• Development of algorithms for automatic extraction
of quantitative data from sources.
• Ensuring the availability of data in user-specified
formats through the search system and catalogs
• Ensure availability of data through the API and the</p>
        <p>Open Archives Initiative (OAI) protocols.
• Development of metadata management services: data
cataloging, packet loading, synchronization,
monitoring and usage statistics, etc.
• Organization of a block for the subject processing of
geological quantitative information.</p>
        <p>The general functional scheme of the System is
shown in figure (see Figure 1).</p>
        <p>The system provides data input in two modes:
automatic and manual.</p>
        <p>The user through the Search Module performs a
query to the System in one of three modes: simple,
extended and spatial. The provisioning module generates
a response and makes it possible to obtain a card of the
quantitative data table with all output information in PDF
and Excel formats (see Figure 2).</p>
        <p>The representation module manages the catalog
system, which is organized according to a thematic
principle. The field catalog contains more than 10 000
titles (the list of Cadastre Deposits of the Russian Federal
Geologic Fund "ROSGEOLFOND"). The catalog of
geological complexes of the Russian Far East includes
540 items.</p>
        <p>The user can access to the Personal account, which
stores the marked sets of quantitative data and the results
of their processing. Through the Personal account is a
manual entry of information. Through the administrator's
unit, all processes and services of the System are
managed and configured; the frequency and time of
services start-up, the status and use of the resource are
monitored.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.4 Block for processing and analyzing of geological quantitative information</title>
        <p>While carrying out geological studies, the need to
apply quantitative methods of processing, analysis and
generalization of data increases. The complexity of the
methods used is increasing from the methods of
elementary statistics to the multidimensional analysis of
data and the construction of models of natural processes.</p>
        <p>Therefore, the actual task is to provide the researcher
with tools for processing and analyzing quantitative data
on the Internet. Moreover, the user-geologist is interested
in the possibility of processing and analyzing personal
data by various methods not on his own PC, but on the
side of remote servers and, if necessary, using
supercomputers.</p>
        <p>In the system under development, a computational
and analytical unit for processing and analyzing
geological quantitative information is organized in the
form of a set of service and analytical functions with the
possibility of user access to the selection of the
processing method; chain processing, including data
loading, transformation formats, method analysis and
visualization of results; a thematic chain that includes a
sequence of analysis methods.</p>
        <p>
          While processing geological tables of quantitative
data, methods of statistical and multidimensional data
analysis, as well as the construction of graphs and maps,
are in demand. The set of possibilities of the computing
environment R is sufficient to select it as a computational
node software. In addition to the basic functions of
statistical data processing and the construction of
elementary graphs, we use the following methods:
cluster analysis, principal components, factor analysis,
discriminant analysis, canonical variables analysis,
linear and nonlinear regression analysis,
multidimensional scaling, as well as specialized sample
visualization packages, for example, "GEOmap",
developed by Jonathan Lees [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and others. The
computational capabilities of a node of a
computationalanalytical block for processing and analyzing
quantitative geological data are extensible, since the
medium R chosen by the author offers a flexible system
for adding author's algorithms.
        </p>
        <p>The "Rserve" extension allows other programs to use
the capabilities of the R language via the TCP/IP
protocol. Each connection has a separate workspace and
a directory for loading data. The computing node is
accessible by IP address (or domain name).</p>
        <p>The block is accessed through the distributed data
services management platform. The platform uses the
Web Processing Service (WPS) standard created by the
Open Geospatial Consortium to exchange information
between spatial data services, but the WPS protocol is
universal and can be used to unify access to compute
nodes with any type of data. (see Figure 3).</p>
        <p>Figure 3 Schema of the data processing node in R
language</p>
        <p>The platform provides a single access interface to all
registered processing algorithms and computing
resources and acts as an intermediary between the user
and external processing systems.</p>
        <p>A script for one’s access to the analysis methods
(principal component analysis) is shown in figure (see
Figure 4). The user downloads the data (see Figure 4(a))
and selects the analysis method (see Figure 4(b)). The
management platform via the WPS process calls remote
processing on the compute node and returns the result in
Excel format (see Figure 4(c)).</p>
        <p>The Computational and Analytical Node of the
System can be used by users or external ISs based on
their own information when accessing the Platform for
Management of Computing Services.</p>
        <p>Currently, the system is in the process of finalization
and testing. After the completion of these works it will
be available for the Internet users.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Conclusion</title>
      <p>Methods and technologies for integrating tables of
quantitative data from different technological,
geographically distributed sources are proposed.</p>
      <p>An integrated method for extracting tables of
quantitative data from the texts of geological scientific
publications is developed.</p>
      <p>An algorithm for automatically generating metadata
in the DataCite format is developed.</p>
      <p>Basing on the developed and adapted methods and
technologies, a system for the integration and processing
of geographically distributed quantitative geological
information is implemented.</p>
      <p>A computational and analytical block for processing
and analyzing geological quantitative information is
developed and implemented.</p>
      <p>Acknowledgments. This work is supervised by
Professor(Geology and Mineralogy) Naumova Vera
Viktorovna, Head of department, Vernadsky State
Geological Museum.</p>
      <p>The work is supported by RFBR, №14-07-00068.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Brase</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>DataCite - A global registration agency for research data</article-title>
          .
          <source>In: Fourth International Conference on Cooperation and Promotion of Information Resources in Science and Technology</source>
          , pp.
          <fpage>257</fpage>
          -
          <lpage>261</lpage>
          . IEEE, Beijing (
          <year>2009</year>
          ). doi:
          <volume>10</volume>
          .1109/COINFO.
          <year>2009</year>
          .66
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Breuel</surname>
            ,
            <given-names>T.M.</given-names>
          </string-name>
          :
          <article-title>Two Geometric Algorithms for Layout Analysis</article-title>
          . In: Lopresti,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kashi</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds)
          <source>Document Analysis Systems V. DAS 2002. Lecture Notes in Computer Science</source>
          , vol
          <volume>2423</volume>
          . Springer, Berlin, Heidelberg (
          <year>2002</year>
          ), doi: 10.1007/3-540-45869- 7_
          <fpage>23</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Data</given-names>
            <surname>Citation</surname>
          </string-name>
          Synthesis Group:
          <article-title>Joint Declaration of Data Citation Principles</article-title>
          . Martone, M. (ed.) San Diego CA:
          <fpage>FORCE11</fpage>
          ; (
          <year>2014</year>
          ) doi:10.25490/a97f-egyk
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>DataCite</given-names>
            <surname>Metadata</surname>
          </string-name>
          Working Group.:
          <article-title>DataCite Metadata Schema for the Publication</article-title>
          and
          <source>Citation of Research Data. Version 4</source>
          .1. DataCite e.V. (
          <year>2017</year>
          ) doi: 10.5438/0015
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Diepenbroek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grobe</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reinke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schlitzer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sieger</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wefer</surname>
          </string-name>
          , G.:
          <article-title>PANGAEA - an information system for environmental sciences</article-title>
          .
          <source>In: Computers &amp; Geosciences</source>
          , vol.
          <volume>28</volume>
          no.
          <issue>10</issue>
          ,pp.
          <fpage>1201</fpage>
          -
          <lpage>1210</lpage>
          . (
          <year>2002</year>
          ) doi:10.1016/S0098-
          <volume>3004</volume>
          (
          <issue>02</issue>
          )
          <fpage>00039</fpage>
          -
          <lpage>0</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6] GEOmap package | R Documentation. https://www.rdocumentation.org/packages/GEO map/versions/2.4-
          <fpage>4</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Kise</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sato</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Iwata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Segmentation of page images using the area voronoi diagram</article-title>
          .
          <source>In: Computer Vision and Image Understanding</source>
          , vol.
          <volume>70</volume>
          , issue 3, pp.
          <fpage>370</fpage>
          -
          <lpage>382</lpage>
          . (
          <year>1998</year>
          ) doi:10.1006/cviu.
          <year>1998</year>
          .0684
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Naumova</surname>
            ,
            <given-names>V. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goryachev</surname>
            ,
            <given-names>I. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyakov</surname>
            ,
            <given-names>S. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belousov</surname>
            ,
            <given-names>A. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Platonov</surname>
            ,
            <given-names>K. A.</given-names>
          </string-name>
          :
          <article-title>Modern technologies of development of the Information infrastructure to support the research on geology of the Russian Far East</article-title>
          .
          <source>In: Information Technology</source>
          , vol.
          <volume>21</volume>
          , no.
          <issue>7</issue>
          ,pp.
          <fpage>551</fpage>
          -
          <lpage>559</lpage>
          . (
          <year>2015</year>
          )
          <article-title>(in Russian)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Nagy</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seth</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Viswanathan</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A prototype document image analysis system for technical journals</article-title>
          .
          <source>In: Computer</source>
          , vol.
          <volume>25</volume>
          , no.
          <issue>7</issue>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>22</lpage>
          . (
          <year>1992</year>
          ).
          <source>doi: 10.1109/2</source>
          .144436
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>O</given-names>
            <surname>'Gorman</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>The document spectrum for page layout analysis</article-title>
          .
          <source>In: IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , vol.
          <volume>15</volume>
          , no.
          <issue>11</issue>
          , pp.
          <fpage>1162</fpage>
          -
          <lpage>1173</lpage>
          , (
          <year>1993</year>
          ). doi:
          <volume>10</volume>
          .1109/34.244677
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>[11] PDFminer. https://www.unixuser.org/~euske/python/pdfmi ner</mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Platonov</surname>
            ,
            <given-names>K. A.</given-names>
          </string-name>
          :
          <article-title>Methods and technologies for creation of the information processing system applied to publications on geology of the Russian Far East</article-title>
          .
          <source>In: RUSSIAN JOURNAL OF EARTH SCIENCES</source>
          , VOL.
          <volume>15</volume>
          ,
          <issue>ES4005</issue>
          ,(
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .2205/2015ES000560
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Platonov</surname>
            ,
            <given-names>K.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumova</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          :
          <article-title>Methods and technologies for geological quantitative information integration</article-title>
          .
          <source>In: Proceedings of Irkutsk State Technical University</source>
          , vol.
          <volume>21</volume>
          , no 21, pp.
          <fpage>67</fpage>
          -
          <lpage>74</lpage>
          , (
          <year>2017</year>
          ). (In Russian) doi: 10.21285/1814- 3520-2017-2-
          <fpage>67</fpage>
          -74
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Sarbas</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>The GEOROC Database as Part of a Growing Geoinformatics Network</article-title>
          . In:
          <article-title>Geoinformatics 2008 - Data to Knowledge</article-title>
          , Proceedings, pp.
          <fpage>42</fpage>
          -
          <lpage>43</lpage>
          . Potsdam, (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Tkaczyk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szostek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fedoryszak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dendek</surname>
            ,
            <given-names>P. J.</given-names>
          </string-name>
          , Bolikowski, Ł.: “CERMINE”
          <article-title>: automatic extraction of structured metadata from scientific literature</article-title>
          .
          <source>In: International Journal on Document Analysis and Recognition</source>
          , pp.
          <fpage>317</fpage>
          -
          <lpage>335</lpage>
          . (
          <year>2015</year>
          ).
          <source>doi: 10.1007/s10032-015-0249-8</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Wilkinson</surname>
            ,
            <given-names>M. D.</given-names>
          </string-name>
          et al.:
          <article-title>The FAIR Guiding Principles for scientific data management and stewardship</article-title>
          .
          <source>In: Sci. Data</source>
          <volume>3</volume>
          :
          <fpage>160018</fpage>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1038/sdata.
          <year>2016</year>
          .
          <volume>18</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>K. Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Casey</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Wahl</surname>
            ,
            <given-names>F. M.</given-names>
          </string-name>
          :
          <article-title>Document Analysis System</article-title>
          .
          <source>In: j-IBM-JRD</source>
          , vol.
          <volume>26</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>647</fpage>
          -
          <lpage>656</lpage>
          . (
          <year>1982</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>