<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Open Science System for Text Mining</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gianpaolo Coro</string-name>
          <email>coro@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giancarlo Panichi</string-name>
          <email>panichi@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pasquale Pagano</string-name>
          <email>pagano@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ISTI-CNR</institution>
          ,
          <addr-line>via Moruzzi 1 Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Text mining (TM) techniques can extract high-quality information from big data through complex system architectures. However, these techniques are usually difficult to discover, install, and combine. Further, modern approaches to Science (e.g. Open Science) introduce new requirements to guarantee reproducibility, repeatability, and re-usability of methods and results as well as their longevity and sustainability. In this paper, we present a distributed system (NLPHub) that publishes and combines several state-of-theart text mining services for named entities, events, and keywords recognition. NLPHub makes the integrated methods compliant with Open Science requirements and manages heterogeneous access policies to the methods. In the paper, we assess the benefits and the performance of NLPHub on the I-CAB corpus1.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Today, text mining operates within the
challenges introduced by big data and new Science
paradigms, which impose to manage large
volumes, high production rate, heterogeneous
complexity, and unreliable content, while ensuring
data and methods longevity through re-use in
complex models and processes chains. Among the new
paradigms, Open Science (OS) focusses on the
implementation in computer systems of the three
"R"s of the scientific method: Reproducibility,
Repeatability, and Re-usability
        <xref ref-type="bibr" rid="ref17 ref24">(Hey et al., 2009; EU
Commission, 2016)</xref>
        . The systems envisaged by
OS, are based on Web services networks that
support big data processing and the open publication
1Copyright © 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
of results. Although text mining techniques
exist that can tackle big data experiments
        <xref ref-type="bibr" rid="ref1 ref2 ref20">(Gandomi
and Haider, 2015; Amado et al., 2018)</xref>
        , few
examples that incorporate OS concepts can be found
        <xref ref-type="bibr" rid="ref28">(Linthicum, 2017)</xref>
        . For example, common text
mining "cloud" services do not allow easy
repeatability of the experiments by different users and are
usually domain-specific and thus poorly re-usable
        <xref ref-type="bibr" rid="ref1 ref2 ref4 ref8">(Bontcheva and Derczynski, 2016; Adedugbe et
al., 2018)</xref>
        . Available multi-domain systems do
not use communication standards
        <xref ref-type="bibr" rid="ref12 ref4 ref43 ref44 ref6 ref8">(Bontcheva and
Derczynski, 2016; Wei et al., 2016)</xref>
        , and the few
OS-oriented initiatives that use text mining focus
specifically on documents preservation and
cataloguing
        <xref ref-type="bibr" rid="ref32 ref33">(OpenMinTeD, 2019; OpenAire, 2019)</xref>
        .
      </p>
      <sec id="sec-1-1">
        <title>In this paper, we present a multi-domain text</title>
        <p>
          mining system (NLPHub) that is compliant with
OS and combines multiple and heterogeneous
processes. NLPHub is based on an e-Infrastructure
(e-I), i.e. a network of hardware and software
resources that allow remote users and services
to collaborate while supporting data-intensive
Science through cloud computing
          <xref ref-type="bibr" rid="ref18 ref26 ref3 ref35 ref42">(Pollock and
Williams, 2010; Andronico et al., 2011)</xref>
          .
Currently, NLPHub integrates 30 state-of-the-art text
mining services and methods to recognize
fragments of a text (annotations) associated with
named abstract or physical objects (named
entities), spatiotemporal events, and keywords. These
integrated processes cover overall 5 languages
(English, Italian, German, French, and Spanish),
requested by the European projects this software
is involved in (i.e.
          <xref ref-type="bibr" rid="ref34 ref39 ref5">(Parthenos, 2019; SoBigData,
2019; Ariadne, 2019)</xref>
          ). These processes come
from different providers that have different
access policies, and the e-I is used both to
manage this heterogeneity and to possibly speed up
the processing through cloud computing.
NLPHub uses the Web Processing Service standard
(WPS,
          <xref ref-type="bibr" rid="ref37">(Schut and Whiteside, 2007)</xref>
          ) to describe
all integrated processes, and the Prov-O XML
ontological standard
          <xref ref-type="bibr" rid="ref27 ref9">(Lebo et al., 2013)</xref>
          to track
the complete set of input, output, and parameters
used for the computations (provenance). Overall,
these features enable OS-compliance and we show
that the orchestration mechanism implemented by
NLPHub adds effectiveness and efficiency to the
connected methods. The name "NLPHub" refers
to the forthcoming extensions of this platform to
other text mining methods (e.g. sentiment
analysis and opinion mining), and natural language
processing tasks (e.g. text-to-speech and speech
processing).
2
2.1
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Methods and tools</title>
      <sec id="sec-2-1">
        <title>E-Infrastructure and Cloud Computing</title>
      </sec>
      <sec id="sec-2-2">
        <title>Platform</title>
        <sec id="sec-2-2-1">
          <title>NLPHub uses the open-source D4Science e</title>
          <p>
            I
            <xref ref-type="bibr" rid="ref27 ref31 ref7 ref9">(Candela et al., 2013; Assante et al., 2019)</xref>
            ,
which currently supports applications in many
domains through the integration of a distributed
storage system, a cloud computing platform, online
collaborative tools, and catalogues of metadata
and geospatial data. D4Science supports the
creation of Virtual Research Environments (VREs)
            <xref ref-type="bibr" rid="ref12 ref43 ref44 ref6">(Assante et al., 2016)</xref>
            , i.e. Web-based
environments fostering collaboration and data sharing
between users and managing heterogeneous data and
services access policies. D4Science grants each
user with access to a private online file system
(the Workspace) that uses a high-availability
distributed storage system behind the scenes, and
enables folders creation and sharing between VRE
users. Through VREs and accounting and security
services, D4Science is able to manage
heterogeneous access policies by granting free access to
open services in public VREs, and
controlled/private access to non-open services in private or
moderated VREs. D4Science includes a cloud
computing platform named DataMiner
            <xref ref-type="bibr" rid="ref11 ref13">(Coro et al., 2015;
Coro et al., 2017)</xref>
            that currently hosts ∼400
processes and makes all integrated processes
available under the WPS standard (Figure 1). WPS
is supported by third-party software and allows
standardising a process’ input, its
parameterisation and output. DataMiner executes the processes
in a cloud computing cluster of 15 machines with
Ubuntu 16.04.4 LTS x86 64 operating system, 16
virtual cores, 32 GB of RAM and 100 GB of disk
space. These machines are hosted by the National
Research Council of Italy and the Italian Network
of the University and Research (GARR). Each
process can parallelise an execution either across
the machines (using a Map-Reduce approach) or
on the cores of one single machine
            <xref ref-type="bibr" rid="ref13">(Coro et al.,
2017)</xref>
            . After each computation, DataMiner saves
on the user’s Workspace- all the information about
the input and output data, and the experiment’s
parameters (computational provenance) using the
Prov-O XML standard. In each D4Science VRE,
DataMiner offers an online tool to integrate
algorithms, which supports many programming
languages
            <xref ref-type="bibr" rid="ref12 ref43 ref44 ref6">(Coro et al., 2016)</xref>
            . All these features
make D4Science useful to develop OS-compliant
applications, because WPS and provenance
tracking allow repeating and reproducing a
computation executed by another user. Also, the possibility
to provide a process in multiple VREs focussing
on different domains fosters its re-usability
            <xref ref-type="bibr" rid="ref13">(Coro
et al., 2017)</xref>
            . In this paper, we will use the
term "algorithm" to indicate processes running on
DataMiner, and "method" to indicate the original
processes or services integrated with DataMiner.
2.2
          </p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Annotations</title>
        <p>NLPHub integrates a number of named entities
recognizers (NERs) but also information
extraction processes that recognize events, keywords,
tokens, and sentences. Overall, we will use
the term "annotation" to indicate all the
information that NLPHub can extract from a text.
The complete list of supported annotations,
languages, and processes is reported in the
supplementary material, together with the list of all
mentioned Web services’ endpoints. The
ontological classes used for NERs annotations come from
the Standford CoreNLP software. Included
nonstandard annotations are "Misc" (miscellaneous
concepts that cannot be associated with none of
the other classes, e.g. "Bachelor of Science"),
"Event" (nouns, verbs, or phrases referring to a
phenomenon occurring at a certain time and/or
space), and "Keyword" (a word or a phrase that
is of great importance to understand the text
content).</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.3 Integrated Text Mining Methods</title>
        <p>NLPHub uses a common JSON format to
represent the annotations of every integrated method.
This format describes the input text, the NER
processes, and the annotations for each NER:
1 "text": "input text",
2 "N ER1": {
3 "annotations":{
4 "annotation1":[
5 {"indices": [i1,i2]},
6 {"indices": [i3,i4]},
7 ...,</p>
        <sec id="sec-2-4-1">
          <title>We integrated services and methods with</title>
          <p>DataMiner through "wrapping algorithms" that
transformed the original outputs into this format.</p>
          <p>We implemented a general workflow in each
algorithm to execute the corresponding integrated
method, which adopts the following steps: (i)
receive an input text file and a list of entities to
recognize (among those supported by the language),
(ii) pre-process the text by deleting useless
characters, (iii) encode the text with UTF-8
encoding, (iv) send the text via HTTP-Post to the
corresponding service or execute the method on the
local machine directly, if possible, and (v) return the
annotation as an NLPHub-compliant JSON
document. In the following, we list all the methods
currently integrated with NLPHub with reference
to Figure 1 for an architectural view.</p>
          <p>
            CoreNLP. The Stanford CoreNLP software
            <xref ref-type="bibr" rid="ref16 ref30">(Manning et al., 2014)</xref>
            is an open-source text
processing toolkit that supports several languages
            <xref ref-type="bibr" rid="ref41">(Stanford University, 2019)</xref>
            . NLPHub integrates
CoreNLP as a service instance running within
D4Science with English, German, French, and
Spanish language packages enabled. Also, the
Tint (The Italian NLP Tool) extension for Italian
            <xref ref-type="bibr" rid="ref4 ref8">(Aprosio and Moretti, 2016)</xref>
            was installed as a
separate service. Overall, two distinct replicated and
balanced virtual machines host these services on
machines with 10 GB of RAM and 6 cores.
          </p>
          <p>
            GATE Cloud. GATE Cloud is a cloud
service that offers on-payment text analysis methods
as-a-service
            <xref ref-type="bibr" rid="ref22 ref23 ref26 ref3 ref35 ref42">(GATE Cloud, 2019a; Tablan et al.,
2011)</xref>
            . NLPHub integrates the GATE Cloud
ANNIE NER for English, German, and French within
a controlled VRE that accounts for users’ requests
load. This VRE ensures a fair usage of the
services, whose access has been freely granted to
D4Science in exchange for enabling OS-oriented
features
            <xref ref-type="bibr" rid="ref38">(SoBigData European Project, 2016)</xref>
            .
          </p>
          <p>
            OpenNLP. The Apache OpenNLP library is an
open source text processing toolkit mostly based
on machine learning models
            <xref ref-type="bibr" rid="ref26 ref3 ref35 ref42">(Kottmann et al.,
2011)</xref>
            . An OpenNLP-based English NER is
available as-a-service on GATE Cloud
            <xref ref-type="bibr" rid="ref22 ref23">(GATE Cloud,
2019b)</xref>
            and is included among the free-to-use
services granted to D4Science.
          </p>
          <p>
            ItaliaNLP. ItaliaNLP is a free-to-use
service - developed by the "Istituto di Linguistica
Computazionale" (ILC-CNR) - hosting a NER
method for Italian that combines rule-based and
machine learning algorithms
            <xref ref-type="bibr" rid="ref16 ref25 ref30">(ILC-CNR, 2019;
Dell’Orletta et al., 2014)</xref>
            .
          </p>
          <p>
            NewsReader. NewsReader is an advanced
events recognizer for 4 languages, developed by
the NewsReader European project
            <xref ref-type="bibr" rid="ref12 ref43 ref44 ref6">(Vossen et al.,
2016)</xref>
            . NewsReader is a formal inferencing
system that identifies events by detecting their
participants and time-space constraints. Two balanced
virtual machines were installed in D4Science for
the English and Italian NewsReader versions.
          </p>
          <p>
            TagMe. TagMe is a service for identifying short
phrases (anchors) in a text that can be linked to
pertinent Wikipedia pages
            <xref ref-type="bibr" rid="ref18">(Ferragina and Scaiella,
2010)</xref>
            . TagMe supports 3 languages (English,
Italian, and German) and D4Science already hosts its
official instances. Since anchors are sequences of
words having a recognized meaning within their
context, NLPHub interprets them as keywords that
can help contextualising and understanding the
text.
          </p>
          <p>
            Keywords NER. Keywords NER is an
opensource statistical method that produces tags clouds
of verbs and nouns
            <xref ref-type="bibr" rid="ref14 ref15">(Coro, 2019a)</xref>
            , which was also
used by the H-Care award-winning human
digital assistant
            <xref ref-type="bibr" rid="ref40">(SpeechTEK 2010, 2019)</xref>
            . Tag clouds
are extracted through a statistical analysis of
partof-speech (POS) tags (extracted with TreeTagger, and the alignment algorithm manages all cases
            <xref ref-type="bibr" rid="ref36">(Schmid, 1995)</xref>
            ) and the method can be applied to through algebraic evaluations, as reported in the
all the 23 TreeTagger supported languages. Key- following pseudo-code:
words NER is executed directly on the DataMiner
machines, and the nouns tags are interpreted as 1 AMERGE Algorithm
keywords for the NLPHub scopes, because - by 2
construction - their sequence is useful to under- 3 For each annotation E:
stand the topics treated by a text. 4 Collect all annotations detected
by the algorithms (intervals
          </p>
          <p>
            Language Identifier. NLPHub also provides with text start and end
a language identification process
            <xref ref-type="bibr" rid="ref14 ref15">(Coro, 2019b)</xref>
            , positions);
should language information not be specified as 5 Sort the intervals by their
input. This process was developed in order to be start position;
fast, easily, and quickly extendible to new lan- 6 For each segment si:
guages. The algorithm is based on an empirical 7 If sj is properly included in
behaviour of TreeTagger (common to many POS si, process the next sj ;
taggers): When TreeTagger is initialised on a cer- 8 If si does not intersect sj ,
tain language, but it processes a text written in brake the loop;
another language, it tends to detect many nouns 9 If si intersects sj , create a
and unstemmed words than verbs and other lexi- new segment sui as the union
cal categories. Thus, the detected language is the of the two segments →
one having the most balanced ratio of recognized substitute sui to si and
and stemmed words with respect to other lexical restart the loop on sj ;
categories. This algorithm is applicable to many 10 Save si in the overall list of
languages supported by TreeTagger and can run merged intervals S;
on the DataMiner machines directly. An estimated 11 Associate S to E;
accuracy of 95% on 100 sample text files covering 12 Return all (E; S) pairs sets.
the 5 NLPHub languages was convincing to use
this algorithm as an auxiliary tool for the NLPHub
users.
2.4
          </p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>NLPHub</title>
        <p>On top of the methods and services described so
far, we implemented an alignment-merging
algorithm (AMERGE) that orchestrates the
computations and assembles their outputs. AMERGE
receives a user-provided input text, along with the
indication of the text language (optionally), and a
set of annotations to be extracted (selected among
those supported for that language). Then, it
concurrently invokes - via WPS - the text
processing algorithms that support the input request, and
eventually collects the JSON documents coming
from them. Finally, it aligns and merges the
information to produce one overall sequence
represented in JSON format. The issue of merging
the heterogeneous connected services’ outputs is
solved through the use of the DataMiner wrapping
algorithms. Another solved issue is the merge of
the different intervals identified by several
algorithms focusing on the same entities. These
intervals may either overlap or be mutually inclusive,</p>
        <sec id="sec-2-5-1">
          <title>Since the AMERGE algorithm is a DataMiner</title>
          <p>
            algorithm, it is published as-a-service with a
RESTful WPS interface. It represents one single
access point to the services integrated with
NLPHub. In order to invoke this service, a client
should specify an authorization code in the HTTP
request that identifies both the invoking user and
the VRE
            <xref ref-type="bibr" rid="ref10">(CNR, 2016)</xref>
            . The available annotations
and methods depend on the VRE. An additional
service (NLPHub-Info) allows retrieving the list
of supported entities for a VRE, given a user’s
authorization code. NLPHub is also endowed with
a free-to-use Web interface (nlp.d4science.
org/hub/), based on a public VRE, operating on
top of the AMERGE process, which allows
interacting with the system and retrieving the
annotations in a graphical format.
3
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        We assessed the NLPHub performance by
using the I-CAB corpus as a reference
        <xref ref-type="bibr" rid="ref29">(Magnini et
al., 2006)</xref>
        , which contains annotations of the
following named-entities categories from 527
Italian newspapers: Person, Location, Organization,
      </p>
      <p>
        Geopolitical entity. NLPHub was executed to
annotate these same entities plus Keywords
(Table 1). The involved algorithms were
CoreNLPTint, ItaliaNLP, Keywords NER, and TagMe.
According to the F-measure, CoreNLP-Tint was
the best at recognizing Persons and
Organizations, whereas ItaliaNLP - the only one supporting
Geopolitical entities - had the highest performance
on Locations and a moderately-high performance
on Geopolitical entities. Overall, the connected
methods showed high performance on specific
entities, but there was not one method
outperforming the others on all entities. AMERGE had lower
but good F-measure and a generally high recall in
all cases, which indicates that the connected
algorithms include complementary and valuable
intervals. The AMERGE-Keywords algorithm had
a generally high recall (especially on
Geopolitical entities), which means that the extracted
keywords include also words from the annotated
entities. The associated F-measures indicate that there
is overlap with several entities. In turn, this
indicates that AMERGE-Keywords could be a
valuable source of information in the case of
uncertainty about the entities that can be extracted from
a text. As a further evaluation, we used Cohen’s
Kappa
        <xref ref-type="bibr" rid="ref21">(Cohen, 1960)</xref>
        to explore the agreement
between the algorithms and the I-CAB
annotations. This measure required estimating the
overall number of classifiable tokens, thus it is more
realistic to refer to Fleiss’ Kappa macro
classifications rather than to the exact values
        <xref ref-type="bibr" rid="ref19">(Fleiss, 1971)</xref>
        .
According to Fleiss’ labels, all NERs generally
have good agreement with I-CAB except for
Locations, which are often reported as Geopolitical
entities in I-CAB. This evaluation also highlights
that AMERGE has good general agreement with
manual annotations, and thus can be a valid choice
when there is no prior knowledge about the
algorithm to use for extracting a certain entity.
We have described NLPHub, a distributed
system connecting and combining 30 text processing
methods for 5 languages that adds Open
Scienceoriented features to these methods. The
advantages of using NLPHub are several, starting from
the fact that it provides one single access
endpoint to several methods and spares installation
and configuration time. Further, it proposes the
AMERGE process as a valid option when the best
performing algorithm for a certain entity
extraction is not known a priori. Also, the
AMERGEKeywords annotations can be used when the
entities to extract are not known. Indeed, these
features would require more investigation, especially
through multiple-language experiments, in order
to define their full potential and limitations.
Finally, NLPHub adds to the original methods
features like WPS and Web interfaces, provenance
management, results sharing, and access/usage
policies control, which make the methods more
compliant the with Open Science requirements.
      </p>
      <p>
        The potential users of NLPHub are scholars
who want to use NERs but also want to avoid
software and hardware-related issues, or automatic
agents that need to automatically extract and
reuse knowledge from large quantities of texts. For
example, NLPHub can be used in automatic
ontology population and - since it also supports Events
extraction - automatic narratives generation
        <xref ref-type="bibr" rid="ref26 ref3 ref31 ref35 ref42 ref7">(Petasis et al., 2011; Metilli et al., 2019)</xref>
        . Future
extensions of NLPHub will involve other text mining
methods (e.g. sentiment analysis, opinion mining,
and morphological parsing), and additional NLP
tasks like text-to-speech and speech processing
asa-service.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Supplementary Material</title>
      <p>Supplementary material is available on D4Science
at this permanent hyper-link.</p>
      <p>The
Arihttps:</p>
      <p>2016.
https:</p>
      <p>The
http:
Knowledge-driven multimedia information
extraction and ontology evolution, pages 134–166.</p>
      <p>Springer-Verlag.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Adedugbe et al.2018]
          <string-name>
            <given-names>Oluwasegun</given-names>
            <surname>Adedugbe</surname>
          </string-name>
          , Elhadj Benkhelifa, and
          <string-name>
            <given-names>Russell</given-names>
            <surname>Campion</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A clouddriven framework for a holistic approach to semantic annotation</article-title>
          .
          <source>In 2018 Fifth International Conference on Social Networks Analysis, Management and Security (SNAMS)</source>
          , pages
          <fpage>128</fpage>
          -
          <lpage>134</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Amado et al.2018]
          <string-name>
            <given-names>Alexandra</given-names>
            <surname>Amado</surname>
          </string-name>
          , Paulo Cortez, Paulo Rita, and
          <string-name>
            <given-names>Sérgio</given-names>
            <surname>Moro</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Research trends on big data in marketing: A text mining and topic modeling based literature analysis</article-title>
          .
          <source>European Research on Management and Business Economics</source>
          ,
          <volume>24</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Andronico et al.2011]
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Andronico</surname>
          </string-name>
          , Valeria Ardizzone, Roberto Barbera, Bruce Becker, Riccardo Bruno, Antonio Calanducci, Diego Carvalho, Leandro Ciuffo, Marco Fargetta,
          <string-name>
            <given-names>Emidio</given-names>
            <surname>Giorgio</surname>
          </string-name>
          , et al.
          <year>2011</year>
          .
          <article-title>e-infrastructures for e-science: a global view</article-title>
          .
          <source>Journal of Grid Computing</source>
          ,
          <volume>9</volume>
          (
          <issue>2</issue>
          ):
          <fpage>155</fpage>
          -
          <lpage>184</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[Aprosio and Moretti2016] Alessio Palmero Aprosio and Giovanni Moretti</source>
          .
          <year>2016</year>
          .
          <article-title>Italy goes to stanford: a collection of corenlp modules for italian</article-title>
          .
          <source>arXiv preprint arXiv:1609</source>
          .
          <fpage>06204</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Ariadne2019] Ariadne</source>
          .
          <year>2019</year>
          . adnePlus European Project. //ariadne-infrastructure.eu/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Assante et al.2016]
          <string-name>
            <given-names>Massimiliano</given-names>
            <surname>Assante</surname>
          </string-name>
          , Leonardo Candela, Donatella Castelli, Gianpaolo Coro, Lucio Lelii, and
          <string-name>
            <given-names>Pasquale</given-names>
            <surname>Pagano</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Virtual research environments as-a-service by gcube</article-title>
          .
          <source>PeerJ Preprints</source>
          ,
          <volume>4</volume>
          :
          <fpage>e2511v1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Assante et al.2019]
          <string-name>
            <given-names>Massimiliano</given-names>
            <surname>Assante</surname>
          </string-name>
          , Leonardo Candela, Donatella Castelli, Roberto Cirillo, Gianpaolo Coro, Luca Frosini, Lucio Lelii, Francesco Mangiacrapa, Valentina Marioli,
          <string-name>
            <given-names>Pasquale</given-names>
            <surname>Pagano</surname>
          </string-name>
          , et al.
          <year>2019</year>
          .
          <article-title>The gcube system: Delivering virtual research environments as-a-service</article-title>
          .
          <source>Future Generation Computer Systems</source>
          ,
          <volume>95</volume>
          :
          <fpage>445</fpage>
          -
          <lpage>453</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Bontcheva and Derczynski2016] Kalina Bontcheva and Leon Derczynski</source>
          .
          <year>2016</year>
          .
          <article-title>Extracting information from social media with gate</article-title>
          .
          <source>In Working with Text</source>
          , pages
          <fpage>133</fpage>
          -
          <lpage>158</lpage>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Candela et al.2013]
          <string-name>
            <given-names>Leonardo</given-names>
            <surname>Candela</surname>
          </string-name>
          , Donatella Castelli, Gianpaolo Coro, Pasquale Pagano, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Sinibaldi</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Species distribution modeling in the cloud</article-title>
          .
          <source>Concurrency and Computation: Practice and Experience.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>[CNR2016] CNR</source>
          .
          <year>2016</year>
          .
          <article-title>gcube wps thin clients</article-title>
          . https://wiki.gcube-system.org/ gcube/How_to_
          <article-title>Interact_with_the_ DataMiner_by_client.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Coro et al.2015]
          <string-name>
            <given-names>Gianpaolo</given-names>
            <surname>Coro</surname>
          </string-name>
          , Leonardo Candela, Pasquale Pagano, Angela Italiano, and
          <string-name>
            <given-names>Loredana</given-names>
            <surname>Liccardo</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Parallelizing the execution of native data mining algorithms for computational biology</article-title>
          .
          <source>Concurrency and Computation: Practice and Experience</source>
          ,
          <volume>27</volume>
          (
          <issue>17</issue>
          ):
          <fpage>4630</fpage>
          -
          <lpage>4644</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Coro et al.2016]
          <string-name>
            <given-names>Gianpaolo</given-names>
            <surname>Coro</surname>
          </string-name>
          , Giancarlo Panichi, and
          <string-name>
            <given-names>Pasquale</given-names>
            <surname>Pagano</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A web application to publish r scripts as-a-service on a cloud computing platform</article-title>
          .
          <source>Bollettino di Geofisica Teorica ed Applicata</source>
          ,
          <volume>57</volume>
          :
          <fpage>51</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Coro et al.2017]
          <string-name>
            <given-names>Gianpaolo</given-names>
            <surname>Coro</surname>
          </string-name>
          , Giancarlo Panichi, Paolo Scarponi, and
          <string-name>
            <given-names>Pasquale</given-names>
            <surname>Pagano</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Cloud computing in a distributed e-infrastructure using the web processing service standard</article-title>
          .
          <source>Concurrency and Computation: Practice and Experience</source>
          ,
          <volume>29</volume>
          (
          <issue>18</issue>
          ):
          <fpage>e4219</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Coro2019a]
          <string-name>
            <given-names>Gianpaolo</given-names>
            <surname>Coro</surname>
          </string-name>
          .
          <year>2019a</year>
          .
          <article-title>The Keywords Tag Cloud Algorithm</article-title>
          . https: //svn.research-infrastructures. eu/public/d4science/gcube/ trunk/data-analysis/ LatentSemanticAnalysis/.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Coro2019b]
          <string-name>
            <given-names>Gianpaolo</given-names>
            <surname>Coro</surname>
          </string-name>
          .
          <year>2019b</year>
          .
          <article-title>The Language Identifier Algorithm. hyper-link.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>[Dell'Orletta</surname>
            et al.2014]
            <given-names>Felice</given-names>
          </string-name>
          <string-name>
            <surname>Dell'Orletta</surname>
            , Giulia Venturi, Andrea Cimino, and
            <given-names>Simonetta</given-names>
          </string-name>
          <string-name>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>T2kˆ 2: a system for automatically extracting and organizing knowledge from texts</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources</source>
          and
          <article-title>Evaluation (LREC-</article-title>
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [EU Commission2016]
          <article-title>EU Commission</article-title>
          .
          <article-title>Open science (open access)</article-title>
          . //ec.europa.eu/programmes/ horizon2020/en/h2020-section/ open-science
          <string-name>
            <surname>-</surname>
          </string-name>
          open-access.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>[Ferragina and Scaiella2010] Paolo Ferragina and Ugo Scaiella</source>
          .
          <year>2010</year>
          .
          <article-title>Tagme: on-the-fly annotation of short text fragments (by wikipedia entities)</article-title>
          .
          <source>In Proceedings of the 19th ACM international conference on Information and knowledge management</source>
          , pages
          <fpage>1625</fpage>
          -
          <lpage>1628</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>[Fleiss1971] Joseph L Fleiss</surname>
          </string-name>
          .
          <year>1971</year>
          .
          <article-title>Measuring nominal scale agreement among many raters</article-title>
          .
          <source>Psychological bulletin</source>
          ,
          <volume>76</volume>
          (
          <issue>5</issue>
          ):
          <fpage>378</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[Gandomi and Haider2015] Amir Gandomi and Murtaza Haider</source>
          .
          <year>2015</year>
          .
          <article-title>Beyond the hype: Big data concepts, methods, and analytics</article-title>
          .
          <source>International Journal of Information Management</source>
          ,
          <volume>35</volume>
          (
          <issue>2</issue>
          ):
          <fpage>137</fpage>
          -
          <lpage>144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Cohen1960]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>1960</year>
          .
          <article-title>A coefficient of agreement for nominal scales</article-title>
          .
          <source>Educational and psychological measurement</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <fpage>37</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>[GATE Cloud2019a] GATE Cloud</source>
          .
          <year>2019a</year>
          . GATE Cloud:
          <article-title>Text Analytics in the Cloud</article-title>
          . https:// cloud.gate.ac.uk/.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [GATE Cloud2019b]
          <article-title>GATE Cloud. 2019b. OpenNLP English Pipeline</article-title>
          . https://cloud.gate. ac.uk/shopfront/displayItem/ opennlp-english-pipeline.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Hey et al.2009]
          <string-name>
            <given-names>Tony</given-names>
            <surname>Hey</surname>
          </string-name>
          ,
          <string-name>
            <surname>Stewart</surname>
            <given-names>Tansley</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kristin M Tolle</surname>
          </string-name>
          , et al.
          <year>2009</year>
          .
          <article-title>The fourth paradigm: dataintensive scientific discovery</article-title>
          , volume
          <volume>1</volume>
          . Microsoft research Redmond, WA.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <article-title>[ILC-CNR2019] ILC-CNR</article-title>
          .
          <year>2019</year>
          .
          <article-title>The ItaliaNLP REST Service</article-title>
          . http://api.italianlp.it/ docs/.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [Kottmann et al.
          <year>2011</year>
          ]
          <string-name>
            <given-names>J</given-names>
            <surname>Kottmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B</given-names>
            <surname>Margulies</surname>
          </string-name>
          , G Ingersoll,
          <string-name>
            <given-names>I</given-names>
            <surname>Drost</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Kosin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Baldridge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Goetz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Morton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Autayeu</surname>
          </string-name>
          , et al.
          <year>2011</year>
          .
          <article-title>Apache OpenNLP. www</article-title>
          .opennlp.apache.org.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [Lebo et al.2013]
          <string-name>
            <given-names>Timothy</given-names>
            <surname>Lebo</surname>
          </string-name>
          , Satya Sahoo,
          <string-name>
            <surname>Deborah</surname>
            <given-names>McGuinness</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Khalid</given-names>
            <surname>Belhajjame</surname>
          </string-name>
          , James Cheney, David Corsar, Daniel Garijo, Stian Soiland-Reyes,
          <string-name>
            <given-names>Stephan</given-names>
            <surname>Zednik</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jun</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Prov-o: The prov ontology</article-title>
          .
          <source>W3C Recommendation</source>
          ,
          <volume>30</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>[Linthicum2017] David S Linthicum</source>
          .
          <year>2017</year>
          .
          <article-title>Cloud computing changes data integration forever: What's needed right now</article-title>
          .
          <source>IEEE Cloud Computing</source>
          ,
          <volume>4</volume>
          (
          <issue>3</issue>
          ):
          <fpage>50</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [Magnini et al.2006]
          <string-name>
            <given-names>Bernardo</given-names>
            <surname>Magnini</surname>
          </string-name>
          , Emanuele Pianta, Christian Girardi, Matteo Negri, Lorenza Romano, Manuela Speranza, Valentina Bartalesi, and
          <string-name>
            <given-names>Rachele</given-names>
            <surname>Sprugnoli</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>I-cab: the italian content annotation bank</article-title>
          .
          <source>In LREC</source>
          , pages
          <fpage>963</fpage>
          -
          <lpage>968</lpage>
          . Citeseer.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [Manning et al.2014]
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          , Mihai Surdeanu, John Bauer, Jenny Finkel, Steven Bethard, and
          <string-name>
            <surname>David McClosky</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The stanford corenlp natural language processing toolkit</article-title>
          .
          <source>In Proceedings of 52nd annual</source>
          <article-title>meeting of the association for computational linguistics: system demonstrations</article-title>
          , pages
          <fpage>55</fpage>
          -
          <lpage>60</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [Metilli et al.2019]
          <string-name>
            <given-names>Daniele</given-names>
            <surname>Metilli</surname>
          </string-name>
          , Valentina Bartalesi, and
          <string-name>
            <given-names>Carlo</given-names>
            <surname>Meghini</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Steps towards a system to extract</article-title>
          .
          <source>In Proceedings of the Text2Story 2019 Workshop</source>
          , page na. Springer.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <source>[OpenAire2019] OpenAire</source>
          .
          <year>2019</year>
          .
          <article-title>European project supporting Open Access</article-title>
          . https://www. openaire.eu/.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <source>[OpenMinTeD2019] OpenMinTeD</source>
          .
          <year>2019</year>
          .
          <article-title>Open Mining INfrastructure for TExt and Data</article-title>
          . https://cordis.europa.eu/project/ rcn/194923/factsheet/en.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <source>[Parthenos2019] Parthenos</source>
          .
          <year>2019</year>
          . Parthenos European Project. //www.parthenos-project.
          <source>eu/.</source>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [Petasis et al.2011]
          <string-name>
            <given-names>Georgios</given-names>
            <surname>Petasis</surname>
          </string-name>
          , Vangelis Karkaletsis, Georgios Paliouras, Anastasia Krithara, and
          <string-name>
            <given-names>Elias</given-names>
            <surname>Zavitsanos</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Ontology population and enrichment: State of the art</article-title>
          .
          <source>In [Pollock and Williams2010] Neil Pollock and Robin Williams</source>
          .
          <year>2010</year>
          .
          <article-title>E-infrastructures: How do we know and understand them? strategic ethnography and the biography of artefacts</article-title>
          .
          <source>Computer Supported Cooperative Work (CSCW)</source>
          ,
          <volume>19</volume>
          (
          <issue>6</issue>
          ):
          <fpage>521</fpage>
          -
          <lpage>556</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [Schmid1995]
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Treetagger - a language independent part-of-speech tagger</article-title>
          .
          <source>Institut für Maschinelle Sprachverarbeitung, Universität Stuttgart</source>
          ,
          <volume>43</volume>
          :
          <fpage>28</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [Schut and Whiteside2007]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Schut</surname>
          </string-name>
          and
          <string-name>
            <given-names>A</given-names>
            <surname>Whiteside</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>OpenGIS Web Processing Service</article-title>
          . OGC project document http://www. opengeospatial.org/standards/wps.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>[SoBigData European Project2016] SoBigData European Project</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deliverable D2.7 - IP principles and business models</article-title>
          . http: //project.sobigdata.eu/material.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <source>[SoBigData2019] SoBigData</source>
          .
          <year>2019</year>
          .
          <article-title>The SoBigData European Project</article-title>
          . http://sobigdata.eu/ index.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <source>[SpeechTEK 20102019] SpeechTEK</source>
          <year>2010</year>
          .
          <year>2019</year>
          .
          <article-title>SpeechTEK 2010 - H-Care Avatar wins People's Choice Award</article-title>
          . http://web.archive. org/web/20160919100019/http://www. speechtek.com/europe2010/avatar/.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [Stanford University2019] Stanford University.
          <year>2019</year>
          .
          <article-title>Stanford CoreNLP - Human Languages Supported</article-title>
          . https://stanfordnlp.github. io/CoreNLP/.
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [Tablan et al.2011]
          <string-name>
            <given-names>Valentin</given-names>
            <surname>Tablan</surname>
          </string-name>
          , Ian Roberts, Hamish Cunningham, and
          <string-name>
            <given-names>Kalina</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>GATE Cloud.net: Cloud Infrastructure for Large-Scale, Open-Source Text Processing</article-title>
          .
          <source>In UK e-Science All hands Meeting.</source>
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [Vossen et al.2016]
          <string-name>
            <given-names>Piek</given-names>
            <surname>Vossen</surname>
          </string-name>
          , Rodrigo Agerri, Itziar Aldabe, Agata Cybulska, Marieke van Erp,
          <string-name>
            <surname>Antske Fokkens</surname>
          </string-name>
          , Egoitz Laparra,
          <string-name>
            <surname>Anne-Lyse</surname>
            <given-names>Minard</given-names>
          </string-name>
          , Alessio Palmero Aprosio,
          <string-name>
            <given-names>German</given-names>
            <surname>Rigau</surname>
          </string-name>
          , et al.
          <year>2016</year>
          .
          <article-title>Newsreader: Using knowledge resources in a cross-lingual reading machine to generate more knowledge from massive streams of news</article-title>
          .
          <source>Knowledge-Based Systems</source>
          ,
          <volume>110</volume>
          :
          <fpage>60</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [Wei et al.2016]
          <string-name>
            <surname>Chih-Hsuan</surname>
            <given-names>Wei</given-names>
          </string-name>
          , Robert Leaman, and
          <string-name>
            <given-names>Zhiyong</given-names>
            <surname>Lu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Beyond accuracy: creating interoperable and scalable text-mining web services</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>32</volume>
          (
          <issue>12</issue>
          ):
          <fpage>1907</fpage>
          -
          <lpage>1910</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>