<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Knowledge Discovery on Scopus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Paolo Fornacciari</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Monica Mordonini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Nonelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Sani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Tomaiuolo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Parma</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>s of his articles. As a case study, we have limited our research to the authors who have published at least one article about Sentiment Analysis, in a decade. Starting from the more relevant terms extracted from abstracts, we then perform a clusterization of users. This shows the emergence of some subtopics of Sentiment Analysis, which are studied by distinct groups of authors.</p>
      </abstract>
      <kwd-group>
        <kwd>Data mining</kwd>
        <kwd>Social Networks</kwd>
        <kwd>Clustering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Since data is being produced at increasing volumes, the computational challenge
to extract useful knowledge from it is also becoming more and more important.
But, having to deal with large quantity of data also means that much latent
information can be extracted and leveraged for further analysis.</p>
      <p>In this work, we deal with information gathered from Scopus, a repository
of metadata about scientific research papers. In fact, the scientific production is
also accelerating, as emerging countries show great ambitions about innovation
and both new and established institutions use citation-based metrics for hiring
researchers and managing career advancements.</p>
      <p>In this growing quantity of data, in particular, we analyze the social graph
of researchers and their research topics. For highlighting communities of
researchers, in particular, we use a directed and weighted social graph, based on
citations. A citation from an article to another corresponds to the creation of
arcs among the respective authors. On the other hand, we are also interested in
obtaining knowledge about the research topics in which each author works. For
this purpose, we apply data mining techniques to the textual abstracts of the
articles he/she has written. In fact, mining topics from abstracts provides more
cue than explicitly assigned keywords, which are often used inconsistently.</p>
      <p>The authors included in our research are those who have published at least
one article about Sentiment Analysis, in a decade. A practical API is provided
by Scopus, for performing such queries. We apply data mining techniques to
gather the most significant terms for each author, and then we use these terms
as features to cluster them. As a result of our analysis, we notice the existence
of some subtopics, which were not obviously foreseeable in advance and which
distinguish di↵erent groups of authors.</p>
      <p>The rest of the paper is organized as follows. The next section briefly
discusses some background topics and related work. Then, Section 3 presents the
methodological and practical aspects of the research. Section 4 shows the result
of this research and some example of latent knowledge that can be extracted from
this kind of data. The final section provides some concluding remarks, about the
perspectives of this kind of research.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Automatic Knowledge Discovery and Data Mining are becoming more and more
important, since data are produced and collected at significant rate. One of the
first research works about these topics is presented in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Before the increase
of Internet usage and the era of Big Data, the authors discussed the need for
a new generation of computational theories and tools to assist humans in the
extraction of useful information (knowledge) from the rapidly growing volumes
of digital data. These methods are central to the field of Knowledge Discovery
in Databases (KDD), which tries to make sense of data.
      </p>
      <p>
        Big Data concerns large-volume, complex, growing data sets with multiple,
autonomous sources. Wu et al. presented the features of the Big Data revolution
and proposed a model to process them, from the data mining perspective [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
The world of social networking and Web is one of the main source of Big Data
and the field of the research in the application of data mining techniques to the
World Wide Web is called Web mining [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        A Social Networking Service (also Social Networking Site, SNS) is an
online platform that is used by people to build social networks or social relations
with other people who share similar personal interests, or activities, or real-life
connections. The main features of SNSs are illustrated in [
        <xref ref-type="bibr" rid="ref11 ref15">11, 15</xref>
        ], while an
introduction to the Social Network Analysis is provided in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], Sentiment
Analysis (SA), carried out by a framework of Bayesian classifiers, is applied to a
Twitter channel with the aim of detecting online communities characterized by
the same mood. SA is one of the techniques for the extraction of opinion-oriented
information, a comprehensive survey of Opinion Mining is presented in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        One of the most important problems in the field of social network analysis is
community detection, aimed at clustering the nodes on the basis of their social
relationships. The description of a new algorithm, named PaNDEMON, able
to exploit the parallelism of modern architectures while preserving the quality
of results, is reported in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], parallel computation is exploited to find
relevant structures in complex systems, while in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] the complexity is managed
in a totally di↵erent way, by using a combined approach of sentiment analysis
and social network analysis in order to detect communities that are e↵ectively
homogeneous within a network.
      </p>
      <p>
        Participation in social networks has long been studied as a social phenomenon
according to di↵erent theories [
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ], and, in particular, the notion of social capital
refers to the person’s benefit due to his relations with other persons, including
family, colleagues, friends and generic contacts. The role of social capital in the
participation in online social networking activities is shown in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] in the various
cases of virtual organizations or virtual teams. Social networking systems are
bringing a growing number of acquaintances online, both in the private and
working spheres and, in businesses, several traditional information systems (such
as CRMs and ERPs) have also been modified in order to include social aspects [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Scopus is an abstract and indexing database with full-text links, produced
by the Elsevier Co. Scopus database provides access to articles and to the
references included in those articles, allowing to search both forward and backward
in time [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Scopus, together with Web of Science, represents one of the most
authoritative bibliographic databases in order to perform bibliometric analyses
and comparisons of countries or institutions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Several works have been conducted in knowledge discovery on bibliographic
databases. Newman constructed networks of scientists by using data from three
bibliographic databases in biology, physics, and mathematics [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. In these
networks the nodes are scientists, and two scientists are connected if they have
coauthored a paper. The author reports various analyses conducted over these
networks, such as the number of papers written by each author, the number of
coauthors, the typical distance between scientists through the network, and the
variation of di↵erent patterns of collaboration between subjects and over time.
      </p>
      <p>
        Rankings based on publications can supply useful data in a comprehensive
assessment process of academic and industrial research. A framework for an
automatic and versatile publications ranking for research institutions and scholars is
shown in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The authors demonstrate that the biggest diculty in developing
their framework was the insucient availability of bibliographic data containing
the institution with which an author was aliated when the paper was published.
      </p>
      <p>
        Using social network analysis to study scientific collaboration patterns has
become an important research topic in information systems research. For
example, research publications in the International Conference on Information Reuse
and Integration provided the data in order to identify the most popular
research topics, as well as the most productive researchers and organizations in
that area [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The underlying idea is to treat scientific growth as a process of
knowledge di↵usion, so the most productive authors (or leaders) are often the
ones who introduce new ideas and thus have a great impact on the research
community.
      </p>
      <p>
        An important type of social network is the co-authorship network, which has
been studied both for social community extraction and social entity ranking. In
general, to construct this kind of network we need to consider the co-authorship
relation between two authors as a collaboration. Instead, in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], the authors
introduce a supportiveness measure on co-authorship networks, with quite
interesting results. They model the collaboration between two authors as a supporting
activity between them, then they treat the supportiveness ranking problem as a
reverse k nearest neighbor (k-RNN for short) searching problem on graphs.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the study of the scientific collaboration and endorsement is performed
by using not only the co-authorship network but also the citation network. The
results show that productive authors tend to directly coauthor with and closely
cite colleagues sharing the same research interests; they do not generally
collaborate directly with colleagues having di↵erent research topics, but they directly
or indirectly cite them; highly cited authors do not generally coauthor with each
other, but closely cite each other.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>Scopus, which is commonly accessed through its web interface, is a bibliographic
database containing abstracts and citations for scientific articles1. It also contains
useful pieces of information about the authors and the institutions to which these
authors are aliated (universities, companies, research centers, ...). The records
stored inside the database can be categorized into three kinds of “information
elements”: abstracts, authors and aliations. These categories can be seen as
three di↵erent classes of objects, each one with its specific sub-attributes (i.e. an
author has an ID, a name, a surname, one or more aliations, and so on).</p>
      <p>To programmatically access the data inside Scopus, Elsevier (the company
running the database) also developed a set of HTTP RESTful APIs, whose
up-to-date specifications can be found on the Elsevier developers portal2. At
the time of writing this article, there are 11 Scopus APIs publicly available.
Among these, for the sake of our research, just 6 have been used: ScopusSearch,
AbstractRetrieval, AuthorSearch, AuthorRetrieval, AliationSearch and
AliationRetrieval. Given the names of those APIs, it is easy to see that for each
class of information elements one can find two kind of APIs: one that can be
used to search among elements of a specific type, and another one that can be
used to retrieve a specific object, given its ID. Each API has its specific query
variables and conventions, which are described inside the already linked Elsevier
developers website. The flagship API, named ScopusSearch, searches among the
abstract objects.</p>
      <p>To manage the HTTP requests and join the JSON responses coming from the
server into bigger datasets, we have built a basic Python library wrapping the 6
main Scopus APIs and a set of Python scripts and Jupyter Notebooks (IPython
Notebooks) which can be executed in series to automate the 5 steps of the
Knowledge Discovery in Databases (KDD) pipeline. Outside of the experiment
here presented, the library and the scripts can be used (and have been used
during our tests) to query the Scopus APIs with any of the available parameters
and search keys, and in particular to reproduce the whole knowledge discovery
process described in the next steps. As said, the code needed to repeat this</p>
      <sec id="sec-3-1">
        <title>1 http://www.scopus.com/</title>
      </sec>
      <sec id="sec-3-2">
        <title>2 http://dev.elsevier.com/</title>
        <p>experiment or try other queries is written in Python and can be downloaded
from a public repository3.</p>
        <p>This knowledge discovery experiment is conducted applying the 5-steps KDD
pipeline (Figure 1) on the Scopus database. As usual, in addition to the 5 KDD
steps, this involves also a 0-step in which the domain (in this case, the Scopus
systems) is analyzed in order to understand what kind of data is available and
how to access it.
The first step of the KDD process, used to extract information, involves selecting
a more specific subset of data from the whole database. Applying this on Scopus,
we have decided to download responses from the ScopusSearch API for a specific
query.</p>
        <p>This API replies to HTTP GET requests with JSON (or XML) payloads
containing a list of Abstracts matching the input search query, up to the limit of
5000 results per single query. Search results can be downloaded from the API in
lists of only 25 abstracts per HTTP request, and must then be joined and saved
in the file system for the next steps.</p>
        <p>This process has been automated and the requests and responses are
managed by the first step of the pipeline. In the experiment, this is used to
download from the ScopusSearch API the data matching the query
“TITLE-ABSKEY({sentiment analysis})”, which searches for all the abstracts having the
exact whole sentence “sentiment analysis” in their “title”, “abstract” or
“keywords” attributes. The query produces 3648 results.</p>
        <p>While doing a very first cleanup of the data, dropping invalid or corrupted
abstracts, the JSON responses are joined creating a JSON array of abstracts.
Each of these objects contains not only the text of the abstract and other data
related to the article, but also some pieces of information about its authors</p>
      </sec>
      <sec id="sec-3-3">
        <title>3 http://www.github.com/valleymanbs/scopus-json/</title>
        <p>and their aliations. A very first cleaning of the results, dropping invalid or
corrupted data, produces the Target Data on which the second step of the KDD
process is run.
3.2</p>
        <sec id="sec-3-3-1">
          <title>KDD Step 2: Preprocessing</title>
          <p>Matching the second step of the classic KDD pipeline, the Target Data is
preprocessed, doing a deeper cleanup of useless information and integrating the
missing data, running the appropriate requests to the AuthorRetrieval and the
AliationRetrieval APIs.</p>
          <p>After this first cleanup and integration process, a search is performed on the
already downloaded articles, in order to obtain the ones with citations. A set of
queries is then run, again on the ScopusSearch API, to download all the articles
citing an article which is already inside the dataset. The newly download data
are joined into the original dataset, after being cleaned and integrated.</p>
          <p>Using the “TITLE-ABS-KEY({sentiment analysis})” dataset, which is
downloaded in the first step of the experiment, a final CSV dataset (10307 abstracts)
is produced, which also contains useful data related to their authors. This last
thing is particularly important, since the next steps will focus on the authors of
articles.
3.3</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>KDD Step 3: Transformation</title>
          <p>Going further in the process of extracting knowledge from data, two di↵erent
kinds of transformation processes are applied to the Preprocessed Data obtained
from step 2, both supported by the code included in the project repository.</p>
          <p>The first transformation stage creates a graph from the previously
preprocessed abstracts dataset: the entries of this dataset represent a social network,
in which a node represents an author and the directed edges represent the
citations between the authors of the articles which formed the dataset coming from
the previous step. The graphs produced in the pipeline are formatted following
Gephi standards, and will be used in future works in order to analyze the social
networks between authors of the same research field. Gephi itself is a well known
open source software for graph analysis4.</p>
          <p>The second kind of transformation consists in the extraction of all the
authors data from the preprocessed dataset of abstracts. This approach is extended
by grouping, under each author, all the text found in the abstracts of his/her
articles. At an intermediate stage, this transformation creates a dataset (Table 1)
in which each line relates to a single author with all the text contained in the
titles, abstracts and keywords attributes found inside the abstracts information
elements downloaded from the Scopus APIs in the first 2 steps of the KDD
pipeline.</p>
          <p>This authors-text dataset is rich in unstructured information and is then
processed to obtain structured numerical data for the data mining step, going</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>4 http://gephi.org/</title>
        <p>author id text
0 10041296000 Media-aware quantitative trading based on publ...
1 10042379000 A context-dependent sentiment analysis of onli...
2 10043514200 The networked cultural di↵usion of Korean wav...
through a TF-IDF vectorization process. TF-IDF is a bag of words technique
to vectorize text data in which a vocabulary is extracted from all the words
contained in a Corpus of several text Documents (after a stemming process).
Each term of the corpus vocabulary serves as an index of the vector that
represents the unstructured text document as a structured indexed numerical vector.
The weights inserted in each vector are calculated weighting each single word
frequency inside the corresponding document against the term frequency inside
all the documents in the corpus, following the formula
wi,j = tf i,j ⇥ log
✓ N ◆
df i
(1)
where tf i,j is the number of occurrences of i in j, df is the number of
documents containing i, and N is the total number of documents.</p>
        <p>In our particular vectorization process the whole sum of text contained in
the dataset makes up the corpus, while all the articles associated to an author
are a single document inside this corpus; from this unstructured text data has
been extracted a vocabulary vector made of the 500 most frequent terms inside
the whole corpus, counting 27956 unique words, stemmed as in Table 2.</p>
        <p>The TF-IDF vectorization is the last stage of the transformation process and
finally produces the Transformed Data on which the KDD process will go on: a
matrix of numerical data in which each line is a vector representing an author.</p>
        <p>In other words, from the TF-IDF vectorization has been obtained a set of
vectors in a 500-dimensional features space: each unique author found in the data
originally downloaded from Scopus is represented in this 500D Vector Space
Model by a vector which weights have been computed using the frequency of
words found in the text written by the corresponding author, compared against
the same words frequency inside all the text written by all the authors in the
dataset.</p>
        <p>Cluster</p>
        <p>BLUE
ECommerce</p>
        <p>RED
Algorithms</p>
        <p>GREEN
Social Media</p>
        <p>Authors
1082
3376
3388</p>
        <p>Terms
reviews, product, features, aspects, users,
mining, rating, modeling, customers, online
classification, word, features, modeling, text,
emotional, learning, language, polarized, lexicons</p>
        <p>social, media, twitter, users, network,
tweeting, mining, systems, predict, topic
In this step the Transformed Data obtained from step 3 has been analyzed with
an unsupervised learning approach, in order to see if analyzing text data, there
are emerging clusters of authors in this field of research.</p>
        <p>To verify this, the vectors coming from the TF-IDF weighting have been
clustered using K-Means and Expectation Maximization. This two methods have
been used to check the results of a clusterer against the other.</p>
        <p>
          Clustering with K-Means and E-M requires to set the number K of clusters
as parameter: after testing several K values, in the experiment we set K=3. This
value has been experimentally found to be the one giving stable and comparable
results between the two clustering algorithms. Along with the experimental
approach, this value has also been found to be the one giving the higher Silhouette
Score [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] when running a Latent Semantic Analysis on the same dataset with
a set of features reduced to 3 dimensions with PCA [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
The unsupervised learning process, used to extract information from text,
eventually leads latent information to emerge. This paves the way for the
identification of the main topics treated by the authors in the chosen research field.
Results obtained for the Sentiment Analysis field are presented in next section.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental results</title>
      <p>After the clustering process, the most frequent words, according to TF-IDF
weights, used by all the authors of each one of the obtained clusters are analyzed.
Two clustering algorithms are used and confronted. The aim is to understand
which topics are treated by the clustered authors and to find relationships
between clusters produced respectively by K-Means (Table 3) and E-M (Table 4).</p>
      <p>The previous step of unsupervised learning process is used to extract
information from text. The obtained data leads to identifying three main
clusters and topics treated by the authors in the “Sentiment Analysis” research
field (Figure 2, 3). According to the highlighted terms, we have labeled these
Cluster</p>
      <p>BLUE
ECommerce</p>
      <p>RED
Algorithms</p>
      <p>GREEN
Social Media</p>
      <p>Authors
1058
4661
2127</p>
      <p>Terms
reviews, product, features, aspects, mining,
users, modeling, customers, rating, online
modeling, emotional, classification, text, features,
word, learning, language, systems, polarized
social, media, twitter, network, users,
tweeting, mining, events, predict, topic
three subtopics (corresponding to the obtained clusters) respectively as:
“ECommerce”, “Algorithms”, “Social media”.</p>
      <p>However, as can be seen from the clustering output and from the occurrence of
the same top-words inside the di↵erent clusters, the clustering process conducted
both with K-Means and E-M tends to be weak. This is probably due to the topic
searched, which is very tight, with low variance between authors (new field, few
articles).</p>
      <p>
        Finally, we have confronted the results of the clusterization process, by terms,
with the communities of authors, obtained by the social graph topology [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In
particular, communities have been identified by the density of weighted and
directed social arcs, obtained from citations among articles. It is apparent, quite
at first sight, that clusters are spread all over the social graph (see Figure 4). In
fact, community detection can distinguish three communities, with sizes
comparable to those of content-based clusters. However, in each community, clusters
are represented with almost the same incidence as in the whole graph. The same
kind of results are obtained for di↵erent runs of the community detection
process, generating more and smaller communities, where clusters are nevertheless
all well present.
      </p>
      <p>This means that, in this case, the subtopics of research and the citation
graph represent mostly orthogonal and independent aspects of analysis. For the
topic of “Sentiment Analysis”, in fact, the social graph of authors has a large
and highly connected core, with low modularity. Thus, the results indicate that
authors often cite other authors, working in di↵erent communities and di↵erent
subtopics.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The availability of large quantity of data often allows analysts to discover latent
knowledge. We applied the techniques of social network analysis, data mining
and clustering to data about scientific research articles. In particular, we used
metadata available in the Scopus repository, to create a social graph of scientific
authors, starting from citations among their articles. Moreover, using data
mining techniques, we have inferred some relevant research topics for each author,
from the textual analysis of the abstracts of his articles.</p>
      <p>In particular, we have limited our analysis to the case study of authors who
have published at least one article about “Sentiment Analysis”, in a decade.
After associating each author with the most relevant topics emerging from the text
analysis of his abstracts, we have performed a clusterization of authors. This
process has finally brought to light some groups of authors, who conduct their
research about distinct application areas of Sentiment Analysis, namely:
algorithms for sentiment analysis; marketing and e-commerce; social media analysis.
Contrasting the results of the clustering process, by relevant terms, and the
community detection process, based on the social graph of citations, indicates that
authors often cite other authors, working in di↵erent communities and di↵erent
subtopics.</p>
      <p>Fornacciari et al.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amoretti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fornacciari</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mordonini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomaiuolo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Local-first algorithms for community detection</article-title>
          .
          <source>In: 2nd International Workshop on Knowledge Discovery on the WEB</source>
          ,
          <source>KDWeb</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Angiani</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fornacciari</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mordonini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomaiuolo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Models of participation in social networks</article-title>
          .
          <source>In: Social Media Performance Evaluation and Success Measurements</source>
          , p.
          <fpage>196</fpage>
          .
          <string-name>
            <given-names>IGI</given-names>
            <surname>Global</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Archambault</surname>
          </string-name>
          , E´.,
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gingras</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , Larivi`ere, V.:
          <article-title>Comparing bibliometric statistics obtained from the web of science and scopus</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>60</volume>
          (
          <issue>7</issue>
          ),
          <fpage>1320</fpage>
          -
          <lpage>1326</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>V.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guillaume</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lambiotte</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lefebvre</surname>
          </string-name>
          , E.:
          <article-title>Fast unfolding of communities in large networks</article-title>
          .
          <source>Journal of statistical mechanics: theory and experiment</source>
          <year>2008</year>
          (
          <volume>10</volume>
          ),
          <source>P10008</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Burnham</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          :
          <article-title>Scopus database: a review</article-title>
          .
          <source>Biomedical digital libraries 3(1)</source>
          ,
          <volume>1</volume>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cooley</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mobasher</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srivastava</surname>
          </string-name>
          , J.:
          <article-title>Web mining: Information and pattern discovery on the world wide web</article-title>
          .
          <source>In: Tools with Artificial Intelligence</source>
          ,
          <year>1997</year>
          . Proceedings., Ninth IEEE International Conference on. pp.
          <fpage>558</fpage>
          -
          <lpage>567</lpage>
          . IEEE (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cristani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fogoroasi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomazzoli</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Measuring homophily</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          . vol.
          <volume>1748</volume>
          (
          <year>2016</year>
          ), https://www.scopus. com/inward/record.uri?eid=
          <fpage>2</fpage>
          -
          <lpage>s2</lpage>
          .
          <fpage>0</fpage>
          -
          <lpage>85012298603</lpage>
          &amp;partnerID=
          <volume>40</volume>
          &amp;md5=
          <fpage>81df100456c2118853ca823496097c79</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cristani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomazzoli</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olivieri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Semantic social network analysis foresees message flows</article-title>
          .
          <source>In: ICAART 2016 - Proceedings of the 8th International Conference on Agents and Artificial Intelligence</source>
          . vol.
          <volume>1</volume>
          , pp.
          <fpage>296</fpage>
          -
          <lpage>303</lpage>
          (
          <year>2016</year>
          ), https://www.scopus.com/inward/record.uri?eid=
          <fpage>2</fpage>
          -
          <lpage>s2</lpage>
          .
          <fpage>0</fpage>
          -
          <lpage>84969287486</lpage>
          &amp;partnerID=
          <volume>40</volume>
          &amp;md5=
          <fpage>6d7a0bb42fd4f45cdb48b8dc1193907a</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Day</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ong</surname>
            ,
            <given-names>C.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            ,
            <given-names>W.L.</given-names>
          </string-name>
          :
          <article-title>An analysis of research on information reuse and integration</article-title>
          .
          <source>In: Information Reuse &amp; Integration</source>
          ,
          <year>2009</year>
          . IRI'09. IEEE International Conference on. pp.
          <fpage>188</fpage>
          -
          <lpage>193</lpage>
          . IEEE (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Scientific collaboration and endorsement: Network analysis of coauthorship and citation networks</article-title>
          .
          <source>Journal of informetrics 5(1)</source>
          ,
          <fpage>187</fpage>
          -
          <lpage>203</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ellison</surname>
            ,
            <given-names>N.B.</given-names>
          </string-name>
          , et al.:
          <article-title>Social network sites: Definition, history, and scholarship</article-title>
          .
          <source>Journal of Computer-Mediated Communication</source>
          <volume>13</volume>
          (
          <issue>1</issue>
          ),
          <fpage>210</fpage>
          -
          <lpage>230</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Fayyad</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piatetsky-Shapiro</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>From data mining to knowledge discovery in databases</article-title>
          .
          <source>AI</source>
          magazine
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <volume>37</volume>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Fornacciari</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mordonini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomaiuolo</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A case-study for sentiment analysis on twitter</article-title>
          .
          <source>In: WOA</source>
          . pp.
          <fpage>53</fpage>
          -
          <lpage>58</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Fornacciari</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mordonini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomauiolo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Social network and sentiment analysis on twitter: towards a combined approach</article-title>
          .
          <source>In: 1st International Workshop on Knowledge Discovery on the WEB</source>
          ,
          <source>KDWeb</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Franchi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poggi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomaiuolo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Blogracy: A peer-to-peer social network</article-title>
          .
          <source>International Journal of Distributed Systems and Technologies (IJDST) 7</source>
          (
          <issue>2</issue>
          ),
          <fpage>37</fpage>
          -
          <lpage>56</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Franchi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poggi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomaiuolo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Social media for online collaboration in firms and organizations</article-title>
          .
          <source>International Journal of Information System Modeling and Design (IJISMD) 7</source>
          (
          <issue>1</issue>
          ),
          <fpage>18</fpage>
          -
          <lpage>31</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. Han,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Pei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>Understanding importance of collaborations in co-authorship networks: A supportiveness analysis approach</article-title>
          .
          <source>In: Proceedings of the 2009 SIAM International Conference on Data Mining</source>
          . pp.
          <fpage>1112</fpage>
          -
          <lpage>1123</lpage>
          . SIAM (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Hofmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Probabilistic latent semantic indexing</article-title>
          .
          <source>In: Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval</source>
          . pp.
          <fpage>50</fpage>
          -
          <lpage>57</lpage>
          . ACM (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          :
          <article-title>Coauthorship networks and patterns of scientific collaboration</article-title>
          .
          <source>Proceedings of the national academy of sciences 101(suppl 1)</source>
          ,
          <fpage>5200</fpage>
          -
          <lpage>5205</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , et al.:
          <article-title>Opinion mining and sentiment analysis</article-title>
          .
          <source>Foundations and Trends R in Information Retrieval</source>
          <volume>2</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , R.N.:
          <article-title>Automatic and versatile publications ranking for research institutions and scholars</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>50</volume>
          (
          <issue>6</issue>
          ),
          <fpage>81</fpage>
          -
          <lpage>85</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Rousseeuw</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>Silhouettes: a graphical aid to the interpretation and validation of cluster analysis</article-title>
          .
          <source>Journal of computational and applied mathematics 20</source>
          ,
          <fpage>53</fpage>
          -
          <lpage>65</lpage>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Sani</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amoretti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vicari</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mordonini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pecori</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cagnoni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serra</surname>
          </string-name>
          , R.:
          <article-title>Ecient search of relevant structures in complex systems</article-title>
          .
          <source>In: AI*IA 2016 Advances in Artificial Intelligence</source>
          . pp.
          <fpage>35</fpage>
          -
          <lpage>48</lpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Scott</surname>
          </string-name>
          , J.:
          <article-title>Social network analysis</article-title>
          .
          <source>Sage</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>G.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Data mining with big data. ieee transactions on knowledge and data engineering 26(1</article-title>
          ),
          <fpage>97</fpage>
          -
          <lpage>107</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>