<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DEVELOPMENT OF RATING SYSTEMS FOR SCIENTOMETRIC INDICES OF UNIVERSITIES</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Kherson State University</institution>
          ,
          <addr-line>27, Universytets'ka St., 73000 Kherson</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Spivakovsky</institution>
          ,
          <addr-line>Vinnik, YuTarasich, KPanova</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article provides a brief overview of the most popular scientometric systems. In our opinion scientific papers are one of the main objects of evaluation of researcher's scientific activity, scientific journals and research organizations. The key idea of the article is description of authors' viewing to develop the automatic analysis system of scientometric indicators, which are various in structure, systems' features and construction on their basis of consolidated ratings. The philosophy of the system is providing open data of different scientometric systems.</p>
      </abstract>
      <kwd-group>
        <kwd>scientific activity</kwd>
        <kwd>information systems</kwd>
        <kwd>scientometric systems</kwd>
        <kwd>bibliometric systems</kwd>
        <kwd>scientometric indicators</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Our time is characterized by the phenomenal acceleration of knowledge accumulation
and the complication of its structure. According to Dell-EMC [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the amount of
produced data is growing more than twice every two years. Based on IDC (International
Data Corporation) report prediction, the global data volume will grow exponentially
to 44 trillion gigabytes by 2020. This tendency is inherent to all areas of human life.
The same problem of storage of information exists in scientometric systems.
      </p>
      <p>Human consciousness is objectively incompetent in processing and storing of
large volume of complex accurate data. Today information technologies are one of
main ways to arrange and create effective tools for organizing the interaction and
processing large amount of information. In our opinion, nowadays in Ukraine the
archaic methods of collection, processing and presentation of information and
scientific activity are adopted.</p>
      <p>Today many information systems attempt to create methods and technologies of
processing and saving information on the activities of scientists. In our opinion,
presentation of information on the university’s scientific activity should be in the
rating form. The rating accumulates several aspects and provides an opportunity to
analyze development in different directions and its changes.</p>
      <p>It is important to note that rating is subjective concept, and based on the principles
of rating, it is possible to model the development of scientific activity at the university
according to its goals, and for the same purpose the influence of a certain element of
the rating can be changed at any time.</p>
      <p>Our system allows automating the processing of information and its presentation,
so we get more accurate result much faster. It is important that result can be obtained
at any moment, and this allows us to get a dynamic picture, that helps to make
decisions related to the scientific activity.</p>
      <p>The availability of information system that would collects, processes and presents
the scientific indicators of organizations is the actual. Therefore, the aim of our work
is to present our experience in developing system of automatic construction of ratings
of scientific organizations based on their scientometric indicators.</p>
      <p>In the article we consider the existing information systems for the processing of
scientific activities (2), describe the key components of our system and basic
principles of its work (3), as well as the basic methods and technologies (4) used for its
implementation.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>
        After analyzing the information systems that run on the activities of scientists,
scientific groups, publishers, etc .., we offer to look for the most interesting projects (in our
opinion)[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <sec id="sec-2-1">
        <title>1. Scopus</title>
        <p>
          Scopus is the largest abstract and citation database of peerreviewed literature,
which indexes more than 7 000 items of scientific, technical and medical journals and
about 4,000 international publishers [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>Scopus is designed to serve the research information needs of researchers,
educators, administrators, students and librarians across the academic community.</p>
        <p>
          Scopus enables researchers to combine their articles under a single profile [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>Our system gets the following attributes from Scopus:
─ author's name;
─ number of publications;
─ scientometric indicators (Hirsch index, citation index, etc.);
─ links to the publications;
─ publication description.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2. Google Scholar</title>
        <p>Google Scholar is freely accessible search system, which indexed the full text of
the scientific publications of all formats and disciplines.</p>
        <p>Google Scholar executes not only informational, but scientometric function. From
the list of results on a hyperlink Search Cited by we can obtain the information how
many and what documents are linked on the publication in data base Google Scholar.</p>
        <p>
          The number in Cited by reflects the degree of authoritativeness and publicity of
publication [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ].
        </p>
        <p>Google Scholar classifies articles in the same way as scientists, by evaluating the
entire text of each article, its author, the publication in which the article appeared, and
the frequency of citing this work in the scientific literature. The most relevant results
are always displayed on the first page.</p>
        <p>Our system gets the following attributes from Google Scholar:
─ scientometric indicators (Hirsch index, citation index, etc.);
─ articles in Google Scholar.</p>
      </sec>
      <sec id="sec-2-3">
        <title>3. Web of Science</title>
        <p>
          Web of Science is an International established data base of Scientific Citation, and
a search platform that combines abstract databases of publications in scientific
journals and patents, including databases of the mutual citation of publications. Web of
Science gives possibility to search among 12 000 magazines and 148 000 materials of
conferences in the field of natural, social, human sciences and arts, which allows to
obtain the most relevant information for your questions. In addition to search, Web of
Science establishes a reference link between the specific research using the cited
materials and thematic links between articles established reputable researchers working
in this field. It is the most extensive database of abstracts. It is available by
subscription [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>4. ORCID</title>
        <p>ORCID (Open Researcher and Contributor ID) is a nonproprietary alphanumeric
code to uniquely identify scientific and other academic authors and contributors [14].
This addresses the problem that a particular author's contributions to the scientific
literature or publications in the humanities can be hard to recognize as most personal
names are not unique, they can change (such as with marriage), have cultural
differences in name order, contain inconsistent use of first-name abbreviations and employ
different writing systems.</p>
        <p>The ORCID offers an open and independent registry intended to be the de facto
standard for contributor identification in research and academic publishing.</p>
        <p>In our system ORCID id is used for unique scientist identification in different
scientific databases and systems.</p>
      </sec>
      <sec id="sec-2-5">
        <title>5. Tutor Network</title>
        <p>Tutor Network is a web service, developed in Kherson State University. It was
developed using such technologies, as ASP.NET MVC, C#, ADO.NET Entity
Framework, JavaScript Framework − JQuery, AJAX and Microsoft SQL Server.</p>
        <p>This service provides an opportunity to take into account all types of scientific
works and publications, which can not be considered by other systems.</p>
        <p>A distinctive feature of this system is displaying information about such scientific
works as manuals, monographs etc.</p>
      </sec>
      <sec id="sec-2-6">
        <title>Technical details will be considered in a subsequent article.</title>
        <p>
          As mentioned in the previous article [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], the research team of Kherson State
University (KSU), which included the authors of the article, took part in a number of
international and national projects whose aim was the development and
implementation of scientific and management processes of analytical information systems and
services. These projects were - Tempus TACIS CP No 20069-1998 "Information
Infra-structure of Higher Education Institutions"; Tempus TACIS MP JEP
230102002 "UnіT-Net: Information Technologies in the University Management Network";
US Department of State Freedom Grant SECAAS-03-GR-214 (DD) "Nothern New
York and Southern Ukraine: New Partnership of University for Business and
Economics Development", etc.
        </p>
        <p>
          During the scientific activity we faced with such problem, as the absence of a
clear mechanism of evaluation of personal contribution in the work of the University,
and incomprehension of the construction of university decisions related to
scientometric. In addition, for the analysis of the scientific indicators of scientists’ group, or a
specific organization, it should be carried out manually. The only option of its partial
automation is rating the organization's profile in Google Scholar. These reasons
motivated us to implement the system of automatic rating construction, its basic principles
were described in the previous article [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>Based on scientometric analysis of the systems described in section 2, we decided
to implement the interaction of our system with such scientometric systems as:
─ Google Scholar, as the most commonly used bibliographic database and it is easy
in using;
─ Scopus, as the largest and the most authoritative abstract database.</p>
      </sec>
      <sec id="sec-2-7">
        <title>On the next step we consider the system architecture.</title>
        <p>3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Structure of the system</title>
      <sec id="sec-3-1">
        <title>The high-level system architecture is shown in Fig.1.</title>
        <p>Parser is search, receiving and transfer the open information of scientometric
indicators of authors and journals provided by Scopus and Google Scholar.</p>
        <p>Parser is made by using xpath queries and regular expressions. Each xpath query
turns to the page of resource. We developed multiplestream parser; each its stream
initializes the parsing of particular resource.</p>
        <p>For interaction with Scopus and Google Scholar parser uses xpath queries, and
API is used for getting information from Web of Science and Tutor Network.</p>
        <p>All data received by parser is stored in the system database. DB of system is
distributed by the data storage. Individual entities of DB are database of scientometric
indicators of researcher and scientific publications.</p>
        <p>Information processing is realized by performing a set of predefined SQL queries.
The main task of the system is automatic construction of consolidated rating of
scientists, research groups, organizations according to indicators of processed
scientometric systems. These indicators are:
1. h-index (Scopus &amp; Google Scholar). The h-index is based on the highest number of
papers included that have had at least the same number of citations;
2. citations (Scopus&amp; Goggle Scholar). Numbers of total citations of documents that
are indexed by the system;
3. publications (Scopus). Total number of documents that are indexed by the system.</p>
        <p>Thus, the main task of the new version was the realization of an automatic
construction of consolidated rating, which allows building a rating for any scientists,
research groups, organizations.</p>
        <p>We determined types of presentation of the results for scientists:
─ profiles of the scientists of the university with generalized information of
scientometric indicators for each database;
─ the rating list of all the scientists of organization;
─ the rating list of all the scientists of organization’s structural subdivision (faculty or
department);
─ the general scientometric information about the university.</p>
      </sec>
      <sec id="sec-3-2">
        <title>The example of system work is shown in the Fig.2.</title>
        <p>The system allows to search a scientist by:
─ ORCID id;
─ Scopus ID;
─ Google Scholar ID;
─ full name of a scientist.</p>
        <p>In response to a search query, the system will return the list of references to
scholars, information about which is in the database, or the message that nothing is found.
After clicking on the link, the personal profile of scientist with generalized
information of scientometric indicators for each database will open.</p>
        <p>On the tab “Rating of Faculties” we get the list of faculties and the highest
Hindex on the faculty. There is an option to select the number of list items that will be
displayed on the page. You can also sort the list of faculties by increasing or
decreasing in alphabetical order or by the value of the H-index. Depending on the selected
tab you can view information from Scopus or Google Scholar. Information on the tab
"Rating of Departments" is displayed similarly.</p>
        <p>On the tab "Rating of Scientists" we see the list of all scientists of the university,
sorted in descending order of such scientific metric indicators as the h-index, the
number of documents in the chosen scientometric database, and the number of
citations.</p>
        <p>Based on the number of scientist’s papers (publications) in scientometric system,
we offer to divide scientists into 4 categories and we defined an equivalent color for
each of them (Fig. 2.):
─ blue ― the number of documents is over 10;
─ green ― the number of documents is in the range of 5 to 10 (inclusive);
─ yellow ― the number of documents is in the range from 1 to 4 (inclusive);
─ red ― no documents.</p>
        <p>On the tab "Scientific metrics", you can find general university statistics, grouped
by scientometric systems, such as:
─ the total number of university staff, registered in selected systems, as well as their
distribution by departments;
─ maximum H-index ;
─ maximum number of documents in selected system;
─ maximum citation index.</p>
        <p>The analysis of the scientific activity of KSU scientists’ shows the highest number
of publications in Scopus has such scholar as Volodymyr Peschanenko (23), and the
maximum number of citations has Alexander Khodosovtsev (98). And the most
hindex has the teachers of the Chair of Botany (7).
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Module of analytics</title>
      <p>As a separate part of the system, module of analytics (written in R language) was
added in the current version. In our system, we use it for processing a scientometric
data array, and graphical representations of statistical data.</p>
      <p>In particular, we used this language opportunity as a time series analysis that
allows the theoretical opportunity to apply models to predict the growth of dynamics of
university rating.</p>
      <p>The diagrams present data showing relation between such scientometric
indicators, as the value of h-index in Scopus and the number of papers in Scopus or Google
Scholar. The diagram can be displayed within the university or faculty, grouped by
departments. The example of diagram is shown in Fig.3.</p>
      <p>As it is evident from the diagrams, the best result is one, located most right and
high.</p>
      <p>Another option for using module in R is possibility to get a variety of statistical
tables, such as the table of universities and the number of the scientists, whose articles
were written in co-authorship with the KSU’s scientists; the table that shows the
general information about set of scientometric indicators of scientists, and total number of
registered users on each department and faculty, etc.</p>
      <p>Thus, the graphical and table representation of statistical data makes the process of
perception of information easier. Moreover, the R language provides an opportunity
to analyze the relations between indicators.</p>
      <p>The main task of developing system was the realization of the possibility of
automated construction of the rating of scientometric indicators for the evaluation of
scientific activity not only in Kherson State University, but in any university. Thus, our
system allows constructing a rating of scientists, research groups and organizations
(as well as their structural subdivisions) by using the API (Application Programming
Interface).</p>
      <p>The system provides access by request in such form:
http: //publication.kspu.edu/api/v1 /teacher? option=
[orcid_id|scopus_id|google_scholar_id|name| | &amp; value= [search_value].</p>
      <p>By specifying the search parameters in the request (some scientometric system
and scientist’s id in it), we will get a list in the json format that looks like this:</p>
      <p>Consequently, API using makes possible to build a rating either for individual
scientists, research groups, and for any university (as well as its structural subdivisions)
by writing its own json parser for processing the received data.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Tools and technologies</title>
      <p>The solutions for automatic construction of ratings require the use of certain products
and technologies:
─ Json</p>
      <p>It is used in the system for the exchange of data for thirdparty systems. Thus, our
system can be a source of data for other resources. It implements the data exchange
via json requests.</p>
      <p>These are universal data structures. Nearly all modern programming languages
support them in any form. It is logical to assume that a data format, independent from
the programming language, should be based on these structures [13].
─ R language</p>
      <p>R is a programming language and free software environment for statistical
computing, data analysis and their graphical representation. R provides a wide variety of
statistical and graphical techniques, and is highly extensible.</p>
      <p>
        Similarly to the previous version, one of the most important algorithms used in the
system is Levenstein algorithm [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        This algorithm is used for solving the problem of determining belonging the
scientist to a particular organization, which arises at changing of the organization's name,
its spelling errors in the article, the change of scientists their place of work, etc. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
The article presents our experience in developing system of automatic construction of
ratings of scientific organizations based on their scientometric indicators in Scopus
and Google Scholar, based on the algorithm of constructing ratings of scientific
publications by the presence/absence in various ratings based on their scientometric
indicators, proposed in our previous article.
      </p>
      <p>In the current version of the system the collection, statistical processing and
presentation of rating of scientometric indicators of scientists, research groups and
organizations are realized. Data source of the system is open information provided by
such scientometric systems as Scopus and Google Scholar.</p>
      <p>Today the system is used to build a consolidated rating of scientists of Kherson
State University and its structural units. It was designed in such a way that it is
possible to deploy it in other universities, and to customize it for their specific individual
goals and tasks. Our system enables to build an automatic rating based on
scientometric indicators by using Application Programming Interface (API).</p>
      <p>The next stage in the development of the system, we see in the realization of its
interaction with Web of Science, as the second in authority international database. The
system and technical details will be considered in more detail in the next article.
10. Ukrainian scientific citation index. Ukrainian Research and Academic Network,
http://uincit.uran.ua/scientists/fronts/about
11. Semantic Scholar. Allen Institute for Artificial Intelligence,
https://www.semanticscholar.org/#subscribe
12. ICI Journals Master List. Index Copernicus,
http://jml2012.indexcopernicus.com/page.php?page=2
13. Introducing JSON. ECMA International, http://www.json.org/
14. What is ORCID?, https://orcid.org/about/what-isorcid/mission</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. Bibliometrics of Ukrainian Science. Social Communications Research Center, http://nbuviap.gov.ua/</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>2. The largest database of peer-reviewed literature</article-title>
          . Elsevier, https://www.elsevier.com/solutions/scopus
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Scientometric databases.
          <source>National library of Ukraine</source>
          ., http://www.nbuv.gov.ua/node/1367
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Science</given-names>
            <surname>Citation</surname>
          </string-name>
          <article-title>Index for scientists</article-title>
          .
          <source>Regional Center of New Information Technologies</source>
          , http://index.petrsu.ru/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Spivakovsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinnyk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Tarasich</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Web Indicators of ICT Use in the Work of Ukrainian Dissertation Committees and Graduate Schools as Element of Open Science</article-title>
          . In: Yakovyna V.,
          <string-name>
            <surname>Mayr</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikitchenko</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zholtkevych</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spivakovsky</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batsakis</surname>
            <given-names>S</given-names>
          </string-name>
          . (eds) Information and Communication Technologies in Education, Research, and
          <string-name>
            <surname>Industrial Applications. ICTERI</surname>
          </string-name>
          <year>2015</year>
          .
          <source>Communications in Computer and Information Science</source>
          , vol
          <volume>594</volume>
          ., pp.
          <fpage>3</fpage>
          -
          <lpage>19</lpage>
          . Springer, Cham (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Spivakovsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinnyk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarasich</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poltoratskiy</surname>
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Design and development of information system of scientific activity indicators</article-title>
          . .
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spivakovsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikitchenko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ginige</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mayr</surname>
            ,
            <given-names>H. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plexousakis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zholtkevych</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burov</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kharchenko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kobets</surname>
          </string-name>
          , V. (Eds.):
          <article-title>ICT in Education, Research and Industrial Applications: Integration, Harmonization and Knowledge Transfer</article-title>
          .
          <source>Proc. 12th Int. Conf. ICTERI</source>
          <year>2016</year>
          , vol.
          <volume>1614</volume>
          , pp.
          <fpage>103</fpage>
          -
          <lpage>110</lpage>
          . CEUR-WS.org (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lowenstein</surname>
          </string-name>
          , V .:
          <article-title>Binary codes with correction for deletions, insertions and substitutions of character</article-title>
          .
          <source>Reports, USSR Academy of Sciences 163.4</source>
          (
          <year>1965</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>International</given-names>
            <surname>Projects</surname>
          </string-name>
          . Kherson State University, http://www.kspu.edu/
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>World</surname>
          </string-name>
          <article-title>'s Data More Than Doubling Every Two Years</article-title>
          . Dell-EMC, https://www.emc.com/about/news/ press/2011/20110628-
          <fpage>01</fpage>
          .htm
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>