<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Prediction Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rand Alchokr</string-name>
          <email>rand.alchokr@ovgu.de</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rayed Haider</string-name>
          <email>rayedhaider95@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yusra Shakeel</string-name>
          <email>yusra.shakeel@kit.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Leich</string-name>
          <email>tleich@hs-harz.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gunter Saake</string-name>
          <email>saake@ovgu.de</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacob Krüger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Eindhoven University of Technology</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Hochschule Harz</institution>
          ,
          <addr-line>Wernigerode</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Karlsruhe Institute of Technology</institution>
          ,
          <addr-line>Karlsruhe</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>METOP GmbH</institution>
          ,
          <addr-line>Magdeburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Otto-von-Guericke University</institution>
          ,
          <addr-line>Magdeburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>77</fpage>
      <lpage>89</lpage>
      <abstract>
        <p>Measuring the success and impact of a scientific publication is an important, thus controversial matter. Despite all the criticism, it is widespread that citation counts is considered a popular indication of a publication's success. Therefore, in this paper, we use a machine learning framework to test the ability of alternative metrics (altmetrics) to predict the future impact of papers reflected in the citation counts. To achieve the experiment, we extracted 7,588 papers from 10 computer science journals. To build the feature space for the prediction problem, 14 diferent altmetric indices were collected, 3 feature selection approaches, namely, Variance threshold, Pearson's Correlation, and Mutual information method, were used to minimize the feature space and rank the features according to their contribution to the original dataset. To identify the classification performance of these features, three classifiers were used: Decision Tree, Random Forest, and Support Vector Machines. According to the experimental data, altmetrics can predict future citations and the most useful altmetrics indications are social media count, tweets, news count, capture count, and full-text view, with Random Forest outperforming the other classifiers.</p>
      </abstract>
      <kwd-group>
        <kwd>Bibliometric</kwd>
        <kwd>alternative metrics</kwd>
        <kwd>machine learning</kwd>
        <kwd>computer science</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        A successful publication is a desirable goal for any researcher, irrespective of their scientific
ifeld. However, judging how successful a published paper is and measuring that success is
considered a critical issue. Furthermore, forecasting scientific impact and success is becoming
an essential regular task for the hiring committees, funding agencies, and department heads for
recruitment decisions and rewards [
        <xref ref-type="bibr" rid="ref26 ref3 ref7">3, 7, 26</xref>
        ]. Through this a merit-based career advancement
scheme is developed, that forecasts the individual’s performance based on past achievements and
projects future performance. However, distilling the contents of each article into an appraisal
nEvelop-O
∗Corresponding author.
CEUR
Workshop
Proceedings
of an individual’s past, present, and future influence and determining an acceptable ranking of
candidates is significantly challenging when presented with candidate pools ranging from a
few hundred for tenure-track positions to thousands for fellowship and grant competitions.
      </p>
      <p>
        In the past decades, researchers have relied heavily on quantitative indicators for evaluating
the scientific success of a given research body. Citation frequency is a well-known criterion for
research evaluation and despite all the criticisms of citations not being a perfect and objective
means of measuring scientific quality, citation counts are still widely referred to as a foremost
indicator of the impact and success of a publication in the scientific community. Recently, there
has been extensive research to investigate the link between a paper’s citations and all possible
factors correlated to it [
        <xref ref-type="bibr" rid="ref10 ref15 ref23 ref24 ref25 ref5 ref6">5, 10, 25, 23, 24, 15, 6</xref>
        ]. Additionally, research is becoming interested in
forecasting the future success of a paper [
        <xref ref-type="bibr" rid="ref13 ref2 ref38 ref39 ref7">7, 13, 38, 39, 2</xref>
        ]. Among these factors, bibliometrics
and altmetrics are deemed of the utmost relevance. While bibliometrics is the traditional indices
reflecting the characteristics as well as the credibility of papers, authors, and publishing venues
(e.g. citations, h-index of the author, cite score of the venue, etc.), altmetrics have been recently
introduced to capture the spread of a publication in various online platforms (e.g. Wikipedia,
Twitter, Facebook). The combination of both bibliometrics and altmetrics, is recommended
by researchers to complement their pros and cons [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Recent studies have looked at the
relationship between bibliometric indicators and altmetrics, taking into consideration
peerreviewed quality evaluation methods [
        <xref ref-type="bibr" rid="ref24 ref27 ref30 ref31 ref32 ref33 ref8">24, 8, 27, 33, 32, 30, 31</xref>
        ]. The application of altmetrics in
research assessment raises the question of whether the data collected by altmetrics is a good
predictor of future success and whether it correlates with citation.
      </p>
      <p>
        On the other hand, the remarkable progress in the field of machine learning (ML) has matured
a plethora of techniques that can eficiently handle various forecasting tasks. In the context of
predicting papers’ citations using bibliometrics and altmetrics, multiple studies formulate the
problem into a regression task that considers continuous values of both features and the output
[
        <xref ref-type="bibr" rid="ref13 ref17 ref2 ref22 ref29">13, 2, 17, 22, 29</xref>
        ], whereas other studies considered classification algorithms that generate
categorical outcome [
        <xref ref-type="bibr" rid="ref38 ref39">38, 39</xref>
        ]. In this paper, we rely on both of these metrics to find out which
altmetrics features contribute to forecasting citation counts. We consider a paper successful if it
achieves a high number of citations. We categorize the publications according to their citation
counts. Belonging to a class of higher ranking hints at a more successful paper. The goal of
this study is to determine which features of altmetric are useful in predicting future highly
cited papers and which machine learning model would be the best for this prediction. In our
experiment, we will use Decision Trees, Random Forests, and Support Vector Machines.
      </p>
      <p>In detail, our main contributions in this paper are as follows:
• We collect an extensive dataset comprising papers from 10 computer engineering journals
from 2010 to 2015. Further, we elicit the papers’ citations and altmetrics, aiming to find
the most promising altmetrics formula to predict the future success of a paper.
• We discuss multiple prediction models and compare their accuracy.</p>
      <p>Through our experiments, we aim to provide a better understanding of the usefulness of
altmetrics to indicate the future success of publications.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Background</title>
      <sec id="sec-3-1">
        <title>Next, we present the background needed to understand this paper.</title>
        <sec id="sec-3-1-1">
          <title>2.1. Evaluation Metrics</title>
          <p>
            Peer Reviewing during the scientific evaluation process of papers is an essential part of
publishing academic research, representing an important quality assurance mechanism [
            <xref ref-type="bibr" rid="ref34">34</xref>
            ]. On
the other hand, bibliometrics which represent the traditional metrics are common measures
that the research community relies on when assessing the scientific impact and quality of a
publication [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]. Such metrics have multiple advantages, they mainly facilitate the examination
of large datasets and help decision-making on individuals, institutions, or research grants [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ].
Citation counts, h-index, and impact factor are among the most important metrics used for
assessing the impact and quality of publications, publishing venues, authors, or research in
general. Citation-based metrics are assumed to directly reflect on the impact and quality of a
publication by implying credibility to the reader and reflecting the total impact of a
publication on a research field [
            <xref ref-type="bibr" rid="ref25">25</xref>
            ]. Despite their potential benefits, bibliometrics have always been
criticized in the context of measuring the impact or quality of research, which they do not
necessarily capture [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ]. However, many studies suggest that using bibliometrics is a helpful
complement to mitigate potential biases during traditional peer review.
          </p>
          <p>
            Altmetrics have been recently introduced as means to assess the impact of a publication based
on publicly available interfaces of various online platforms [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ]. These metrics allow researchers
to track the impact of publications beyond traditional bibliographic metrics and help them in
catching the buzz and spread of their research to a broader audience by calculating quantitative
values of user interactions on social platforms, for instance, Wikipedia, Twitter, Facebook, the
number of downloads, views, or read times. It is known that altmetrics may not accurately
represent scientific quality, they lack the evidence, are dificult to measure, commercialized,
and easily manipulated [
            <xref ref-type="bibr" rid="ref23 ref36">36, 23</xref>
            ], however, based on the mentioned benefits, many researchers
argue that altmetrics can serve as an impact indicator and a complement to traditional metrics
[
            <xref ref-type="bibr" rid="ref15 ref23 ref24">23, 24, 15</xref>
            ]. Researchers recommend using both kinds of metrics when assessing the impact or
quality of a publication to complement their pros and cons [
            <xref ref-type="bibr" rid="ref23">23</xref>
            ].
          </p>
          <p>In conclusion, we rely on both kinds of metrics to measure the success of a publication. We
consider a paper successful and has an impact if it has achieved a high number of citations.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>2.2. Predictive Algorithms</title>
          <p>
            By definition, machine learning is a branch of computer science that grew out of artificial
intelligence research into pattern recognition and computational learning theory [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. It is
the learning and building of algorithms that can learn from and make predictions on datasets.
There are three types of machine learning algorithms: 1) Supervised learning algorithms: with
two types: classification and regression, 2) Unsupervised learning algorithms: association,
clustering, and dimensionality reduction, and 3) reinforcement learning.
          </p>
          <p>
            Supervised learning, is defined as learning from labeled training data. The training data is
learned using a supervised learning algorithm, which then creates a prediction function. For
unseen occurrences, the predictive function will be utilized to determine the class label. Linear
regression, Logistic Regression, CART, Naïve Bayes, and K-Nearest Neighbors (KNN) — are
examples of supervised learning. Also, Bagging with Random Forests, Boosting with XGBoost,
and Multilayer Perception (basic ANN). To start with, Naive Bayes applies the assumption of
independence between every set of features, meaning that all features contribute independently
to the probability of the target’s outcome [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]. XGBoost is a scalable tree-boosting system,
it is used widely by data scientists nowadays [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ]. A classifier is an example of a supervised
learning algorithm. Machine learning algorithms that tackle the categorization problem are
known as classifiers.
          </p>
          <p>
            A classification issue is described as a task of determining class labels for fresh observations
based on a training batch of data with a known class label. ANN is a helpful model for
classification, clustering, pattern recognition, and prediction in many fields [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. Random Forest
inputs and random features produce good results in classification—less so in regression. Finally,
The K-Nearest Neighbors (KNN) has often been used in pattern recognition problems. According
to the existing literature, there have been various studies done to investigate the factors that
influence citation and studies that attempted to forecast and estimate future citations.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Related Work</title>
      <p>
        According to existing literature, various studies investigated the factors that influence citation,
while others attempted to forecast and estimate future citations. Some of these studies utilized
the early citation counts to predict the publication‘s future success [
        <xref ref-type="bibr" rid="ref2 ref29 ref35">35, 2, 29</xref>
        ]. Their results
agree on the impact early citations and other related factors have on predicting highly cited
publications. Social media metrics started to gain interest in research. For instance, tweets
had a weak ability to positively predict high citation counts across several disciplines [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. In
the computer science domain, multiple classification methods were used to check whether
the future success of articles depends on bibliometrics or altmetrics, and the results show
that both contribute equally with PCA achieving the best performance [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ]. Another study
investigated altmetrics specifically using the ”Altmetric Attention Scores”, but this time to
predict the retraction of the articles. The results show that roughly one-fourth of the retractions
are properly predicted using five alternative metrics Copiello [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Another study by Akella
et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] used atmetrics social media features to predict early and long-term citation counts
using several classifiers and regressors, their main results indicate that Mendeley readership
plays a crucial role in determining the early citations. We built our experiments on theirs, but
ifrst by determining the most influential features.
      </p>
      <p>We present an overview of the related work in Table 1 collected by conducting a literature
search on Scopus1 digital library. For each study, we display the type of prediction, and feature
selection methods that provide the necessary background information to guide our experiments.
Overall, the literature demonstrates that researchers have explored a variety of machine learning
algorithms and features in their eforts to predict the academic influence of research publications.
neural networks and ensemble models performed
better, with high predicted accuracy and F-1
scores. Mendeley readership plays a crucial role
in determining the early citations
Paper Potential Index (PPI) model and multi-feature model PPI model outperforms the multi-feature model in
terms of range-normalized RMSE and it better
interprets changes in citation without requiring
parameter adjustments. In terms of Mean Absolute
Percentage Error and Accuracy, the multi-feature
model outperforms the PPI model; nevertheless,
its predictive performance is more dependent on
parameter modification
Stepwise multiple regression used to select appropriate features and Regression model works well in this situation
to build a regression model for explaining the relationship between where bibliometrics have high predictability
comcitation impact and the chosen features (external features of a paper, pared to other features and that the regression
authors, journal, citations) model works well in this situation
Linear regression, sentiment analysis of the publications tweets (posi- A weak positive prediction of high citation counts
tive, negative, neutral) of 6,482,260 tweets, July 2011 to June 2016, user’s across 16 broad disciplines in Scopus, number of
profile, types of journals, citation count, subjects unique Twitter users improved the adjusted
Rsquared value of regression analysis in several
disciplines
Quantile regression, utilized citations to predict publication‘s future Both predictors (i.e., impact factor and early
cisuccess (Impact factor of the publication and the First 1-year citation tations) contribute to the accurate prediction of
counts) are used as predictors long-term citation impact
CART, Naive Bayes, Maximum Entropy Markov, bibliometrics: author, Maximum Entropy Markov model had a better
co-author, venue of publication prediction of the average number of citations
whereas CART performed better for predicting
an average relative increase in citations. They
concluded that an excellent paper will be cited
regardless of the paper’s publishing time and a
high-quality paper will have a high influence
Logistic regression, support vector machine modules, cross-validation, It is feasible to accurately predict future citation
AUC, HITON, and Markov Blanket Algorithm alongside with citation counts with a mixture of content-based and
biblioclassifications, all features, only content features, bibliometric, and metric features using machine learning methods
only the impact factor
Linear regression, 8 years citation window to evaluate the Impact factor,
early citations and compared the correlation of metrics (peer review,
bibliometrics) with the success of scholarly publications
Deep learning CNN prediction models, biblio-features.
CS=Computer Science, E=Engineering, Bio=Biomedical, Lib=Library, Inf=Information, S= Science, M=Mathematics, Re=Rehabilitation,
PM=Physical Medicine, Me=Medical, L=Life,WoS=Web of Science, APS=Applied Physics Statistics, CPS=Computational Science, AM=Applied
Mathematics, P=Physics,Ch=chemical, Mut=Mutiple fields. Doc=Documentation</p>
      <p>Fields
CS
CS
CS,WoS
CS
CS
InfS,LibS</p>
    </sec>
    <sec id="sec-5">
      <title>4. Experiments</title>
      <p>In this section, we describe how we elicited and analyzed the data to achieve our goal. Our
experiment consists of four main phases: (1) Data collection, (2) Feature selection, (3) Classification
predictive models, and (4) Evaluation of the models.</p>
      <sec id="sec-5-1">
        <title>4.1. Data Collection</title>
        <p>2https://service.elsevier.com/app/answers/detail/a_id/14880/supporthub/scopus/
3https://service.elsevier.com/app/answers/detail/a_id/14884/supporthub/scopus/kw/SNIP/
4https://service.elsevier.com/app/answers/detail/a_id/14883/supporthub/scopus/kw/sjr/
5https://plumanalytics.com/
6https://plumanalytics.com/learn/about-metrics/social-media-metrics/
many, we chose tweet count and Facebook count. Mentions count7 is another PlumX category
that includes blog posts, comments, reviews, and Wikipedia links about the publication from
various resources such as Reddit, Slideshare, Vimeo, YouTube, and Github. The three most
important subcategories of Mentions count are news, blog, and reference counts. Moving to
the third category, Capture count8 which tracks users’ actions like bookmarking, marking as
favorite, reading, and exporting the paper. It also includes multiple subcategories such as reader
count which gathers its data from CiteULike, Goodreads, Mendeley, and SSRN, and export/saves
count. The last PlumX metrics category is Usage count9 that points out some usage statistics
like links count, abstract or full-text view count and more.</p>
        <p>The collected papers were ranked in descending order according to their citation counts and
then categorized into 2 categories, the highly cited papers (HCPs) and the Low cited papers
(LCPs) using the following process:
• Calculate the average of all citation counts
• Divide the papers according to their citations compared to the average into the following:
– (HCPs) these papers have a number of citations that were greater than or equal
to the average citation counts of the venue.
– (LCPs) these papers have a number of citations less than the averaged citations
of all papers of that venue.</p>
        <p>The goal of this categorization of the papers into two groups is to characterize the various
stages of paper growth, with HCPs being assigned to the successful ones. Despite the simplicity
of this two-class classifier setting, the evaluation of which indications better forecast the future
success of papers could be clearly captured based on it.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Feature Selection</title>
        <p>The term “feature selection” refers to the process of minimizing the number of input features
that can be utilized to describe the interrelationships between them. It eliminates features
that are redundant or useless. Irrelevant features give no valuable information about the data,
whereas redundant features deliver no additional information than the currently selected
features. In this paper, three diferent feature selection techniques were used to measure the
importance of each feature on the dataset:
Variance Threshold (VAR) a fundamental baseline technique to feature selection is
the VAR approach. It removes features with low variance or those whose variance is less than a
particular threshold. The premise is easy to grasp. Calculate the standard deviation of each
sample feature value and if the number is less than the threshold, filter and then eliminate. By
default, all zero-variance characteristics are turned of. A variance of 0 shows that the sample
feature’s value has remained unchanged.</p>
        <p>Var[] = (1 − )
Pearson’s Correlation (PC) Correlation-based Feature (CFS) is a well-known similarity
measure that evaluates the correlation between features and classes, as well as between features
and other features, to determine the significance of characteristics. In this paper, the
importance of the features’ subset was determined by CFS using Pearson’s correlation equation. The
covariance is cov (X, Y). It can be used for binary classification and regression issues, with a
range of (-1, 1) from a negative to a positive correlation. It’s a fast statistic that ranks features
according to their absolute correlation coeficient with the aim. Between a feature X and the
target Y, the Pearson correlation coeficient is:
  =
cov( ,  )</p>
        <p>( ) 
 ( ;  ) =  ( )– ( | )
Mutual Information Gain (MI) is a metric for measuring how much information one random
variable possesses about another. That is the mutual information between X and Y and may be
thought of as a measure of X’s amount of knowledge of Y (or Y’s amount of knowledge of X).
Therefore, it can be defined as:
Where I (X; Y) represents mutual information between X and Y, H(X) represents entropy for X,
and H (X | Y) represents conditional entropy for X given Y.</p>
      </sec>
      <sec id="sec-5-3">
        <title>4.3. Identification of the Prediction Algorithm</title>
        <p>To evaluate the robustness of the feature subsets created using the three feature selection
techniques, three machine learning algorithms based on classification were applied to the
collected features. We tested the following supervised machine learning methods.</p>
        <p>
          Decision Trees (DT) are trees that classify instances by sorting them based on feature values
[
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. Two entities, decision nodes, and leaves can be used to explain the tree. Each leaf in
a decision tree indicates a value that the node might adopt, whereas each node represents a
feature in an instance to be categorized. From the root to the leaf, a path is traced and sorted
according to feature values. In this study, we have used 5 leaves.
        </p>
        <p>
          Random Forest (RF) is a classifier made up of h(x,k), k=1,..., where k is independently
identically distributed random vectors, and each tree votes for the most popular class at input
x with a single unit vote [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The RF classifier in this paper is made up of eight trees, each of
which was developed using the classification and regression tree (CART) technique. Each case
of a fresh dataset is handed down to each of the eight trees in order to categorize it. The forest
picks the class with the most votes out of eight to be the case’s final class label.
        </p>
        <p>
          Support Vector Machines (SVM) is a technique, a sparse kernel decision machine that
builds its learning model without calculating posterior probabilities. This is a relatively new
supervised machine learning method. According to Gonzalez-Abril et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], SVM conducts
classification by creating an N-dimensional hyperplane that best divides the data into two
groups. It has been demonstrated that increasing the margin and establishing the maximum
feasible distance between the separating hyperplane and the instances on either side reduces
the predicted generalization error.
        </p>
      </sec>
      <sec id="sec-5-4">
        <title>4.4. Evaluation of Classification Models</title>
        <p>We have three classifiers to select from to answer a specific classification issue, therefore, we need
to assess the quality of each (prediction accuracy). To achieve that, we use a confusion matrix
that describes the number of correctly and incorrectly predicted examples by the classification
model. Table 4 depicts the binary classification problem’s confusion matrix, which is a particular
contingency table with two dimensions: actual and predicted. Each metric is a critical indicator
of how well a model performed in relation to a set of criteria. The percentage of valid predictions
correctly categorized by the model is known as model accuracy (Eq. 1):
Accuracy is the fraction of positive results predicted by the model that is really positive (Eq. 2):</p>
        <sec id="sec-5-4-1">
          <title>Model recall is the proportion of relevant outcomes retrieved (Eq. 3):</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Results and Discussion</title>
      <p>The first step was selecting the most promising features using the three diferent feature selection
methods, namely Variance Threshold (VAR), Pearson’s Correlation (PC), and Mutual Information
  (  ) =
 () =</p>
      <p>(  )
(  +   )
(MI). The feature selection was done on the entire dataset. Our results show that there are 9
features that are considered the most significant to reflect the original dataset among the 14
features collected for the prediction task in Table 3. The outcome for the selected feature subset
for each feature selection technique is shown in Table 5.</p>
      <p>The indices Social count(X0), Tweets(X1), News(X4), Capture(X7), and Text view(X13) appear
in all three feature subsets among the nine features in Table 5, indicating that these five features
constitute the dataset’s fundamental characteristics and are the most important representative
of the original dataset. That is, these five features are the ones that play the most important roles
in deciding which papers will become highly cited. Indices Social count(X0), Tweets(X1) show
how many times a paper has been discussed in a certain term and shared on the social network,
this indicates how social media metrics help research dissemination. News(X4) represents the
number of times a paper is referenced in the news media. Moreover, Capture(X7) captures the
interest in the publication on the internet overall. Another useful feature is Text view(X13),
which is the number of times a publication has been seen in full detail.</p>
      <p>The second step was splitting our dataset into test and training sets. We sample our training
set while holding out 30% of the data for testing (evaluating) our classifier. This method can
approximate how well our model will perform on new data.</p>
      <p>The performance of these features in predicting future highly cited papers was then tested
using the three classification models mentioned previously, Decision tree (DT), Random Forest
(RF), and Support Vector Machines (SVM). A code project using Python and its libraries like
SKlearn and numpy was developed to test these models and their performance later using the
performance measures mentioned earlier. Table 6 illustrates the final classification performance
of each feature selection method outcome under each of the three classifiers and the average
classification accuracy (Acc), precision (Prc), and recall (Rcl) are shown in the last row. Obviously,
each classifier has a significant classification performance for each of the feature subsets. The
feature subset picked by Pearson’s correlation (PC) and Mutual information (MI) has obtained
the best precision with 0.97 respectively trained by Random Forest and in terms of precision,
all three feature selection techniques have the same number with 0.96. Whereas the Variance
threshold (VAR) has a maximum recall of 0.99, which was evaluated using Random Forest.
Regardless of the classification model or feature selection approach, the average classification
accuracy is equal to or more than 0.9. Although there is little variation in accuracies, the findings
show that the features derived by the three feature selection approaches are stable and helpful
to classify and forecast future highly cited papers. Furthermore, the results reveal that Random
Forest, in particular, fared best, compared to the Decision Tree and Support Vector Machines.</p>
      <p>The current study’s drawback is that it only looked at 10 journals in the field of computer
science. The findings from this small corpus would not apply to papers in other fields.
Additionally, the current study is exclusively based on PlumX altmetrics correlated with Scopus and
we narrowed our attention to only three machine-learning algorithms. Other algorithms, such
as neural networks, and XGBoost might be investigated in the future. However, the results
serve as a point of reference for future evaluations of prediction-related studies. We provide the
dataset along with one prediction model code as an example for further analysis10.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusion</title>
      <p>In this paper, we build several experiments based on previous research that investigated metrics
and their potential power to predict citation counts. Focusing on the computer science domain
and aiming to find the most promising formula of altmetrics to predict the future success
of a paper measured in the number of citations, we first performed several feature selection
techniques to choose the most important feature subset that better represents the original
dataset. An extensive dataset comprising papers from 10 computer engineering journals (7,588)
was collected, altmetrics and citation counts for each paper were extracted. Furthermore,
altmetrics were evaluated using a feature space with 14 feature indices to determine the most
promising dataset using Variance threshold, Pearson’s correlation, and Mutual information,
and later the classification performance of the feature subsets was verified using three types of
classifiers: Decision tree, Random forest, and Support vector machines. Finally, we evaluated
these prediction models and compared their accuracy. The results show that Random forest
surpasses the other classification methods and we conclude that altmetrics are a valuable
predictor for highly cited papers, specifically these five altmetrics features: social media count,
tweets, news, capture, reader count, and text view.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Abiodun</surname>
            ,
            <given-names>O.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jantan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Omolara</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dada</surname>
            ,
            <given-names>K.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohamed</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arshad</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <year>2018</year>
          .
          <article-title>State-of-the-art in artificial neural network applications: A survey</article-title>
          . Heliyon .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Abramo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>D'Angelo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Felici</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Predicting publication long-term impact through a combination of early citations and journal impact factor</article-title>
          .
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Acuna</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allesina</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kording</surname>
            ,
            <given-names>K.P.</given-names>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Future impact: Predicting scientific success</article-title>
          .
          <source>Nature .</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Akella</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alhoori</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kondamudi</surname>
            ,
            <given-names>P.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freeman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <year>2021</year>
          .
          <article-title>Early indicators of scientific impact: Predicting citations with altmetrics</article-title>
          .
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Aksnes</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langfeldt</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wouters</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Citations, citation indicators, and research quality: An overview of basic concepts and theories</article-title>
          . SAGE Open .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Alchokr</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krüger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shakeel</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saake</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leich</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <year>2022</year>
          .
          <article-title>Peer-reviewing and submission dynamics around top software-engineering venues: A juniors' perspective</article-title>
          , in: International Conference on Evaluation and Assessment in Software Engineering.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Predicting the citations of scholarly paper</article-title>
          .
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Bornmann</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leydesdorf</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <year>2014</year>
          .
          <article-title>How to improve the prediction based on citation impact percentiles for years shortly after the publication date</article-title>
          ?
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <year>2001</year>
          .
          <article-title>Random forests</article-title>
          .
          <source>Machine Learning .</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Carlsson</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <year>2009</year>
          .
          <article-title>Allocation of research funds using bibliometric indicators - asset and challenge to swedish higher education sector</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guestrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <year>2016</year>
          .
          <article-title>Xgboost: A scalable tree boosting system</article-title>
          ,
          <source>in: International Conference on Knowledge Discovery and Data Mining.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Copiello</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>2020</year>
          .
          <article-title>Other than detecting impact in advance, alternative metrics could act as early warning signs of retractions: tentative findings of a study into the papers retracted by plos one</article-title>
          .
          <source>Scientometrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Daud</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Che</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <year>2014</year>
          .
          <article-title>Using machine learning techniques for rising star prediction in co-author network</article-title>
          .
          <source>Scientometrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Edgar</surname>
            ,
            <given-names>T.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manz</surname>
            ,
            <given-names>D.O.</given-names>
          </string-name>
          ,
          <year>2017</year>
          . Machine Learning. Syngress.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Eysenbach</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
          <article-title>Can tweets predict citations? metrics of social impact based on twitter and correlation with traditional metrics of scientific impact</article-title>
          .
          <source>Journal of Medical Internet Research .</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qu</surname>
          </string-name>
          , H., Cheng, Y.,
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            ,
            <surname>Zhan</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ,
          <year>2021</year>
          .
          <article-title>The prediction of asymptomatic carotid atherosclerosis with electronic health records: A comparative study of six machine learning models</article-title>
          .
          <source>BMC Medical Informatics and Decision Making .</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aliferis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <year>2008</year>
          .
          <article-title>Models for predicting and explaining citation count of biomedical articles</article-title>
          .
          <source>AMIA Symposium .</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Galligan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyas-Correia</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>2013</year>
          .
          <article-title>Altmetrics: Rethinking the way we measure</article-title>
          .
          <source>Serials Review .</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Gonzalez-Abril</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angulo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velasco-Morente</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Català</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2005</year>
          .
          <article-title>Unified dual for bi-class SVM approaches</article-title>
          . Pattern Recognition .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Hassan</surname>
            ,
            <given-names>S.U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aljohani</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Idrees</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarwar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawaz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martínez-Cámara</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ventura</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herrera</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <year>2020</year>
          .
          <article-title>Predicting literature's early impact with sentiment analysis in twitter</article-title>
          .
          <source>Knowledge-Based Systems .</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Holden</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenberg</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <year>2005</year>
          .
          <article-title>Tracing thought through time and space: A selective review of bibliometrics in social work</article-title>
          .
          <source>Social Work in Health Care .</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>A deep learning methodology for citation count prediction with large-scale biblio-features.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Lutz</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2014</year>
          .
          <article-title>Do altmetrics point to the broader impact of research? an overview of benefits and disadvantages of altmetrics</article-title>
          .
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Nuzzolese</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciancarini</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gangemi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peroni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poggi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Presutti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Do altmetrics work for assessing research quality? Scientometrics .</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Patro</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aggarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
          <article-title>How honest is the h-index in measuring individual research output? Journal of postgraduate medicine</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Penner</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>R.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petersen</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fortunato</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>2013</year>
          .
          <article-title>On the predictability of future impact in science</article-title>
          .
          <source>Scientific reports 3</source>
          ,
          <fpage>3052</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Poggi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciancarini</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gangemi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nuzzolese</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peroni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Presutti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Predicting the results of evaluation procedures of academics</article-title>
          . PeerJ Computer Science .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Quinlan</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <year>1986</year>
          .
          <article-title>Induction of decision trees</article-title>
          .
          <source>Machine Learning .</source>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Ruan</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
          </string-name>
          , J., Cheng, Y.,
          <year>2020</year>
          .
          <article-title>Predicting the citation counts of individual papers via a BP neural network</article-title>
          .
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Shakeel</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alchokr</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krüger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leich</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saake</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <article-title>2022a. Altmetrics and citation counts: An empirical analysis of the computer science domain</article-title>
          , in: Joint Conference on Digital Libraries.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Shakeel</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alchokr</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krüger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leich</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saake</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <year>2022b</year>
          .
          <article-title>Are altmetrics useful for assessing scientific impact? a survey</article-title>
          ,
          <source>in: International Conference on Management of Digital EcoSystems.</source>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Shakeel</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alchokr</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krüger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leich</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saake</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <year>2022c</year>
          .
          <article-title>Incorporating altmetrics to support selection and assessment of publications during literature analyses</article-title>
          ,
          <source>in: International Conference on Evaluation and Assessment in Software Engineering.</source>
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Shakeel</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alchokr</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krüger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saake</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leich</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <year>2021</year>
          .
          <article-title>Are altmetrics proxies or complements to citations for assessing impact in computer science?</article-title>
          , in: Joint Conference on Digital Libraries.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Siler</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bero</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Measuring the efectiveness of scientific gatekeeping</article-title>
          .
          <source>Proceedings of the National Academy of Sciences .</source>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Stegehuis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Litvak</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waltman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Predicting the long-term citation impact of recent publications</article-title>
          .
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Thelwall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <year>2020</year>
          .
          <article-title>The pros and cons of the use of altmetrics in research assessment</article-title>
          .
          <source>Scholarly Assessment Reports .</source>
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Thelwall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nevill</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <year>2018</year>
          .
          <article-title>Could scientists use altmetric.com scores to predict longer term citation counts</article-title>
          ?
          <source>Journal of Informetrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabási</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <year>2013</year>
          .
          <article-title>Quantifying long-term scientific impact</article-title>
          .
          <source>Science .</source>
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Which can better predict the future success of articles? Bibliometric indices or alternative metrics</article-title>
          .
          <source>Scientometrics .</source>
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>P.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <year>2014</year>
          .
          <article-title>Citation impact prediction for scientific papers using stepwise regression analysis</article-title>
          .
          <source>Scientometrics</source>
          <volume>101</volume>
          ,
          <fpage>1233</fpage>
          -
          <lpage>1252</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>