<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Changchun, China and Online
∗ Corresponding author.
zhouhc@clas.ac.cn (H. Zhou); huangxiaorong@clas.ac.cn
(X. Huang); puhj@clas.ac.cn (H. Pu) zhang@smail.nju.edu.cn (Q. Zhang)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>May Generative AI Be a Reviewer on an Academic Paper?⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haichen Zhou</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiaorong Huang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hongjun Pu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qi Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nanjing University</institution>
          ,
          <addr-line>Nanjing, 210023</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Science Library (Chengdu), Chinese Academy of Sciences</institution>
          ,
          <addr-line>Chengdu, 610041</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The application of artificial intelligence (AI) to academic evaluation is one of the important topics within the academic community. The widespread adoption of technologies such as Generative AI (GenAI) and Large Language Models appears to have introduced new opportunities for academic evaluation. The question of whether GenAI has the capability to perform academic evaluations, and what differences exist between its abilities and those of human experts, becomes the primary issue that needs to be addressed first. In this study, we have developed a set of evaluation criteria and processes to investigate on 853 post peer-reviewed papers in the field of cell biology, aiming to observe the differences in scoring and comment styles between GenAI and human experts. We found that the scores given by GenAI tend to be higher than those given by experts, and the evaluation texts lack substantive content. The results indicate that GenAI is currently unable to provide the depth of understanding and subtle analysis provided by human experts.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;academic evaluation</kwd>
        <kwd>Generative AI</kwd>
        <kwd>large language models</kwd>
        <kwd>Copilot</kwd>
        <kwd>ChatGPT 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>How to use AI for more objective, accurate, and efficient
academic evaluation has become an important research
topic[1][2]. Generative AI (GenAI) is a novel technology
that uses artificial intelligence to generate content in
various forms[3][4]. In the context of academic evaluation,
GenAI provides a new possibility for automating
academic evaluations by generating academic
evaluation content[ 5 ]. Comparing the evaluations of
human experts and those generated by GenAI is a very
intuitive way to better understand the effectiveness and
reliability of GenAI. However, there is still a lack of
research on the quality of the content generated by
GenAI and whether there are differences between it and
the content generated by human experts. Figuring out
these issues provides a basis for us to answer whether
GenAI can match the depth of understanding and subtle
analysis provided by human experts, what areas GenAI
excels in, and where it may need further improvement.
Hence, we focus on analyzing the difference between
human expert evaluations and GenAI evaluations. We
aim to answer the following research questions:
RQ1: Can GenAI conduct academic evaluations?
RQ2: What differences exist between the scoring
results of GenAI and human experts?
RQ3: What differences exist between the
evaluation text features of GenAI and human
experts?</p>
    </sec>
    <sec id="sec-2">
      <title>2. Method</title>
      <p>Our research methodology includes the following steps:
1.
2.</p>
      <p>Select papers from H1 Connect (connect.h1.co)
as research cases and establish selection
criteria.</p>
      <p>Generate a list of papers to be collected.</p>
      <p>Obtain data such as Paper Title, DOI, Expert
Score, Review Text, etc., to form the original
dataset.</p>
      <p>Design evaluation dimensions and scoring
system for the research field of our dataset.</p>
      <p>Generate the Copilot question template
(Prompt).</p>
      <p>Use the template to ask questions and collected
Copilot scores and evaluation text data.</p>
      <p>Compare the differences in scores and texts
between Copilot and experts.</p>
      <sec id="sec-2-1">
        <title>2.1. Data Preparation</title>
        <p>To minimize the influence of various factors on the
evaluation results, such as the differences in evaluation
standards for papers in different fields, newly published
papers not yet receiving sufficient attention, and
differences in evaluation preferences among different
experts, we have limited the research field to Cell
Biology. We focused on papers from cell biology
published in 2020 that received one evaluation. We
collected data on 853 papers (as of May 2022) from H1
Connect. H1 Connect is a leading platform for
researchers and clinicians seeking expert opinions and
insights on the latest life sciences and medical research.
We collected key information about the papers,
including paper title, authors, journal, DOI, PMID,
recommended score etc.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Question Template Design</title>
        <p>We designed an evaluation system specifically for the
field of Cell Biology to enhance the relevance and
reliability of the content generated by Copilot. By
summarizing the review principles of top journals in this
field, such as “Nature Reviews Molecular Cell Biology”,
“Trends in Cell Biology”, “The Journal of Cell Biology”,
“Nature Cell Biology”, and “Journal of Molecular Cell
Biology”, we extracted the following evaluation
dimensions. Copilot will be required to evaluate each
paper on these dimensions, provide a recommendation
score, and finally give a comprehensive evaluation.</p>
        <p>After several rounds of testing, the final template
(prompt) for querying Copilot has been established as
follows, with the inclusion of PubMed ID to assist
Copilot in accurately targeting information on the
internet:</p>
        <p>I have summarized a set of criteria for evaluating
academic papers:</p>
        <p>•Originality: The paper must report novel,
innovative and influential research that does not repeat or
plagiarize existing work.</p>
        <p>•Accuracy: The paper must follow high standards of
experimental design, data analysis and result presentation,
without errors, biases or misleading.</p>
        <p>•Conceptual advance: The paper must provide a
deep understanding and mechanistic explanation of an
important problem or area, not just superficial or
incremental improvements.</p>
        <p>•Timeliness: The paper must reflect the current hot
topics in the scientific community.</p>
        <p>•Significance: The paper must have immediate or
long-term impact and implications.</p>
        <p>I also have a recommended scoring system: 1 star
(Good), 2 stars (Very Good) , 3 stars (Exceptional) .You're
acting as a scientist. I'll give you a PubMed ID for the paper.
First, please display the title of the paper and search the
web site. No abstract is required. Second, please according
to my criteria and scoring system for evaluation and
scoring the paper; Third, please according to my scoring
system for the overall evaluation and scoring of the paper.
Pubmed ID:XXXXXXXX</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Collection of Evaluation Results</title>
        <p>The process of collecting Copilot evaluation results is
shown in Figure 1:</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Result and Discussion</title>
      <p>RQ1: Can GenAI conduct academic evaluations?
GenAI can conduct academic evaluations and produce
readable results in form.</p>
      <p>RQ2: What differences exist between the scoring
results of GenAI and human experts?
Observing the score distribution ratio (Figure 2), papers
scored 3 stars by experts only account for 15%, while
those scored 2 stars and 1 star are both around 40%.
However, Copilot’s 3 stars evaluations account for more
than 60%, 2 stars account for 32.72%, and 1 star is less
than 1%.</p>
      <p>Comparing the scoring results of Copilot and expert
(Figure 3), we observe that: Copilot’s scoring results are
higher, indicating that it tends to give higher scores to
most papers. The average score given by Copilot is 2.68
stars, while the average score given by experts is 1.76
stars. This is consistent with the experimental results of
Mike Thelwall on 51 papers[2]. The fact that over 60% of
papers are scored 3 stars by Copilot suggests that it may
not yet possess the core ability to accurately distinguish
high-value academic papers.</p>
      <p>RQ3: What differences exist between the
evaluation text features of GenAI and human
experts?
From the perspective of sentences, both the number of
sentences and the average sentence length in the Copilot
text are less than/shorter than those in the expert text,
but the difference is not significant (Figure 5).</p>
      <p>From a lexical perspective, the overall proportion of
word types between the two is not significantly different,
with Copilot tending to use more adjectives (Figure 6
and Appendix A). The high-frequency words used by
experts better reflect professionalism and specificity,
such as “cell”, “protein”, and “cancer”. In contrast, the
high-frequency words used by Copilot are more general,
such as “significant” (Table 1).</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>In this study, we collected post peer-review scores and
text data for 853 papers in the field of Cell Biology from
the H1 Connect website. We also obtained the scores
and evaluation text data for each of these papers from
Copilot, based on our designed evaluation criteria and
process. By comparing the evaluation results of Copilot
and experts through quantitative analysis and text
mining methods, we found that Copilot can score and
evaluate under specific prompts. From the scoring
perspective, there is a significant difference in the
scoring patterns between Copilot and experts: the
former tends to give higher star, but the high proportion
of 3-star reveals that it does not have enough ability to
judge the actual value of the paper. From the text
perspective, Copilot’s shorter sentences and generic
wording indicate that its evaluation is only at the stage
of imitating the features of the evaluation text, and it
cannot yet carry out substantive evaluations of the
originality, accuracy, and other core elements of the
paper.</p>
      <p>Overall, GenAI, represented by Copilot, is currently
unable to provide the depth of understanding and subtle
analysis provided by human experts. It should still not
be used for academic evaluation at this stage, as its
overevaluative nature may lead to the proliferation of
lowquality academic results[6]. This study presents several
limitations: Firstly, a disparity exists between experts
and Copilot, with the latter unable to access complete
paper texts, unlike experts. Secondly, the author did not
perform iterative testing nor utilized the mean outcomes
of various expert assessments for analysis. Third, the
evaluation criteria of Copilot and experts are not
consistent. The limitations bear potential inaccuracies
for the research outcomes. Consequently, we aim to
address these deficiencies in the subsequent phase of
analysis.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work was supported by the key project of
innovation fund from National Science Library
(Chengdu), the Chinese Academy of Sciences
(E3Z0000902). We sincerely appreciate the insightful
comments and constructive suggestions provided by the
reviewers, which have significantly contributed to the
improvement of our manuscript.</p>
    </sec>
    <sec id="sec-6">
      <title>Appendix</title>
    </sec>
    <sec id="sec-7">
      <title>A. Abbreviations and Examples of SpaCy Parts-of-Speech</title>
      <p>







</p>
      <p>ADJ--adjective *big, old, green*
ADP--adposition *in, to, during*
ADV--adverb *very, tomorrow, where *
AUX--auxiliary *is, has (done), will (do) *
CCONJ--coordinating conjunction *and, or,
but*
DET--determiner *a, an, the*
NOUN--noun *girl, cat, tree, air, beauty*
PUNCT--punctuation *., (, ), ?*
VERB--verb *run, runs, running, eat, ate,
eating*</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>W.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cao</surname>
          </string-name>
          , et al,
          <article-title>Can large language models provide useful feedback on research papers? A large-scale empirical analysis</article-title>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2310.01783.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Thelwall</surname>
          </string-name>
          , Can ChatGPT evaluate research quality?,
          <year>2024</year>
          . URL: http://arxiv.org/abs/2402.05519.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bloomberg</surname>
          </string-name>
          ,
          <article-title>Generative AI to Become a $1.3 Trillion Market by 2032</article-title>
          , Research Finds,
          <year>2023</year>
          . RUL: https://www.bloomberg.com/company/press/gen erative-ai
          <article-title>-to-become-a-1-3-trillion-market-by2032-research-finds/ Gartner, Understand and Exploit GenAI with Gartner's New Impact Radar</article-title>
          ,
          <year>2024</year>
          . URL: https://www.gartner.com/en/articles/understandand-exploit
          <article-title>-gen-ai-with-gartner-s-new-impactradar.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>J. de Winter</surname>
          </string-name>
          ,
          <article-title>Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstracts</article-title>
          ,
          <source>J. Scientometrics</source>
          . (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>M. B. Garcia</surname>
          </string-name>
          ,
          <article-title>Using AI tools in writing peer review reports: should academic journals embrace the use of ChatGPT ?</article-title>
          ,
          <source>Annals of biomedical engineering 52</source>
          ,
          <year>2024</year>
          :
          <fpage>139</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>