<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Relationship between Team Diversity and Innovation Performance in Interdisciplinary Research Teams within the Field of Artificial Intelligence: Decision Tree Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Junwan Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chenchen Huang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shuo Xu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>College of Economics and Management, Beijing University of Technology</institution>
          ,
          <addr-line>Beijing, 100124</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Interdisciplinary research teams are crucial in solving complex problems by providing creative solutions that singlediscipline teams cannot achieve. Previous studies have primarily focused on the linear relationship between independent variables and team innovation performance, neglecting the non-linear aspect. To address this gap, this paper examines the non-linear relationship between diverse factors and the innovation performance of interdisciplinary research teams in artificial intelligence. By utilizing the Classification and Regression Tree (CART) model, the study reveals that activity diversity and interdisciplinary research team innovation performance exhibit a U-shaped relationship in terms of “novelty” innovation performance. Furthermore, this relationship is influenced by research interest diversity. Specifically, low research interest diversity leads to low innovation performance as activity diversity increases. Meanwhile, research interest diversity emerges as the most critical factor impacting innovation performance. The importance of member diversity, institutional diversity, and activity diversity on innovation performance should not be ignored. Through decision tree analysis, this paper extends research on the multifactor combination, complex nonlinear relationships, and multipath influence mechanism of team diversity on interdisciplinary research teams' innovation performance.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Interdisciplinary research team</kwd>
        <kwd>Team diversity</kwd>
        <kwd>Innovation performance</kwd>
        <kwd>Classification and regression tree (CART) model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Growing globalization and intense market competition have
transformed the scientific model and increased the number of
specialized research teams. In order to enhance the efficiency and
quality of scientific research, more research teams are
transitioning from single teams to diversified teams[1].
Interdisciplinary research has become a necessary choice in this
context[2], allowing teams to draw upon a wide range of
disciplines and expertise to address complex scientific questions.
The diversity of team members from various disciplines and
backgrounds contributes to greater knowledge and innovative
results[3]. Thus, effective interdisciplinary collaboration is crucial
for achieving scientific and innovative breakthroughs[4].</p>
      <p>Previous research highlights the significance of diversity as a
crucial factor influencing the success of interdisciplinary research
teams[5]. Diversity can be broadly categorized as demographic
diversity and task-related diversity[6]. Demographic diversity
encompasses variations in team members’ demographic attributes,
such as age, gender, and institutional backgrounds[7].
Taskrelated diversity pertains to the diverse qualities that team
members bring to their academic or professional pursuits,
including workplace functions, knowledge, and education[8].
Creating successful teams with demographic and task-related
diversity is not a straightforward process of simply combining
individuals from different disciplines. Horwitz et al.[6] discovered
that while demographic diversity did not significantly impact team
performance, task-related diversity positively influenced it.
Diverse teams struggle with issues such as gender differences,
team conflict, and collaboration[9]. Given these contradictory
findings, our focus is on investigating the impact of both
demographic diversity and task-related diversity on the innovation
performance of interdisciplinary research teams, while analyzing
the varying importance of different diversity factors. Existing
studies have primarily focused on exploring the linear relationship
between independent variables and team innovation performance,
overlooking the nonlinear aspect. The nonlinear relationship
between these characteristics and team innovation performance,
especially in the context of demographic diversity and task-related
diversity in interdisciplinary research teams, remains unclear.</p>
      <p>This paper aims to investigate the impact and decision-making
mechanisms of team diversity on the innovation performance of
interdisciplinary research teams. Firstly, team innovation
performance is divided into novelty and impact[10]. Second, we
will investigate the influence of team diversity on the
interdisciplinary research team’s innovation performance in terms
of both the demographic diversity and task-related diversity of
team members. From a social categorization perspective, we
assume that gender diversity, national diversity, and institutional
diversity are included in the demographic diversity in this context.
Meanwhile, relying on the informational decision-making
perspective, we hypothesize that task-related diversity includes
sociability diversity, activity diversity, research interest diversity,
and member diversity. Specifically, we address the following
research questions in this paper: RQ1: What is the complex
relationship structure among demographic diversity, task-related
diversity, and the innovation performance of interdisciplinary
research teams? RQ2: What combinations of characteristics
promote high levels of team innovation performance? RQ3: What
diversity characteristics should researchers focus on to enhance
the innovation performance of interdisciplinary research teams?
2.</p>
    </sec>
    <sec id="sec-2">
      <title>Data and methods</title>
      <p>This paper investigates the impact of team diversity on
interdisciplinary research teams’ innovation performance. The
basic process is shown in Figure 1. Firstly, raw data is processed
to form authors’ collaborative relationship data and measure each
author’s collaborative tie strength. Second, stable collaborative
relationships are identified using a pre-set threshold (super tie).
Then, members of interdisciplinary research teams are identified,
and the diversity index of each team is measured. Again, team
innovation performance is divided into novelty and impact to be
measured. Finally, by using team diversity as the conditional
attribute and innovation performance as the decision attribute, the
impact of team diversity on the interdisciplinary research team’s
innovation performance is explored using the CART model.
Data Processing
Paper publication
data</p>
      <p>Author
collaborative
relationship</p>
      <p>Collaborative
Networks
Team</p>
      <p>Diversity
gender
diversity
activity
diversity
institutional national
diversity diversity
research interest diversity
classDifieccaitsiioonnrteresuelt x1
classDifieccaitsiioonnrteresuelt x2</p>
      <p>······
······</p>
      <p>This paper focuses on empirical research in the field of
artificial intelligence (AI). The dataset is derived from the
information of the most influential scholar award winners on the
AMiner website (https://www.aminer.cn/ai2000) 2023 AI 2000
annual list. There are three reasons for selecting these scholars as
the subjects of the study. Firstly, AI research is inherently
interdisciplinary[11]. Second, since its launch in 2006, the
AMiner platform has already been used by many researchers[12].
Third, since 2017, the AMiner platform has been publishing the
annual AI 2000 most influential scholar list. The purpose of this
list is to annually rank the 2000 scholars who are expected to be
highly cited in the field of AI over the next ten years (2020-2029).
Firstly, we obtain the dataset from the AMiner website, which
covers information on the most influential scholars in the 2023 AI
2000 annual list. The dataset consists of 195 selected scholars and
their collaborators, with five of these scholars receiving awards in
two or more subfields. The public information of the selected
scholars is obtained from their personal websites and academic
social networking platforms. The papers of selected scholars are
downloaded from the Web of Science database. Finally, 25,285
papers are submitted by 195 selected scholars.</p>
      <p>Second, we conduct community detection on the author
cooccurrence network for each selected scholar using the Louvain
algorithm. This algorithm divides the nodes in the network into
different communities based on modularity metrics, which assess
the collaborative relationships between the nodes. The algorithm
identifies strong connections within the same community and
sparser connections between different communities. In the
network, each co-author is represented as a node, and edges
represent collaborations between selected scholars and co-authors
who have published papers together. Figure 2 shows the author
co-occurrence network for Silver, D selected scholars. Different
colors in the figure represent various communities determined by
modularity, with the community to which Silver and other
selected scholars belong shown in purple.</p>
      <p>Next, we calculate the collaborative tie strength among the
nodes in each community to filter out the core collaborators of
each selected scholar. Collaborative tie strength, also known as a
“super tie”, has been extensively studied in scientific collaborative
networks[13]. It represents a long-term and stable collaboration,
similar to life partners, characterized by high intensity, close ties,
and long durations[14]. To identify the core collaborators among
the 195 selected scholars, we calculate the super tie for each
community. Specifically, when a member’s collaborative tie
strength exceeds his or her community’s super tie threshold, that
member is referred to as a super tie collaborator and core team
member. Equations 1 and 2 demonstrate the formula for
calculating the super tie[14]. Additionally, we employ the
Anderson-Darling test to examine the distribution of collaborative
intensity   /  among the members. Our analysis indicates that
the statistical distribution  (  ) of all members’ collaborative
intensity conforms to an exponential distribution, with the average
collaborative strength of members being 2.83.</p>
      <p>
        〈  〉 =  −1 ∑ =1   (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
   = (〈  〉 − 1) ln   (2)
      </p>
      <p>Where the collaborative tie strength   is defined as the
cumulative number of papers co-authored by the selected scholar
in community  and scholar  over the time between their first and
last paper.   represents the number of different co-authors of
selected scholars in the community  . 〈  〉 represents the average
collaborative tie strength   . Each scholar  with   &gt;    is
labelled as a super tie collaborator of community  .</p>
      <p>Eventually, we identify 195 research teams and their 1,217
core members. Figure 3 shows the distribution of team sizes for
195 teams. The largest team size is 57 members, and 165 teams
are smaller than 10 members, which represents 85% of all teams.
Previous research defines interdisciplinary teams as groups of
scientists from different disciplines who collaborate to address
complex problems[15]. To verify the interdisciplinarity of these
teams, we utilize a method that maps member affiliations to
disciplinary classifications[16] to more accurately determine the
disciplinary backgrounds of the members. Specifically, we extract
secondary institutions from each member’s address, retain the
disciplinary terms in the secondary institution names, and match
these terms to the discipline field in the OECD classification
scheme. In this way, each member’s institution can be precisely
matched to his or her research discipline. The results indicate that
165 teams have members from two different disciplinary
backgrounds, 27 teams have members from three different
disciplinary backgrounds, and 3 teams have members from four
different disciplinary backgrounds, thus reinforcing that the 195
teams in this study are interdisciplinary research teams. Table 1
demonstrates the distribution of members from different
disciplinary fields. Specifically, 71.18% of the members are from
the field of computer and information science, 21.80% are from
the fields of electrical engineering, electronic engineering, and
information engineering, while other fields encompass
environmental engineering, nanotechnology, and physical
sciences, among others. To explore the factors influencing the
interdisciplinary research teams’ innovation performance, we
download each team member’s papers from the Web of Science
database and collect 91,025 papers from all teams.
We calculate the degree of team novelty using the novelty index
proposed by Lee et al.[10]. This index measures the novelty of a
team’s paper based on the rarity of prior citation pairs. The
calculation involves two steps.</p>
      <p>
        Complete the first step of the operation on the paper level. (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
List all paired reference combinations for each paper. (2) Record
the corresponding journal pairs. (3) Aggregate the pairs of journal
combinations published from year t-2 to year t as   set. The time
window from year t-2 to year t is chosen to ensure data robustness.
We then calculate the commonness value using Equation 3[10]. (4)
This equation assigns each paper a range of commonness values.
The commonness values of each paper are ranked, and the 10th
percentile is taken as the commonness value of the paper. Using
the 10th percentile instead of the minimum value helps reduce
noise and increase the reliability of the measure. (5) The
commonness value is transformed using a natural logarithm to
obtain an approximately normally distributed variable. The final
novelty value for that paper is obtained by adding a negative sign.
      </p>
      <p>Commonnes s =      =     ×      ×  =
journal  .   is the number of journal pairs in   set that contain
journal  , and   is the number of all journal pairs in   set.</p>
      <p>Calculating novelty at the team level is the next stage. The
number of paper publications by each team is counted and divided
by the team size to calculate the team’s novelty value.</p>
      <p>
        We measure a team’s impact using forward citations[10].
High-impact papers are defined as those in the top 1% of citation
distribution. This definition follows Uzzi et al.[17] and considers
the use of short citation time windows can lead to the incorrect
identification of highly cited papers[18]. First, the process of
identifying high-impact papers is completed at the paper level. (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
Rank all papers from highest to lowest citation count. (2) By using
a five-year moving window, we define papers in the top 1% of the
rankings from year t-5 to year t as high-impact papers. (3) We use
a dummy variable to indicate whether each paper is a high-impact
publication, assigning a value of 1 if it is and 0 otherwise. Next,
the number of high-impact papers per team is counted and divided
by the team size to finally obtain the team’s impact value.
2.3.2
      </p>
      <sec id="sec-2-1">
        <title>Independent variables</title>
        <p>Gender diversity refers to the subjective or objective similarities
and differences between team members in terms of gender[19].
National diversity is defined as having team
members from
different national backgrounds, which introduces sociological
categorization
and
the
potential
for
diverse
cognitive
perspectives[20]. Institutional diversity refers to the presence of a
variety of members in different institutions[21]. All of the above
demographic diversity indicators mentioned above are measured
using the Simpson index, which is calculated using Equation 4.
 = 1 − ∑ =1</p>
        <p>2</p>
        <p>Where  is the total number of categories,   is the percentage
of members of the group  . The higher the  value, the greater is
the value of the diversity.</p>
        <p>Using the AMiner platform, we algorithmically obtain data on
sociability, activity, and research interest diversity indices to
assess the academic proficiency of members. Definitions and
formulas for these indicators are provided[22].</p>
        <p>The sociability index is derived from considering both the
number of scholars’ collaborators and their collaborative papers,</p>
        <p>Where, in  years ( belongs to near  years),   is a group of
papers published by scholars in n years, weight(n) =  this year−n,
and the following principles are applied to the values of  and  :
if the current month is in the first half of the year (month&lt;July), it
is set</p>
        <p>= 4 and  = 0.75; if the present month is at the second
half, it is set  = 3 and  = 0.85.
as shown in Equation 5.</p>
        <p />
        <p>Where the #
as shown in Equation 6.</p>
        <p>activity(A) = ∑</p>
        <p>ℎ 
( ) = 1 + ∑
ℎ 
ℎ ( )   (#
 ) (5)
 is the number of papers co-authored
between scholar and co-authors.</p>
        <p>The activity index measures a scholar’s frequency and number
of recent publications, along with the significance of each paper,
( )</p>
        <p>IS(  ) × weight(n)
(4)
(6)</p>
        <p>To assess the diversity and differences in a team’s overall
sociability and activity, we utilize Equation 4 to calculate the
diversity of these two evaluation indicators.</p>
        <p>The research interest diversity of scholar is based on the
breadth of the field of interest. Using the topic model, we identify
each scholar’s field of study and assign their papers to relevant
topics. The   (t) topic distribution is obtained by Equation 7, and
the research interest diversity is defined as the threshold of the
distribution of the   ( ), which is calculated by Equation 8.
  (t) = #
  
#
research interest diversity; MD = member diversity.
The CART model is a supervised machine learning model
proposed by Breiman[24]. It is a classification regression method
generated based on the regression of the fork decision properties.
It is commonly used in data analysis and evaluation, including
project performance assessment[25]. When studying the factors
that influence the innovation performance of interdisciplinary
research teams, we often face the challenge of dealing with
multivariate and nonlinear relationships. Traditional regression
models, while useful in revealing relationships between
independent and dependent variables, may have limitations in
handling complex relationships. These models rely on the least
squares method, which requires rigorous hypothesis testing and
variable control[26]. In the context of understanding how
diversity impacts the teams’ innovation performance, it is crucial
to consider multiple independent variables and potential
interactions among them. Incorrect selection of control variables
or omission of important variables may lead to biased results in
regression analysis. In contrast, the CART model offers a more
flexible and robust solution[27]. It constructs a decision tree using
recursive binary splitting, dividing data subsets into smaller
subsets[24]. This nonparametric approach eliminates the need for
rigorous hypothesis testing and variable control, as it
automatically selects partitioning rules based on the actual
distribution and characteristics of the data[28]. As a result, the
CART model can capture nonlinear relationships and higher-order
interactions[29], providing a more accurate understanding of the
impact of diversity on team innovation performance.
2.5.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Design of the CART Decision Tree</title>
        <p>The paper categorizes the data based on innovation performance
and divides it into training and test samples in an 8:2 ratio.
Alongside the CART model, several baseline models are trained
on the dataset, and their performance on the test set is compared
to select the optimal model. These baseline models include the
C4.5 model, CART model, Random Forest model, and Gradient
Boosting Tree model. The C4.5 model is similar to the CART
model but uses information entropy as the partitioning criterion.
The random forest model combines multiple decision trees to
optimize classification, while the gradient boosting tree model
iteratively trains weak classifiers and combines them into strong
classifiers. To enhance the model’s performance, the paper
utilizes the grid search method to systematically explore various
hyperparameter combinations and determine the optimal
parameter configurations. The accuracy of each model, after grid
search, exceeds 0.6, indicating that over 60% of the samples are
correctly predicted by the model. Among the models, the CART
model demonstrates superior performance with accuracy rates of
0.73 and 0.68 in measuring the novelty and impact of the team,
respectively. Based on these results, the CART model is selected
for use in this paper.
3.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion</title>
      <sec id="sec-3-1">
        <title>Model Result Analysis</title>
        <p>Through the CART model, a decision rule table for innovation
performance can be constructed between on the variables of
“novelty” and “impact”. Table 4 shows that a high decision result
indicates that the novelty and impact of the team with the current
decision rule are higher than the median novelty and impact of all
interdisciplinary research teams, respectively.</p>
        <p>Table 4</p>
        <p>Innovation Performance Decision Rules
Novelty
“impact”, there is a higher proportion of interdisciplinary teams
with a low innovation performance rating than in the “novelty”
innovation performance. Secondly, activity diversity and member
diversity split as root nodes, which are key factors affecting the
novelty and impact of a team, respectively. Third, the confidence
coefficients for most of the decision rules are above 60%,
indicating that the weight of the sample size supporting the
current decision rule in the leaf node’s sample size is 60% or more.
This suggests that the results are highly interpretable.</p>
        <p>Figure 6 shows that there are a total of three rules to determine
whether a team has high or low innovation performance. The
CART model divides data into approximately two branches based
on whether the most important feature “activity diversity” is less
than or equal to -0.40. The results show that with lower activity
diversity, team members can focus more on their original thinking,
enhancing team innovation performance without the need to be
concerned about publication frequency and quantity. When
activity diversity is higher, an increase in research interest
diversity contributes to teams achieving high levels of innovation
performance. Horizontal comparison reveals that activity diversity
is a crucial factor influencing innovation performance.
Interdisciplinary research teams with low activity diversity are not
influenced by research interest diversity, whereas interdisciplinary
research teams with high activity diversity are impacted by
research interest diversity.</p>
        <p>≤</p>
        <p>&gt;
to prioritize the diversity of institutions represented within the
teams.</p>
        <p>≤</p>
        <p>&gt;
The model’s final characteristic importance for explanatory
variables is presented in Figure 8. Among the factors affecting a
team’s novelty, research interest diversity has the highest
characteristic importance of 0.97, while activity diversity has a
characteristic importance of 0.03. Among the factors affecting a
team’s impact, research interest diversity has the highest
characteristic importance of 0.48. Member diversity and
institutional diversity closely followed with 0.31 and 0.22,
respectively. This result shows that research interest diversity is
most strongly associated with interdisciplinary research teams'
innovation performance. This shows that team members' diverse
expertise backgrounds enable knowledge integration and
reconfiguration, which are crucial elements in innovation
performance[21].</p>
        <p>AD
MD
SD
RID
ND
ID
GD
This paper finds a U-shaped relationship between activity
diversity and team innovation performance in “novelty”
innovation performance. However, this relationship is impacted
by research interests diversity. Specifically, interdisciplinary
Novelty
Impact
0.97
teams with low activity diversity are able to improve their
innovation performance independently of research interest
diversity. In contrast, low research interest diversity leads to low
innovation performance when activity diversity increases. In
terms of “impact” innovation performance, increasing member
diversity and managing the range of research interests can be
beneficial. In addition, interdisciplinary research teams with low
member diversity need to focus on the institutional diversity of
team members, as institutional diversity has a positive impact on
the team’s effectiveness.</p>
        <p>In the evaluation of various factors, research interest diversity
emerges as the most significant determinant of innovation
performance in interdisciplinary research teams. This implies a
close association between research interest diversity and the
team’s ability to innovate. Researchers with distinct research
themes bring a diverse range of knowledge and contribute to the
reconfiguration of knowledge by identifying and integrating
insights from different fields[30].</p>
        <p>Managers should consider research team diversity when
developing it. To create a healthy innovation environment, they
should pay attention to the heterogeneity of different
organizational and disciplinary backgrounds to which team
members belong and strive to optimize the level of knowledge
diversity in the team. Furthermore, managers should encourage
and promote activity diversity. Meanwhile, the costs of too much
research interest diversity need to be noted to avoid the
phenomenon of “too much of a good thing”. This is because
knowledge diversity among collaborative members in different
research fields is often considered a double-edged sword. As
Wang et al have found, increasing knowledge diversity leads to a
decline in social influence after a certain peak[21].</p>
        <p>Our study has limitations. Firstly, it focuses solely on the
AMiner platform in AI, limiting its scope and generalizability.
Future studies should broaden the research fields. Secondly, the
sample of 195 interdisciplinary teams may not fully reflect AI
team diversity and complexity, potentially suffering from
sampling error. A more representative sample is needed. Thirdly,
while we focused on diversity within teams, future research could
explore diversity in other research team activities. Lastly, our
team recognition method overlooks member turnover dynamics.
Future studies should introduce a dynamic analysis of member
flow for more accurate core member identification.</p>
      </sec>
      <sec id="sec-3-2">
        <title>ACKNOWLEDGMENTS</title>
        <p>This work was supported partially by the National Natural
Science Foundation of China (Grant No. 72174016). Our gratitude
also goes to the anonymous reviewers and the editor for their
valuable comments.
[24]
[25]</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] [2] [3]
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
          </string-name>
          , H.,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Collaboration diversity and scientific impact</article-title>
          .
          <source>arXiv preprint arXiv</source>
          ,
          <year>1806</year>
          .
          <volume>03694</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Van Noorden</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Interdisciplinary research by the numbers</article-title>
          .
          <source>Nature</source>
          ,
          <volume>525</volume>
          (
          <issue>7569</issue>
          ),
          <fpage>306</fpage>
          -
          <lpage>307</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>Greve</surname>
            ,
            <given-names>H. R.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Superman or the fantastic four? Knowledge combination and experience in innovative teams</article-title>
          .
          <source>Academy of management journal</source>
          ,
          <volume>49</volume>
          (
          <issue>4</issue>
          ),
          <fpage>723</fpage>
          -
          <lpage>740</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>