<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>How to Measure Information Cocoon in Academic Environment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jia Yuan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guoxiu He</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yunhan Yang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Education, The University of Hong Kong</institution>
          ,
          <addr-line>Hong Kong, SAR</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Economics and Management, East China Normal University</institution>
          ,
          <addr-line>Shanghai</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>When individuals face an abundance of information, they often selectively choose data that reinforces their existing beliefs, ignoring opposing views and creating an 'information cocoon'. This phenomenon is not limited to social media; it is also relevant in academic circles. This study introduces a novel method for measuring information cocoons in academia from two main perspectives: depth and breadth. We utilised two models, BERTopic and Sentence-BERT, to help quantify the depth and breadth of the study. The results of the study show that the degree of information cocoon in the overall citation network is on a decreasing trend, and the information exchange in academia is gradually open and innovative. Secondly, there are diferences in the information cocoon between disciplines, and disciplines with diferent cocoon sizes have their own characteristics, whose uniqueness and complexity need to be taken into full consideration in the assessment. In addition, the study also found that there is a non-linear pattern between the number of citations of scholarly literature and its information cocoon performance. These results stress the need to understand and address information cocoon dynamics in academia, promoting strategies for a more inclusive and diverse scholarly collaborations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;information cocoon</kwd>
        <kwd>academic environment</kwd>
        <kwd>research depth</kwd>
        <kwd>research breadth</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In the era of big data, the explosive growth and overload
of information have led to increased network dependence,
fragmentation, and selective exposure in people’s
information behavior [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In information dissemination, the public
only pays attention to what they choose and the field that
makes them happy. Over time, they will confine themselves
to a cocoon like cocoon room [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. When people in a
positive feedback loop, they are mainly exposed to content they
have already agreed with, which afects the diversity of
information acceptance [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Any environment that generates information is likely to
have an information cocoon, including academia. Within
this system, scholars’ interaction with information can lead
to the formation of an information cocoon. This manifests
when researchers excessively consume similar information
over time, resulting in issues like information narrowing,
group polarization, reduced innovation, and research
bottlenecks. This prompts questions: How prevalent is the
information cocoon in academia? How can it be measured?
And what variations exist among diferent groups?</p>
      <p>
        Previous studies have extensively examined the
formation, impact, and ways to break out of information
cocoons[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However, there has been limited
systematic research on measuring information cocoons, especially
within academic environments. Furthermore, most studies
have focused on social media platforms, with few addressing
academic settings[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ][
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Motivated by the existing research gaps, our primary
objective is to propose a comprehensive method for
measuring the information cocoon within academic environments.
Specifically, we aim to quantify the evolutionary changes
in the value of information cocoon within academia. To
achieve this, we decompose the information cocoon into
two key components: research depth and research breadth.
In order to accurately quantify these aspects, we utilize
BERTopic and sentence-BERT techniques. Furthermore, we
intend to analyze the variations in information cocoons
across diferent groups, encompassing various disciplines
and citation levels.</p>
      <p>Our analysis uncovers a downward trend in the value of
the information cocoon, accompanied by disparities among
diferent groups. These findings provide comprehensive
and practical insights into the phenomenon of information
cocoon within academia. It serves as a timely reminder for
scholars to critically examine their perspectives and take
proactive measures to avoid succumbing to the pitfalls of an
information cocoon. By doing so, scholars can efectively
optimize the information environment within academia for
enhanced research outcomes.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Theoretical Foundation</title>
      <sec id="sec-2-1">
        <title>2.1. Information cocooning in an academic context</title>
        <p>
          In the academic career of scholars, it is crucial to balance
continuous horizontal and vertical development. Horizontal
development allows scholars to cover a wide range of fields,
while vertical development allows them to conduct in-depth
research in specific areas. Focusing only on horizontal
development may lead to a superficial understanding of fields
and a lack of expertise, while pursuing only vertical
development may limit the breadth of knowledge. Therefore,
scholars need to maintain in-depth study of specific fields
throughout their careers, while gaining a broad
understanding of other fields, to enhance their ability to solve complex
problems, foster a spirit of innovation, and promote the
allround development of academic research. Such a balance
not only captures the essence of the problem and provides
insights, but also integrates knowledge from diferent fields
to produce comprehensive and diverse results for academic
research[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>Scholars with larger information cocoons tend to perform
poorly in terms of depth or breadth of research, as evidenced
by their inability to break through to innovation in a
particular research direction, or limitations in their research
areas.</p>
        <p>To this end, evaluating the extent of information cocoons
requires a comprehensive consideration of both depth and
breadth. Only by simultaneously addressing these two
dimensions can researchers better transcend the constraints
imposed by information cocoons. Conversely, focusing
solely on one dimension or conducting superficial analyses
may lead to limitations and misconceptions regarding
information. Thus, this paper is grounded in this rationale to
devise methodologies and propose corresponding metrics.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Pretrained Language Model</title>
        <sec id="sec-2-2-1">
          <title>2.2.1. Sentence-BERT</title>
          <p>
            Sentence BERT is a modified version of the pre-trained BERT
network that incorporates siamese and triplet network
structures. By leveraging these structures, Sentence-BERT
generates semantically meaningful sentence embeddings that
can be compared using cosine-similarity [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]. In recent years,
Sentence-BERT has brought about a signicfiant
transformation in NLP applications by capturing sentence meaning
with unprecedented accuracy [
            <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
            ]. Building upon this
advancement, our study utilizes Sentence-BERT to extract
valuable sentence features from document titles. This
facilitates the calculation of similarity, enabling a comprehensive
assessment of information correlation among documents.
By employing this approach, we achieve a more precise
measurement of document relevance, providing a robust
foundation for subsequent analyses.
          </p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2.2.2. BERTopic</title>
          <p>
            BERTopic is a topic modeling technique that leverages a
pre-trained transformer-based language model to generate
document embeddings. These embeddings are then
clustered, and a class-based TF-IDF process is employed to
generate topic representations[
            <xref ref-type="bibr" rid="ref11">11</xref>
            ]. BERTopic has demonstrated
its ability to generate coherent topics, incorporating both
traditional models and retaining competitiveness in subject
modeling. By harnessing the power of BERTopic, we can
accurately identify the themes addressed in scholarly
literature titles, enabling a more precise understanding of the
research scope. This allows us to assess the distribution
of these themes efectively, thereby facilitating an in-depth
evaluation of the research breadth in our study.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        3.1. Data
For data collection, this study utilized the Semantic Scholar
Open Research Corpus (S2ORC), an extensive dataset
comprising 81.1 million academic papers from diverse
disciplines. The corpus includes comprehensive metadata,
abstracts, and parsed references. S2ORC serves as a
centralized repository that aggregates papers from hundreds of
academic publishers and digital archives, resulting in the
largest publicly available collection of machine-readable
academic text to date [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>From the S2ORC dataset, we extracted papers published
within the timeframe of 2010 to 2021. The data collection
process encompassed capturing various information,
including article titles, first authors, reference titles, publication
dates, and citation counts. To ensure a suficient level of
academic expertise among scholars, we removed duplicate and
incomplete entries. Additionally, first authors with fewer
than six publications during the specific period were
excluded. As a result of this rigorous selection process, our
ifnal dataset consists of 107,775 articles.</p>
      <sec id="sec-3-1">
        <title>3.2. Measures</title>
        <sec id="sec-3-1-1">
          <title>3.2.1. Research Depth</title>
          <p>To measure the research depth, we developed two distinct
metrics named "re_depth" and "self_depth". The "re_depth"
metric quantifies the research depth by assessing the
disparity between the target paper and its reference list. On the
other hand, the "self_depth" metric quantifies the research
depth by evaluating the distinction between the target paper
and previous studies published by the same author.</p>
          <p>In the first place, citing relevant literature is of utmost
importance for authors, as it allows them to build upon
existing knowledge and propose new insights. The disparities
in knowledge between their research and the sources they
cite indicate the level of innovation within their study. This
concept of aggregated knowledge at the topic or field level
allows us to observe the macro-scale evolution of
knowledge, emphasizing the critical nature of citing behavior. We
believe that a greater diference between the target literature
and the cited sources reflects a higher level of innovation
in the target paper, indicating a deeper level of research.
Therefore, we utilize the variance between the target
publication and the cited sources as an indication of the research
depth within the target paper. To quantify this variance, we
utilized the Sentence-BERT model to assess title
similarities between a paper and its reference list. The calculation
formula is:</p>
          <p>Ref_depth = 1 −
∑︀
=1  (, )

(1)</p>
          <p>Here,  (, ) represents the similarity between the
paper and each reference, while n denotes the total number of
references to this paper.  refers to an article,  refers to
the i_th reference of .</p>
          <p>Furthermore, the depth of research becomes evident
through the evolving trajectory and intensity of
individual scholars’ pursuits. Each presentation of research
findings signifies a continuous journey of self-challenge and
breakthrough. Scholars who achieve breakthroughs in
research depth often showcase distinctions from prior
research. These distinctions can manifest in the exploration
of new topics or the acquisition of fresh insights within
the same problem domain. To precisely evaluate this depth
of inquiry, Sentence-BERT was employed in this study to
quantify the divergence of each publication authored by
the same individual from their prior research. The dataset
was organized accordingly, categorized by author, and
arranged chronologically in reverse order of publication.
Subsequently, the similarity of each paper to the three papers
preceding its publication time was calculated. The
calculation formula is as follows:
Self_depth =
{︃
1 −
∑︀+=3+1 (, ) ,  + 3 ≤</p>
          <p>3
0,  + 3 &gt; 
(2)</p>
          <p>In this context,  (,  ) represents the similarity
between two specific papers authored by the same individual.
The variables  and  correspond to distinct documents
authored by the same individual.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.2.2. Research Breadth</title>
          <p>To assess the research breadth, we introduced two distinct
metrics: "ref_breadth" and "self_breadth". The "ref_breadth"
metric quantifies the research breadth by examining the
number of topics within the references of the target
paper. Conversely, the "self_breadth" metric quantifies the
research breadth by evaluating the diversity of topics
addressed within the target paper.</p>
          <p>
            Initially, we hypothesized that the number of topics
covered by the references serves as an indicator of the research
breadth within the literature[
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. To capture this valuable
information, we collected the reference titles associated
with each paper and employed the BERTopic model to
classify each reference title into specific topics. Consequently,
we recorded the number of topics for each group of
references as "ref_topic_counts". The formula for calculating the
"ref_breadth" metric is as follows:
          </p>
          <p>Ref_breadth =
ref_topic_counts
10
(3)
By dividing the ref_topic_counts by 10, we harmonized this
value with the scale of other indicators utilized in this study.</p>
          <p>
            Furthermore, the "ref_breadth" metric ofers insights into
whether scholars have explored diverse fields of knowledge
throughout their research endeavors. Through the
utilization of BERTopic modeling, each paper is assigned
probabilities for belonging to various topic groups. In this study,
the Gini coeficient is employed to quantify the breadth of
research interests. The Gini coeficient is a widely used
measure to assess the level of inequality within a dataset
or distribution[
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. A higher coeficient indicates a more
uneven distribution of probabilities among literature topics,
implying a focus on a singular topic and suggesting a
narrower breadth. Conversely, a lower coeficient signifies a
more evenly distributed probability of theme allocation,
suggesting a broader range of diverse themes. The calculation
formula for the Gini coeficient is as follows:
 =
∑︀
=1
∑︀
=1 | −  |
22 ¯
(4)
Subsequently, the "self_breadth" metric is derived as follows:
Self_breadth = 1 −
          </p>
          <p>Gini
(5)
In the formulas,  represents the number of papers, 
denotes the -th paper, and ¯ represents the average value
across all papers.</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>3.2.3. Cocoon Value</title>
          <p>In accordance with our definition of the information cocoon
within an academic context, a reduction in both research
depth and breadth indicates that an article is confined to
a singular aspect of information, thereby increasing the
likelihood of information cocooning. Conversely, an
expansion in both research depth and breadth implies that
an article holds the potential to transcend the information
cocoon. Therefore, the expression for the cocoon value is as
follows:  represents the sum of the four aforementioned
indicators.</p>
          <p>Cocoon = Avg {(1 −</p>
          <p>Mi)}
(6)</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Analysis</title>
      <sec id="sec-4-1">
        <title>4.1. The Evolved Information Cocoon in</title>
      </sec>
      <sec id="sec-4-2">
        <title>Academic Context</title>
        <p>Our initial objective is to examine the phenomenon of
information cocooning within the entire academic environment
over time. To achieve this, we calculated the research depth
and breadth values on an annual basis, followed by
computing the cocoon value for each year. The corresponding
findings are illustrated in Figures 1 and 2. Figure 1 demonstrates
that the changes in two depth indicators (represented by
the blue and orange lines) remain relatively stable, whereas
there is a noticeable increase in the “ref_breath" indicator
(depicted by the red line). Furthermore, Figure 2 presents
the overall cocoon value, which exhibits a decreasing trend
over the years. This trend signifies a continuous opening up
and innovation of information in the academic environment,
reflecting a positive phenomenon.
Subsequently, we conducted an investigation into the
variability of information cocooning across diferent research
ifelds and presented our findings in Figures 3 and 4. To
ensure the reliability of our results, we excluded disciplines
with limited data and focused on disciplines with larger
volumes for analysis. Our analysis reveals notable trends
within specific disciplines. In Figure 4, art, economics, and
computer science exhibit the lowest levels of information
cocooning. This is evident from their higher values in
research depth and self_breadth indicators, as depicted by the
pink, shallow purple, and dark purple bars in Figure 3. On
the other hand, geography, business, and engineering tend
to have larger information cocoons, as indicated by their
lower research depth and breadth values in Figure 3.
Furthermore, disciplines such as education, law, and sociology
demonstrate the ability to partially break through the
information cocoon, thanks to their relatively higher values on
one or more of the four metrics. For example, the discipline
of law exhibits a higher "ref_breadth" value, albeit with a
lower "ref_depth" value.</p>
        <p>These findings emphasize the importance of considering
the uniqueness and complexity of each discipline when
developing strategies or policies to overcome the information
cocoon at the field level. Scholars within each field should
also take into account the specific characteristics of their
discipline’s information cocoon when designing their research
studies.</p>
        <p>discipline</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Information Cocoon at Diferent</title>
      </sec>
      <sec id="sec-4-4">
        <title>Citations Levels</title>
        <p>Finally, our objective is to examine potential variations
among papers at diferent citation levels. To accomplish this,
we gathered citation data for each paper in three
representative fields: art, education, and geography, which exhibit
high-level, mid-level, and low-level information cocoons,
respectively. Subsequently, we classified all papers into four
groups based on their citation counts: Group A comprises
papers with the highest number of citations over 300; Group
B consists of papers with citations ranging from 100 to 300;
Group C includes papers with citations between 10 and 100;
and Group D encompasses papers with the lowest citation
class
art
business
computer_science
economics
education
engineering
geography
law
sociology
0.6
count, less than 10. Then, we calculated the research depth
value and research breadth value within each group and
presented the results in Figures 5 and 6.</p>
        <p>The analysis revealed consistent trends in the metric
values across papers in groups A, B, C, and D. A clear pattern
emerged between the number of citations and the degree of
information cocooning. It was observed that the most highly
cited papers generally exhibited higher levels of research
depth and breadth, indicating their comprehensive
exploration of a specific area along with extensive coverage of
related domains. In group B, which comprised highly cited
papers, there was a focus on academic hotspots, attracting
scholars with broad interests; however, the depth of analysis
may have been comparatively limited. On the other hand,
less-cited papers demonstrate a narrower research breadth
but exhibit a significant level of depth. These papers, which
often delved into niche issues or possessed a high degree
of depth, may have faced challenges in gaining acceptance
due to their specialized nature.</p>
        <p>In conclusion, our findings suggest that while extensive
research can lead to a considerable number of citations,
studies that exhibit both depth and breadth tend to have a
greater impact. It is crucial to recognize that a lower
number of citations does not necessarily imply lower quality.
Instead, such papers may possess a high level of depth or
focus on niche topics, holding potential for further
development. Therefore, instead of solely emphasizing citation
counts, evaluating research based on both research depth
and breadth can provide more informative insights.
Consequently, research depth and breadth can serve as indicators
for scholars to reflect upon the information cocoons, as well
as for the academic community to assess the influence of
research.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>In our study, we present an index and methodology for
quantifying the scale of information cocoons within
academic environments and classify them accordingly.
The key findings of this paper can be summarized as
follows. Firstly, we observe a gradual breakdown of
information cocoon within the overall academic landscape,
indicating a trend towards greater comprehensiveness and
innovation. Secondly, disparities exist in terms of research
depth, breadth, and information cocooning across diferent
art</p>
      <p>edu
discipline
geo
geo
disciplines. Lastly, it is worth noting that while some papers
may accumulate citations through diverse research, it is the
papers that possess both research depth and breadth that
have the potential to truly become influential. Additionally,
it is important to consider that papers with a low citation
count may be a result of delving deeper into niche topics,
rather than indicating lower research quality. Therefore,
scholars should adeptly leverage extensive and intricate
academic information, continuously evaluating whether
their research processes are constrained by information
cocoons. Communities can utilize the research depth and
breadth metrics proposed in this study to efectively assess
the impact of the research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <source>Research on the Formation Mechanism of Information Cocoon and Individual Diferences among Researchers Based on Information Ecology Theory</source>
          , Frontiers in Psychology (
          <volume>13</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Sanstan</surname>
          </string-name>
          , (
          <year>2008</year>
          ).
          <article-title>Information Utopia: How People Procitation A(&gt;300) B</article-title>
          (
          <volume>100</volume>
          ,300] C(
          <volume>10</volume>
          ,100] D[
          <volume>0</volume>
          ,10]
          <string-name>
            <surname>citation</surname>
            <given-names>A</given-names>
          </string-name>
          (&gt;300) B(
          <volume>100</volume>
          ,300] C(
          <volume>10</volume>
          ,100] D[
          <volume>0</volume>
          ,10] duce Knowledg”, Translated by Bi Jingyue Beijing: Law Publishing House(
          <volume>8</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Falck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Boyer</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2022</year>
          .
          <article-title>Online Filters and Social Trust: Why We Should Still Be Concerned about Filter Bubbles</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] “Bursting Your (Filter) Bubble: Strategies for Promoting Diverse Exposure,”</article-title>
          <source>in Proceedings of the 2013 Conference on Computer Supported Cooperative Work Companion</source>
          , San Antonio Texas USA: ACM, February
          <volume>23</volume>
          , pp.
          <fpage>95</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Nikolov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>D. F. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flammini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2015</year>
          . “Measuring Online Social Bubbles,”
          <source>PeerJ Computer Science (1)</source>
          , p.
          <fpage>e38</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Cinelli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Francisci Morales</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galeazzi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quattrociocchi</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Starnini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2021</year>
          . “
          <article-title>The Echo Chamber Efect on Social Media,”</article-title>
          <source>Proceedings of the National Academy of Sciences (118:9)</source>
          , p.
          <fpage>e2023301118</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Sutherland</surname>
            ,
            <given-names>K. A.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Holistic academic development: Is it time to think more broadly about the academic development project?</article-title>
          <source>International Journal for Academic Development</source>
          ,
          <volume>23</volume>
          (
          <issue>4</issue>
          ),
          <fpage>261</fpage>
          -
          <lpage>273</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Reimers</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks</article-title>
          .
          <source>Conference on Empirical Methods in Natural Language Processing.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          .
          <article-title>North American Chapter of the Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Rath</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Chow</surname>
            ,
            <given-names>J.Y.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Worldwide city transport typology prediction with sentence-BERT based supervised learning via Wikipedia</article-title>
          . ArXiv, abs/2204.05193.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Grootendorst</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>BERTopic: Neural topic modeling with a class-based TF-IDF procedure</article-title>
          .
          <source>ArXiv, abs/2203</source>
          .05794.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Lo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>L. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kinney</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>S2ORC: The Semantic Scholar Open Research Corpus, in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</article-title>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schluter</surname>
          </string-name>
          , and J. Tetreault (eds.), Online: Association for Computational Linguistics, July, pp.
          <fpage>4969</fpage>
          -
          <lpage>4983</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and Han,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>“Breadth and Depth of Citation Distribution,”</article-title>
          <source>Information Processing &amp; Management (51:2)</source>
          , pp.
          <fpage>130</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Loet</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caroline</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            , &amp;
            <surname>Lutz</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>Interdisciplinarity as Diversity in Citation Patterns among Journals: Rao-Stirling Diversity, Relative Variety, and the Gini coeficient</article-title>
          .,
          <source>arXiv: Digital Libraries</source>
          ,
          <volume>13</volume>
          .1:
          <fpage>255</fpage>
          -
          <lpage>269</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>