<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Citing Foreign Language Sources : an Analysis of the S2ORC Dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marc Bertin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iana Atanassova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Université de Franche-Comté</institution>
          ,
          <addr-line>CRIT 30 rue Mégevand, F-25000 Besançon</addr-line>
          ,
          <institution>France Institut Universitaire de France (IUF)</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1918</year>
      </pub-date>
      <fpage>66</fpage>
      <lpage>76</lpage>
      <abstract>
        <p>. BIR 2023 : 13th International Workshop on Bibliometric-enhanced Information Retrieval at ECIR 2023, April 2,</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The English language has gradually gained a dominant position in science [1]. At the same
time, the question of the use of diferent languages in publications remains a subject of scientific
debate, which also finds an echo at the political level. It has been shown by Kirillova in 2019 [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ]
that the publication of articles in English in a given country is related to its scientific activity as
well as to the size of the country. These results reflect the desire of large countries to maintain
their national language as the language of science, while small non-English speaking countries
aim to reach international standards.
      </p>
      <p>
        In addition, Moskaleva and Akoev [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ] showed that non-English native language publications
are less read and cited than those in English outside the home country. They also observed that
the ranking of journals is correlated with the share of English publications for multilingual
journals. In terms of content, Di Bitetti compared the abstracts of a bilingual journal showing
that there were no diferences in aspects such as quality, general interest, etc. between articles
published in English and in other languages in these journals [
        <xref ref-type="bibr" rid="ref3">4</xref>
        ].
      </p>
      <p>
        The work of Smirnova and Lillis in 2022 is based on a corpus comparing research articles
written in Russian with those written in English in the following disciplines : philosophy,
sociology and economics [
        <xref ref-type="bibr" rid="ref4">5</xref>
        ]. At the micro level, the article analyses the changes in citations
in the English and Russian texts. At the macro level, the article raises questions about what is
considered ”citation-worthy” in diferent geolinguistic contexts and considers the consequences
of citation brokerage and knowledge production practices and circulation on a global scale.
The work of Angulo en 2021 [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ] shows that science based on a single language, particularly
English, can be a barrier to knowledge transfer that can lead to bias in the provision of global
models. Experience shows that including non-English sources can reduce bias in understanding
and enrich scientific knowledge. Faced with this hegemony, some publishers have begun to
advocate bilingual journals [
        <xref ref-type="bibr" rid="ref6">7</xref>
        ].
      </p>
      <p>
        In this article we address the problem of multilingualism in publications through a
corpusbased experiment using the Semantic Scholar Open Research Corpus (S2ORC) [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ]. Our objective
is to examine foreign language (non-English) references in this corpus, and observe their nature
and distribution.
      </p>
      <p>
        S2ORC is a large dataset of full-text peer-reviewed research articles, mainly written in English.
We chose this dataset for our experiment because it focuses on English research articles and
spans many academic disciplines. S2ORC includes articles from a wide range of scientific
ifelds including medicine, biology, chemistry, engineering, computer science, physics, materials
science, mathematics, psychology, economics, political science, business, geology, sociology,
geography, environmental science, art, history and philosophy. Other large datasets of research
articles exist, such as the thematic dataset COVID19 [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ], Climate change [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ] or the multilingual
dataset ISTEX [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ], which have diferent coverage.
      </p>
      <p>When studying references to foreign language sources, we should take into account the
problem of the multiplicity of scripts. Indeed, if we look at an article written in English, it may
contain references in other languages that are expressed either in the native alphabet of the
foreign language or in Latin characters. Traditionally, two types of operations have been used
to romanise languages with non-Latin scripts. On the one hand, transliteration is the operation
that consists in replacing each grapheme of one writing system by a grapheme or group of
graphemes of another system, regardless of the pronunciation. On the other hand, transcription
is the opposite of transliteration. In transcription, each phoneme of a language is replaced by a
grapheme or group of graphemes from one writing system to another. Transliteration is based
on national and international standards 1.</p>
      <p>1. E.g. Armenian : ISO 9985 : 1996 ; Macedonian, Turkish, Russian, Ukrainian, Belarusian, Bulgarian, non-Slavic
languages in Cyrillic, based on ISO 9 : 1995 ; Chinese with NF ISO 7098 : 1992 ; Georgian, using ISO 9984 : 1996 ;
Greek with ISO 843 : 1997 ; Hebrew, Yiddish or Syriac, which is a Hebrew script based on the NF ISO 259-2 : 1995
standard ; or Thai, which uses the ISO 11940 : 1998 standard.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methods</title>
      <sec id="sec-2-1">
        <title>2.1. Dataset</title>
        <p>We use the full S2ORC version 1 dataset, which contains approximately 81 million open access
articles published up to 2020. While the dataset is intended to include only English language
articles, a closer look reveals that a small proportion of the articles are in other languages.
We show this below. The dataset is available in json format. Each article is identified by its
paper_id. For our experiment, we extracted the metadata of articles (title and year) and the
metadata of the bibliographic references they contain (titles, years, journals, etc.).</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Processing Pipeline</title>
        <p>The language of each article and bibliographic reference was detected using Google’s Compact
Language Detector v3 (gcld3) Python library 2, which implements a neural network model for
language identification. We used the title of the reference as input. Several outputs are provided
by gcld3, including the code 3 of the most likely language and the likelihood score for that
language. The titles of the bibliographic references were used for the language detection.</p>
        <p>In order to obtain a good quality sample, we discarded references for which the probability
score was lower than 0.95. This is the case, for example, when the title is too short to identify the
language with certainty. References with missing metadata, e.g. missing year, were also ignored
(less than 1 % of all references). Furthermore, papers and references with a year before 1950 or
after 2020 were ignored, as such values are most likely due to typing errors in the dataset (less
than 0.5% of all references).</p>
        <p>The metadata for the citing paper (title and year) was obtained by using its paper_id in
S2ORC. The language of the citing paper was determined in the same way as for the references,
using its title.</p>
        <p>The quality of the language detection varies for some of the poorly endowed languages. The
choice of the gcld3 library was done after testing several other libraries for language detection,
e.g. spacy lang-detect. We found that for our task, gcld3 provides better results and is much
faster. We manually evaluated language detection by examining a subset of 100 titles for each
language. If the observed precision was below 50 % we excluded that language 4 from our
experiment. For the majority of the languages that remained the precision is above 75%.</p>
        <p>The gcld3 recognises cases of Latin transliteration and assigns a language code followed
by -Latn, e.g. ru-Latn stands for Russian text transliterated with Latin characters. In our
experiment, we do not make a distinction between references that are transliterated and those
that are written in the native script.</p>
        <p>2. https://github.com/google/cld3
3. The language codes use the IETF BCP 47 language tags, which combine several standards : ISO 639, ISO 15924,
ISO 3166-1 and UN M.49. F. The subtags are maintained by the IANA Language Subtag Registry.</p>
        <p>4. This applies to the following languages : af, ar, az, be, bg-Latn, ceb, co, cy, eo, et, eu, fil,
fy, ga, gd, gl, ha, haw, hi, ht, ig, jv, kk, ku, ky, lb, mg, mi, mk, mn, ms, mt, ne, ny, sm,
sn, so, sq, st, su, sw, uz, vi, xh, yi, yo, zu.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Composition of the Linguistic Groups</title>
        <p>In order to study how the diferent language groups are represented in the dataset, we
have classified the languages with their language tags into language groups. The summary is
presented in table 1, which contains all language tags identified as present in the dataset. Only
English was not associated with its group (the Indo-European Germanic languages), but we
have separated it from the other languages because it is the dominant language of the corpus.
Languages that were excluded from the experiment because of the poor quality of the language
detection are presented in red.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>As a result of the language detection, after applying the above criteria, we obtained a subset
of S2ORC containing a total of 5.9 million articles with 109.9 million references for which the
language was detected with a probability score above 0.95.</p>
      <sec id="sec-3-1">
        <title>3.1. Languages present in the corpus</title>
        <p>Although the dataset is intended to contain only English research articles, we found 35,463
articles that were identified as being written in other languages, out of a total of 5,934,799
articles (0.60 %). Figure 1 shows a bar chart of the number of articles found for each language.</p>
        <p>Most of the articles are in Latin, and this result may be biased because titles in the biomedical
domain may be misclassified as Latin text. For the remaining languages, we manually sampled
and checked the presence of such articles in the dataset. In our study, we have taken this result
as an opportunity to work with a subset of multilingual (non-English) research articles and
study their references. Overall, we found that the dataset contains articles in 32 languages
(including English) and references to sources in 35 languages (including English).</p>
        <p>As might be expected, the vast majority of references in the dataset are to English sources.
They account for 107,437,865 references out of a total of 109,838,938 references (97.81 %). The
remaining 2,401,073 references are to non-English sources. Figure 2 shows a bar chart of the
number of references to sources in the diferent languages.</p>
        <p>To gain a better understanding of the types of foreign language sources cited, we extracted
titles of non-English sources cited in English articles and translated some of them for analysis
and discussion. The translation is intended to be informative. The results are shown in the
table 2. Languages that were excluded from the experiment because of the poor quality of
the language detection are presented in red. We can see that the titles correspond to studies
that have local significance and dimension, relate to a particular country or geographical area,
and are written in their native language. For example, the title in haw (Hawaiian) represents a
study on the history of ”Kahalu’u and Keauhou”. The example in sq (Albanian) is a study of
hydrocarbon energy in Albania, and the example in mk (Macedonian) is about geological studies
in the villages of Kosel and Pesočan.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Evolution over time</title>
        <p>We have studied the distribution of citations to non-English sources with respect to the year
of reference and with respect to the year of publication of the citing article.</p>
        <p>Figure 3 shows the relative proportion of non-English sources cited by articles published
between 1950 and 2020. The numbers on the horizontal axis represent the average number of
sources cited per article in our dataset. The nine most frequent languages are shown in diferent
colours, and the last group contains all other languages. We can see that between 1960 and 1980,
the non-English references cited were mostly in French and German. Since 2000, however, their
share has been decreasing, with the appearance of a large number of sources cited in Portuguese
and Spanish. Overall, the relative share of non-English references has increased since 2000.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Use of English in relation with other languages</title>
        <p>To study how foreign language sources are cited in English articles, we looked at the subset of
references from English articles to non-English sources. Figure 4 shows the relative proportions
of diferent language groups cited in English articles. The left-hand side of the figure shows the
linguistic groups of the citing articles, while the right-hand side shows the linguistic groups of
the references. The majority of references are from the Indo-European Romance and Germanic
groups, which are the closest languages to English. The Indo-European Slavic languages come
next, and all the other groups have very small proportions of citations.</p>
        <p>Finally, we examined the subset of references from non-English articles. The results are shown
in figure 5. Again, the left side of the figure represents the language groups of the citing articles,
while the right side represents the language groups of the references. It is interesting to note
the dominance of English in this group, as we see that the majority of citations in these articles
are to English sources. In addition, articles written in all other language groups cite a majority
of English sources. In particular, a significant proportion of publications in Indo-European
Romance languages (other than English) cite sources in the same language group, and the same
is true for the other Indo-European groups.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion and Limitations</title>
      <p>The results presented in this study may be biased by several diferent factors. Firstly, the
quality of language detection may vary between diferent languages, particularly for poorly
endowed languages which may have a low recognition rate. Manual cleaning and evaluation
would be required to produce better quality data. Secondly, the use of titles as proxies to detect
the language of a publication is not optimal, as titles can be very short or contain technical
terms that may wrongly point to Latin or English. However, given the size of the dataset and
the data available for references, this was the only way to identify the languages of the sources.
Another possibility would be to use titles and abstracts. This study could therefore be improved
by harvesting abstracts for the references in the dataset.</p>
      <p>
        The choice of the S2ORC dataset introduces an important bias related to the English language.
As S2ORC is intended to include only English language research articles, the subset of
nonEnglish language articles that we analysed may not be representative of these languages, as it is
made up of publications that are included in S2ORC. The S2ORC dataset was constructed by
retaining only papers identified as English using the cld2 tool run over titles and abstracts [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ].
This introduces a bias towards English and other languages that use the Latin alphabet. We can
suppose that such a tool is likely to perform better when it comes to excluding all non-latin
non-English articles from the dataset than the articles written in Latin script. This may lead to
an over-representation of Indo-European Romance languages. Furthermore, as the extraction
pipeline used for S2ORC is optimised for the English language, and the extraction of citations
to foreign language papers has not been evaluated for the creation of the dataset. Thus, it can
be expected that the extracted references are biased towards English language references.
      </p>
      <p>In general, the choice of sources to be cited in a publication and their languages is a complex
process influenced by many factors, such as the languages spoken by co-authors, the subject of
the study, and so on. The editorial requirements of some journals may favour English language
references or require the titles of foreign language references to be translated into the language
of the publication.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Perspectives</title>
      <p>We carried out an analysis of the Semantic Scholar Open Research Corpus (S2ORC) and
identified the languages of the research articles and their references based on their titles. While
the vast majority are in English, we found articles in 44 diferent languages and references in 54
languages. We observed their linguistic groups and their distribution over time between 1950
and 2020. The results allow us to observe the dominance of English in science, where the vast
majority of citations in non-English publications are to English sources. We also show that the
relative share of non-English citations will increase from 2020 onwards.</p>
      <p>Evaluating and improving the quality of language detection can provide us with more reliable
results in the future. In the long run, this work aims to propose a typology of foreign language
references in order to better understand the way they are used in publications. Indeed, one of
the issues that interest bibliometricians today is the reflection on the context of citation. The
classification of citation contexts according to multilingual criteria has not been addressed in
the literature to our knowledge. To this end, an extension of this approach will focus on national
and other multilingual corpora. It also seems relevant to study such references in terms of their
place in publications and their linguistic contexts.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
    </sec>
    <sec id="sec-7">
      <title>Références</title>
      <p>This work was supported by grant number ANR-20-CE38-0003-01 and grant number
ANR21-CE38-0003-01.</p>
      <p>[1] P. S. Rao, The role of english as a global language, Research Journal of English 4 (2019)
65–79.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>O. V.</given-names>
            <surname>Kirillova</surname>
          </string-name>
          ,
          <article-title>Publication language and the journal scientometric indicators in global citation databases</article-title>
          ,
          <source>Science Editor and Publisher</source>
          <volume>4</volume>
          (
          <year>2019</year>
          )
          <fpage>21</fpage>
          -
          <lpage>33</lpage>
          . doi :
          <volume>10</volume>
          .24069/
          <fpage>2542</fpage>
          -0267-2019-1-2-
          <fpage>21</fpage>
          -33.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>O.</given-names>
            <surname>Moskaleva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Akoev</surname>
          </string-name>
          ,
          <article-title>Non-english language publications in citation indexes - quantity and quality</article-title>
          , CoRR abs/
          <year>1907</year>
          .06499 (
          <year>2019</year>
          ). URL : http://arxiv.org/abs/
          <year>1907</year>
          .06499.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Di Bitetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Ferreras</surname>
          </string-name>
          ,
          <article-title>Publish (in english) or perish : The efect on citation rate of using languages other than english in scientific publications</article-title>
          ,
          <source>Ambio</source>
          <volume>46</volume>
          (
          <year>2017</year>
          )
          <fpage>121</fpage>
          -
          <lpage>127</lpage>
          . doi :
          <volume>10</volume>
          .1007/s13280-016-0820-7.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Smirnova</surname>
          </string-name>
          , T. Lillis,
          <article-title>Citation in global academic knowledge making : A paired text history methodology for studying citation practices in english and russian</article-title>
          ,
          <source>Journal of English for Research Publication Purposes</source>
          <volume>3</volume>
          (
          <year>2022</year>
          )
          <fpage>78</fpage>
          -
          <lpage>108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Angulo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Diagne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ballesteros-Mejia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Adamjy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Akulov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Capinha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Dia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Dobigny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. G.</given-names>
            <surname>Duboscq-Carra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Golivets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Haubrock</surname>
          </string-name>
          , G. Heringer,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kirichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kourantidou</surname>
          </string-name>
          , C. Liu,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Nuñez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Renault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Taheri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. N.</given-names>
            <surname>Verbrugge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Watari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Courchamp</surname>
          </string-name>
          ,
          <article-title>Non-english languages enrich scientific knowledge : The example of economic costs of biological invasions</article-title>
          ,
          <source>Science of The Total Environment</source>
          <volume>775</volume>
          (
          <year>2021</year>
          )
          <article-title>144441</article-title>
          . doi :
          <volume>10</volume>
          .1016/j.scitotenv.
          <year>2020</year>
          .
          <volume>144441</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Rosselli</surname>
          </string-name>
          ,
          <article-title>Moving towards english</article-title>
          ,
          <source>Acta Neurológica Colombiana</source>
          <volume>36</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          . doi :
          <volume>10</volume>
          .22379/24224022270.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kinney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weld</surname>
          </string-name>
          ,
          <article-title>S2ORC : The semantic scholar open research corpus, in : Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>4969</fpage>
          -
          <lpage>4983</lpage>
          . doi :
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>447</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chandrasekhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Reas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Burdick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Eide</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Funk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Katsis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Kinney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Merrill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mooney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Murdick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sheehan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stilson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Wade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. X. R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wilhelm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Raymond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Weld</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kohlmeier</surname>
          </string-name>
          , CORD-
          <volume>19</volume>
          : The COVID-19 open research dataset,
          <source>in : Proceedings of the 1st Workshop on NLP for COVID-19 at ACL</source>
          <year>2020</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          . URL : https://www.aclweb.org/anthology/
          <year>2020</year>
          . nlpcovid19-acl.1.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Grundmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Krishnamurthy</surname>
          </string-name>
          ,
          <article-title>The discourse of climate change : A corpus-based approach, Critical approaches to discourse analysis across disciplines 4 (</article-title>
          <year>2010</year>
          )
          <fpage>125</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cuxac</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Collignon, ISTEX, un projet national d'archives documentaires : au-delà de l'accès au texte intégral, l'enrichissement des données par méthodes de fouille de textes, in : Analyser la science : les bibliothèques numériques comme objet de recherche in 85ème Congrès ACFAS</article-title>
          , Montréal, Canada,
          <year>2017</year>
          . URL : https://hal.science/hal-01869036.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>