<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mapping bibliographic metadata collections: the case of OpenCitations Meta and OpenAlex</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elia Rizzetto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvio Peroni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Research Centre for Open Scholarly Metadata, Department of Classical Philology and Italian Studies, University of Bologna</institution>
          ,
          <addr-line>Bologna</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This study describes the methodology and analyses the results of the process of mapping entities between two large open bibliographic metadata collections, OpenCitations Meta and OpenAlex. The primary objective of this mapping is to integrate OpenAlex internal identifiers into the existing metadata of bibliographic resources in OpenCitations Meta, thereby interlinking and aligning these collections. Furthermore, analysing the output of the mapping provides a unique perspective on the consistency and accuracy of bibliographic metadata, offering a valuable tool for identifying potential inconsistencies in the processed data.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Bibliographic collection</kwd>
        <kwd>entity mapping</kwd>
        <kwd>OpenCitations</kwd>
        <kwd>OpenAlex</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Open bibliographic metadata collections play a pivotal role in enabling reproducible studies in
the fields of bibliometrics, scientometrics and science of science and permit transparent
procedures in the context of research assessment exercises, thus enabling the implementation
of norms and guidelines that intend to reform the research assessment around the world, such
as the Coalition for Advancing Research Assessment (CoARA1). As the volume and diversity
of scholarly publications continue to expand, the need for comprehensive and interoperable
bibliographic databases becomes increasingly pronounced.</p>
      <p>
        This study delves into the process of mapping entities between two important open
bibliographic metadata collections, OpenCitations Meta [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and OpenAlex [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These mapping
processes are a critical step towards enabling researchers, institutions, and platforms to access
and utilise information seamlessly across diverse collections. In our work, the primary
objective of this mapping is to integrate OpenAlex internal identifiers into the existing
metadata of bibliographic resources (BRs) in OpenCitations Meta, thereby interlinking and
aligning these collections. This paper presents the results of the mapping and provides details
on the methodology adopted to accomplish this task. By shedding light on the complexities
inherent in aligning bibliographic metadata collections, we aim to contribute valuable insights
into the challenges and opportunities associated with such endeavours.
      </p>
      <p>Furthermore, the study investigates the mapping process's implications to assess the
quality of the involved datasets. Analysing the output of the mapping provides a unique
perspective on the consistency and accuracy of bibliographic metadata, offering a valuable
tool for identifying potential inconsistencies in the processed data. The importance of such
considerations lies in their capacity to enhance data quality, fortify interoperability, and foster
a more cohesive scholarly metadata landscape.</p>
      <p>The rest of the paper is structured as follows. In Section “Material and methods”, we
introduce the processed data and the mapping methodology. Then, in Section “Results”, we
present the result of the mapping analysis. Section “Discussions” discusses some of the most
relevant outcomes, highlighting the broader implications of mapping large bibliographic
metadata collections for data integration, quality enhancement, and improved interoperability
within the scholarly domain. Finally, in Section “Conclusions”, we conclude the paper by
sketching out some future works.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Material and methods</title>
      <p>The following subsections analyse multi-mapped and non-mapped BRs in more detail.
2.1. Data</p>
      <p>
        OpenAlex is a collection of scholarly metadata curated and published by OurResearch13,
and initiated in response to the discontinuation of the Microsoft Academic Graph (MAG) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
It features five types of entities providing rich metadata: Works (such as journal articles,
books, and datasets), Sources (i.e. where works are contained, such as journals, conferences,
and repositories), Authors, Institutions, and Concepts. Metadata include external persistent
identifiers (PIDs): DOI, PMID, PMCID, and MAG ID for Work entities (journal articles,
proceeding papers, etc.); ISSN, Wikidata ID14, MAG ID and Fatcat ID15 for Source entities
(journals, books, etc.). Within OpenAlex, entities are identified with a persistent ID scheme,
i.e. the OpenAlex ID. Data is published under CC0 license and accessible via a REST API, a
web-based GUI, or as downloadable snapshots of the whole database (JSON-Lines files) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        In the scope of this paper, the most relevant differences between OC Meta and OpenAlex
concern the number of BRs in the two collections, the data sources they use, and some
differences in the data models:
 OpenAlex is the largest open scholarly data collection, currently comprising
246,844,573 Works and 249,408 Sources, for a total of 247,093,981 BRs. The latest
version of OpenCitations Meta includes 105,953,699 BRs.
 Data in OpenAlex is provided mainly by Crossref and inherited by the now-ceased
Microsoft Academic Graph, but it also includes data from PubMed [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the Directory
of Open Access Journals (DOAJ) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Unpaywall [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], arXiv [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], Zenodo [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the
ISSN International Centre16, and the Internet Archive’s General Index17. OC Meta’s
sources are Crossref, the National Institute of Health Open Citation Collection
(NIHOCC, providing PubMed data) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], OpenAIRE [15], and the Japan Link Center (JaLC)
[16]18.

      </p>
      <p>In OpenAlex, Works can only have one ID value per each ID scheme, and Sources
admit a list of up to two ISSNs or a single literal value for each of the other ID
schemes. On the contrary, in the OCDM, and therefore in OC Meta, there are no limits
on the number of possible values for each ID scheme. This substantial difference in
how the two collections represent their data implies that, for example, if a journal
article has been assigned two DOIs, they can be linked to the same entity (and the
same OMID) in OC Meta, but not in OpenAlex. Another noteworthy difference is that
OpenAlex does not support ISBNs, while OC Meta does.</p>
      <sec id="sec-2-1">
        <title>2.2. Mapping process</title>
        <p>The process leading to the mapping of these two collections is explained as follows. Initially,
two tables are produced, which contain the internal IDs of the collections to be mapped with
each other. The first table is produced by parsing the CSV dump of OC Meta, and, for each
row, contains the OMID, external PIDs, and type for each BR in OC Meta that has external
PIDs. The other table, produced from the JSON-Lines copy of the OpenAlex database, links
each external PID in OpenAlex to the OpenAlex ID to which it is associated.</p>
        <p>The table containing OpenAlex data is converted into a local SQL database. Then, the table
containing OC Meta BRs to be mapped is iterated line by line, and each PID associated with
each entity is looked up in the database containing PID-OpenAlex ID associations. The result
consists of three additional tables:
13https://ourresearch.org/
14https://www.wikidata.org/wiki/Wikidata:Identifiers
15https://fatcat.wiki/
16https://www.issn.org/
17https://archive.org/details/GeneralIndex
18The data provided by JaLC is not included in the dump version processed for the mapping described by the present work (v5),
but is included in the latest version (v6).</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.3. Multi-mapped BRs analysis: methodology</title>
        <p>The mapping revealed that mapped entities in different datasets might go beyond a simple 1
to 1 alignment. Indeed, it is possible that one BR in OC Meta shares one or more external PIDs
with more BRs in OpenAlex. These cases will be referred to as multi-mapped BRs.</p>
        <p>Such cases, after being saved separately from the rest of the results, have first been checked
manually by investigating sample resources, inspecting their full metadata in both datasets,
making use of external APIs (Crossref [17] and DataCite [18]) and accessing the documents’
location on the web via their PIDs. This study led to proposing an ad hoc categorisation, to
frame the causes of such multi-mapping scenarios. We applied such categorisation to the
instances of multi-mapped BRs by using heuristics to understand which category applies to
each specific case.</p>
        <p>The categories for OC Meta BRs that are multi-mapped to OpenAlex Works are the
following:
1. Category A includes cases where two or more Works among the ones that are
multimapped to a single OC Meta BR share at least one external PID. Given that external
PIDs, such as DOIs, should be uniquely assigned to a BR, having more than one entity
with the same external PID in the OpenAlex dataset means that there are either
duplicate entities or errors in the metadata.
2. Category B includes cases where the same entity in OC Meta is mapped to different
versions of the same publication, each represented by a Work entity in OpenAlex –
e.g. in the case of having a version of record and one or more preprint and/or
postprint versions. Preprints and postprints are hosted in a preprint server or a digital
repository. DOIs of preprints or postprints are determined by considering the DOI
prefix and looking it up on a list of DOI prefixes reserved for institutions that manage
preprint servers or digital repositories for non-peer-reviewed publications.
3. Category C includes cases where the same entity in OC Meta is mapped to exactly 2
different Works in OpenAlex, and neither is a preprint or postprint version. The most
likely causes for this scenario are errors in the data source used by OC Meta, bugs in
OC Meta software, or different DOIs intentionally linked to the same OC Meta entity.
4. Category D includes cases where the same entity in OC Meta is mapped to multiple
preprint versions of the same publication, each represented by a Work entity in
OpenAlex. This typology is similar to category B, but it only includes preprint
versions and detects them by checking for version number (e.g. “/v1”) in the DOI
value.
5. Category E includes cases where the same entity in OC Meta is mapped to multiple
preprint versions of the same publication, each represented by a Work entity in
OpenAlex. This typology is similar to categories B and D, but detects preprint versions
by analysing the DOI value and checking if it contains semantic indicators that
associate the DOI with a preprint server (e.g. “/arxiv” or “/zenodo”).
6. Category F includes cases where the multi-mapped OpenAlex Works include a version
of record, together with one or more Works of type “peer-review”, “letter”, “editorial”,
“erratum”, or “other”. For example, the DOI for an erratum notice and a DOI for the
journal article that is being corrected may be wrongly assigned the same OMID in OC
Meta, due to errors in the data source.</p>
        <p>OC Meta BRs that are multi-mapped to OpenAlex Sources fit only into one category, “A”,
which groups cases where two or more multi-mapped OpenAlex Sources share at least one
ISSN.</p>
        <p>The categorisation process (represented as pseudocode in Listing 1) takes as input:
1. Multi-mapped BRs in the form of a table where each row represents the association of
one BR in OC Meta with n BRs in OpenAlex, storing an OMID in the omid field and a
list of OpenAlex IDs in the openalex_id field;
2. A list of 80 DOI prefixes that are assigned by Crossref and DataCite to organisations
or institutions that manage preprint servers or digital repositories hosting
non-peerreviewed versions.
3. A list of strings that, when found inside a DOI value, indicate that the associated
publication is hosted in a preprint server (e.g. “/arxiv”, “/preprints”, “/osf.io”).
4. A SQL database storing full metadata of the OpenAlex BRs involved in the
multimapping.</p>
        <p>The process differentiates between OpenAlex Works and OpenAlex Sources. For rows
storing Works, the process includes querying the database for external PIDs associated with
each Work. If any PID is associated with multiple Works in the row, the categorisation is
labelled with “A”. Subsequently, each multi-mapped Work is examined. If version-marked
DOIs are present, the categorisation is labelled with “D. Otherwise, an assessment is made for
DOI prefixes associated with preprint servers, leading to categorisations such as “B” for
preprint server association, “E” for preprint indicators, “F” for meeting specific OpenAlex
database criteria, and “C” for rows with only two Works.</p>
        <p>For rows storing Sources, the process involves querying the database for ISSNs associated
with each Source. If any ISSN is associated with multiple Sources in the row, the
categorisation is labelled with “A”.</p>
        <p>Rows that remain unclassified after these steps are marked as unclassified.</p>
        <p>Listing 1
Pseudocode representing the process for multi-mapped categorization.</p>
        <p>FUNCTION categorizationProcess(table, doiPrefixes, preprintIndicators,
database):</p>
        <p>FOR EACH row IN table:</p>
        <p>IF Works IN row.openalex_id:
externalPIDs = queryDatabaseForExternalPIDs(row)
IF hasDuplicates(externalPIDs):</p>
        <p>row.category = "A"
ELSE:</p>
        <p>FOR EACH work IN row.openalex_id:</p>
        <p>IF work.hasDOIs():</p>
        <p>IF hasVersionMarkedDOI(work, versionedDOIregex):</p>
        <p>row.category = "D"
ELSE IF isPublishedByPreprintOrganization(work, doiPrefixes) AND
(work.isSubmittedVersion() OR work.isAcceptedVersion()):
row.category = "B"
ELSE IF containsPreprintIndicator(work, preprintIndicators):</p>
        <p>row.category = "E"
ELSE IF allDOIsHaveSamePrefix(work):</p>
        <p>IF work.isPeerReview() OR work.isEditorial() OR
work.isErratum() OR work.isLetter():
row.category = "F"
ELSE IF countWorksInRow(row) == 2:</p>
        <p>row.category = "C"
ELSE IF Sources IN row.openalex_id:
issns = queryDatabaseForISSNs(row)
IF hasDuplicates(issns):</p>
        <p>row.category = "A"
ELSE:</p>
        <p>row.category = "non classified"</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.4. Non-mapped BRs provenance analysis: methodology</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>Table 1 shows the number of processed BRs for both datasets and the general results of a
quantitative analysis of the mapping output. As mentioned above, a BR entity in OC Meta can
be mapped to a BR entity in OpenAlex only if both entities are associated with at least one
external PID in common. Thus, the BRs in the OC Meta CSV dump that are theoretically
mappable to at least one entity in OpenAlex are 90,270,131, and the set of OpenAlex BRs to
which an OC Meta BR can be mapped amounts to 159,039,790 resources. Of the 90,270,131
mappable resources in the OC Meta CSV dump, most (approximately 97%) map to at least one
resource in OpenAlex. However, a small number of these (173,513, roughly 0.2%) align (i.e.
share external PIDs) with more than one entity in OpenAlex (multi-mapped BRs). At the same
time, and vice versa, there is a consistent number of BRs in OC Meta (5,722,979) that do not
uniquely map to a BR in OpenAlex, meaning that there are also cases where two or more BRs
in OC Meta are aligned with the same entity in OpenAlex. These latter cases will be referred
to as inverted multi-mapped BRs. Finally, 18,133,712 BRs in OC Meta do not map to any
resource in OpenAlex, whether because, after being processed, they have been found not to
have any corresponding entity in OpenAlex despite having external PIDs (2,963,534 BRs);
because they do not have any external PID (9,000,386 BRs); or because they are not included in
the CSV dump files, thus were not processed. Concerning the latter scenario, it is worth
mentioning that the OC Meta software, when producing CSV dump files from the triplestore,
does not represent journal issues and journal volumes as table rows. However, almost all BRs
of these types lack external PIDs, with their OMID being the only persistent identifier.
No. of processed BRs with PIDs also supported by OpenAlex (stored in CSV files) 90,270,131</p>
      <sec id="sec-3-1">
        <title>OC Meta</title>
        <p>OpenAlex
105,953,699</p>
        <p>99,270,517
245,207,435
159,039,790
87,605,238
5,722,979</p>
        <p>173,513
18,133,712
Number of BRs in dump
Number of BRs with PIDs supported also by OC Meta</p>
      </sec>
      <sec id="sec-3-2">
        <title>Mapping OC Meta → OpenAlex</title>
        <p>No. of BRs in OC Meta mapped to exactly one BR in OpenAlex (1:1)
No. of BRs in OC Meta, which map to the same BR in OpenAlex as at least one
other BR in OC Meta (n:1, where n&gt;1)
No. of multi-mapped BRs in OC Meta (1:n, where n&gt;1)
No. of non-mapped BRs in OC Meta</p>
        <sec id="sec-3-2-1">
          <title>3.1. Multi-mapped entities</title>
          <p>Multi-mapped BRs have been analysed with respect to the number of OpenAlex entities
mapped to a single BR in OC Meta. As shown in the distribution histogram in Figure 1, most
cases involve two OpenAlex IDs per OMID (91.5%), followed by cases involving 3 OpenAlex
IDs per OMID at a much lesser proportion (6.2%). The remaining cases (more than 3 OpenAlex
IDs per OMID) are significantly less frequent, with values lower than 1.3%. It should also be
mentioned, though, that some multi-mapped BRs are connected to a particularly high number
of OpenAlex IDs: there are isolated cases of OC Meta BRs being mapped to more than 100
entities in OpenAlex, and even an outlier case involving 1,051 OpenAlex IDs. Such examples,
though not common, may also help reveal potential anomalies or inconsistencies in both
datasets.
19These could potentially include mappings that the categorization heuristics failed to catch, or concern general errors in the data
sources used by OC Meta and/or OpenAlex.
there have been changes in the journal name. While OC Meta tends to prioritise the
fundamental continuity of the journal entity – regardless of variations in names, the number
of ISSNs, or diverse publication media – OpenAlex occasionally encounters challenges in
consolidating all ISSNs under a single entity. In Example 1, for the journal identified as
“br/06602375171”, the “Journal of Health”20, OC Meta has two ISSNs, each assigned to a
different entity in OpenAlex (S2764583335, associated with the online ISSN, and S4210187171,
associated with the print ISSN).</p>
          <p>omid
br/06602375171</p>
          <p>openalex_id
S2764583335 S4210187171
(Example 1)</p>
          <p>book
book chapter
&lt;unspecified&gt;
proceedings</p>
          <p>article
proceedings</p>
          <p>report
reference</p>
          <p>book
reference</p>
          <p>entry
web content</p>
          <p>dataset
dissertation</p>
          <p>series
standard
book section
journal</p>
          <p>A
39,758
38,179
27
341
607
477
8
13
99
0
2
1
0
0
0
0
4
9,421
8,722</p>
          <p>B
1
8
502
10
24
146
1
0
7
0
0
0
0
0
0</p>
          <p>C
3,8984
35,744
581
1,112
609
452
230
155
7
22
14
38
9
4
6
1
0
12,496
10,196</p>
          <p>D
31
21
1,753
108
13
335
37
1
0
0
1
0
0
0
0</p>
          <p>E
1,376
1,030
265
16
0
4
1
0
0
1
0
1
0
0
0
0
58</p>
          <p>Unclassi</p>
          <p>fied
64,132
50,579
8,511
2,002
1,503
666
508
167
69
57
47
10
9
3
1
0
0</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2. Non-mapped entities</title>
          <p>A
4,076
4,057
17
2</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Unclassified 2,383</title>
        <p>2,345
38
0
The entities in OC Meta that have not been mapped to any entity in OpenAlex (i.e.
nonmapped entities) have been analysed with regard to the source they have been provided by.
Provenance information is available as RDF data for the great majority of non-mapped
entities, with only 2094 being left out. Approximately 83% of non-mapped entities do not have
any other PID than their OMID, therefore they cannot be mapped until any other PID also
supported by OpenAlex is associated with them in OC Meta data. Table 4 illustrates a
representative sample of the results of provenance analysis, concerning the ten most frequent
bibliographic entity types among non-mapped entities: it shows how many non-mapped
entities derive from each source or set of sources, and entities are grouped by the type of BR
and by the presence/absence of other PIDs besides OMID.
yes
no
yes
no
yes
no
yes
proceed journal
ings issue
book
journal
volume dataset
unspeci journal referen</p>
        <p>fied article ce book report journal
5,383,11 5,064,03 2,521,88 1,576,74 1,242,10 1,419,21 253,284 188,453 135,997 103,263
5 0 6 4 1 2
31
0
11,487
786
0
0
yes
yes
yes
yes
yes
no
yes
no
yes
17
3,847</p>
        <p>0
3,307
120
12
13
0
0
0
0
28
2,125
0
27
0
1
0
0
0
190
3,521
0
0
0
0
0
0
0
0
6
0
165
16
0
287
0
0
0
0
0
0
0
0
0
0
374
2,246
31
57
0
7
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
115
0
0
0
0
0
0
0
0
0
3
0
6
13
21
0
0
0
0
3
0
0
0
0
0
0</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
      <p>The mapping process and the analysis of its results concerned the study and use of a great
amount of data from the involved databases, requiring, for example, the consideration of all
bibliographic entities in their entirety. This study highlighted problems and inconsistencies
within the used datasets. First, concerning OC Meta, the process provided an opportunity to
conduct counts of the number of entities contained in the CSV and JSON-LD files comprising
the dump. This highlighted a discrepancy between the number of BRs contained in the
triplestore and the number of BRs actually reported in the dump files constructed from the
triplestore. Additionally, it was observed that this numerical difference is also reflected in the
RDF files containing provenance information.</p>
      <p>The analysis of multi-mapped and the count of inverted multi-mapped BRs posed
interesting questions as well. A comparison between OC Meta and OpenAlex from the
perspective of two different data models helped emphasise that both collections have duplicate
entities, i.e. resources sharing the same external PID (e.g., DOI or ISSN, which should be
uniquely assigned), with at least one other resource within the collection. In the case of
multimapped BRs, it was further found that the alignments of a single OMID to multiple OpenAlex
IDs could be attributed partly to natural diversities between data models, partly to errors in
data sources, and partly to errors in the software used to populate the collection. Generally,
OC Meta tends to erroneously group various expressions of a resource (preprints, postprints,
and versions of record) into a single entity, propagating errors present in data sources, even
when there should be two separate entities (e.g. in the case of a version of record and its
preprint, which should be two separate entities according to the OCDM). In contrast,
OpenAlex generally tends to have separate entities due to limits on the number of possible
values for each ID scheme and more intensive data correction activities made possible by the
use of web crawlers.</p>
      <p>From the perspective of OC Meta, while some of these multi-mapped cases result from data
representation choices, others are the result of errors often originated from sources (especially
in cases where an OMID is aligned to a very high number of OpenAlex IDs).</p>
      <p>Regarding non-mapped BRs, we observed that, despite OpenAlex formally including a
greater number of entities than OC Meta, approximately 5 million OMIDs are not associated
with any corresponding OpenAlex ID. This is partly because some of the resources counted as
non-mapped BRs (15,170,179 BRs) were not included in the CSV files that the mapping process
takes as input; therefore they were not processed at all during the mapping phase. Of the
other 2,963,533 non-mapped resources, those with one or more external PIDs are particularly
interesting, as one would expect them to have at least one corresponding entity in OpenAlex.</p>
      <p>In this regard, it should be noted that, in the case of the 108,658 non-mapped books from
Crossref, many resources likely have only ISBNs among the external IDs, which are not
supported by OpenAlex and therefore cannot be used for mapping. Another interesting case is
the set of dataset resources from DataCite, totalling 1,238,173 entities, which can be explained
by the fact that DataCite is not among the sources used by OpenAlex. More generally, the
15,061,152 non-mapped BRs without external PIDs underscore the unique contribution made
by OC Meta by assigning a persistent identifier, i.e. OMID, to entities that would otherwise
lack one. Indeed, the OCDM permits to represent journal issues and journal volumes as
firstclass entities, while they are typically represented only as metadata associated with journal
articles (as is the case for OpenAlex).</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>The results of the mapping of OpenCitations Meta bibliographic resources to OpenAlex
bibliographic resources have provided valuable insights into the integration of bibliographic
metadata entities, showcasing that the majority of processed OC Meta resources are
successfully mapped with exactly one entity in OpenAlex. This achievement is significant, as
it allows for the direct ingestion of OpenAlex IDs into the metadata of the corresponding
bibliographic resources in OC Meta. This seamless integration enhances the
interconnectedness and interoperability of these two substantial bibliographic collections.</p>
      <p>However, challenges were encountered in the case of multi-mapped BRs, leading to the
decision to temporarily exclude them from being included in OC Meta. While this choice
poses a limitation, the analysis of these multi-mapped entities has proven instrumental in
identifying inconsistencies within both datasets. Furthermore, the examination of
nonmapped resources, considering their type and provenance, has underlined the impact of using
different data sources and different identifiers in the collections to map, resulting in quite a
significant limitation of the mapping coverage.</p>
      <p>Addressing the limits and inconsistencies revealed by the mapping results, OpenCitations
has proactively taken measures to rectify errors and enhance the quality of its data,
particularly in the production process of dump files. Future developments are envisioned, e.g.
to further refine the management of scenarios involving bibliographic resources being
associated with multiple values for the same ID scheme (e.g. multiple DOIs for the same
journal article). Improvements like these aim to bolster the robustness of the mapping process
as well as the quality of the data, ensuring a more accurate and comprehensive representation
of bibliographic entities.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This project has been partially funded by the European Research Council Executive Agency
under service contract ERCEA/2023/VLVP/0007 and the European Union’s Horizon Europe
framework programme under grant agreement No. 101095129 (GraspOS Project).
[15] C. Atzori, A. Bardi, P. Manghi, and A. Mannocci, ‘The OpenAIRE Workflows for Data
Management’, in Digital Libraries and Archives, vol. 733, C. Grana and L. Baraldi, Eds.,
in Communications in Computer and Information Science, vol. 733. , Cham: Springer
International Publishing, 2017, pp. 95–107. doi: 10.1007/978-3-319-68130-6_8.
[16] M. Hara, ‘Introduction of Japan Link Center (JaLC)’. ORCID, 2020. doi:
10.23640/07243.12469094.V1.
[17] G. Hendricks, D. Tkaczyk, J. Lin, and P. Feeney, ‘Crossref: The sustainable source of
community-owned scholarly metadata’, Quant. Sci. Stud., vol. 1, no. 1, pp. 414–427, Feb.
2020, doi: 10.1162/qss_a_00022.
[18] J. Brase, ‘Datacite - A Global Registration Agency for Research Data’, SSRN Electron. J.,
2010, doi: 10.2139/ssrn.1639998.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Massari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mariani</surname>
          </string-name>
          , I. Heibi,
          <string-name>
            <given-names>S.</given-names>
            <surname>Peroni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Shotton</surname>
          </string-name>
          , '
          <article-title>OpenCitations Meta'</article-title>
          .
          <source>Jun</source>
          .
          <volume>28</volume>
          ,
          <year>2023</year>
          . doi: https://doi.org/10.48550/arXiv.2306.16191.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Priem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Piwowar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Orr</surname>
          </string-name>
          , '
          <article-title>OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts'</article-title>
          ,
          <source>presented at the 26th Internation Conference on Science and Technology Indicators</source>
          , arXiv,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .48550/ARXIV.2205.
          <year>01833</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Daquino</surname>
          </string-name>
          et al., '
          <article-title>The OpenCitations Data Model'</article-title>
          ,
          <source>in The Semantic Web - ISWC</source>
          <year>2020</year>
          ,
          <string-name>
            <given-names>J. Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Tamma</surname>
          </string-name>
          , C. d'Amato,
          <string-name>
            <given-names>K.</given-names>
            <surname>Janowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Seneviratne</surname>
          </string-name>
          , and L. Kagal, Eds., in Lecture Notes in Computer Science. Cham: Springer International Publishing,
          <year>2020</year>
          , pp.
          <fpage>447</fpage>
          -
          <lpage>463</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -62466-8_
          <fpage>28</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Daquino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Massari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Peroni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Shotton</surname>
          </string-name>
          , '
          <article-title>The OpenCitations Data Model'</article-title>
          . figshare,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .6084/M9.FIGSHARE.
          <volume>3443876</volume>
          .V8.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] 'OpenCitations Meta CSV dataset of all bibliographic metadata'</article-title>
          . doi: https://doi.org/10.6084/m9.figshare.
          <volume>21747461</volume>
          .
          <year>v5</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Peroni</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Shotton</surname>
          </string-name>
          , '
          <article-title>OpenCitations, an infrastructure organization for open scholarship'</article-title>
          ,
          <source>Quant. Sci. Stud</source>
          ., vol.
          <volume>1</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>428</fpage>
          -
          <lpage>444</lpage>
          , Feb.
          <year>2020</year>
          , doi: 10.1162/qss_a_
          <fpage>00023</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I.</given-names>
            <surname>Heibi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Peroni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Shotton</surname>
          </string-name>
          , '
          <article-title>Software review: COCI, the OpenCitations Index of Crossref open DOI-to-DOI citations'</article-title>
          ,
          <source>Scientometrics</source>
          , vol.
          <volume>121</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>1213</fpage>
          -
          <lpage>1228</lpage>
          , Nov.
          <year>2019</year>
          , doi: 10.1007/s11192-019-03217-6.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sinha</surname>
          </string-name>
          et al.,
          <article-title>'An Overview of Microsoft Academic Service (MAS) and Applications'</article-title>
          ,
          <source>in Proceedings of the 24th International Conference on World Wide Web, Florence Italy: ACM</source>
          , May
          <year>2015</year>
          , pp.
          <fpage>243</fpage>
          -
          <lpage>246</lpage>
          . doi:
          <volume>10</volume>
          .1145/2740908.2742839.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Canese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jentsch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Myers</surname>
          </string-name>
          , '
          <article-title>PubMed: The Bibliographic Database'</article-title>
          , in The NCBI Handbook, 2nd ed.,
          <year>2013</year>
          , p.
          <fpage>9</fpage>
          . [Online]. Available: https://www.ncbi.nlm.nih.gov/books/NBK153385/
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Morrison</surname>
          </string-name>
          , '
          <article-title>Directory of Open Access Journals (DOAJ)'</article-title>
          , Charlest. Advis., vol.
          <volume>18</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>28</lpage>
          , Jan.
          <year>2017</year>
          , doi: 10.5260/chara.18.3.25.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Dhakal</surname>
          </string-name>
          , 'Unpaywall',
          <string-name>
            <given-names>J.</given-names>
            <surname>Med</surname>
          </string-name>
          . Libr. Assoc., vol.
          <volume>107</volume>
          , no.
          <issue>2</issue>
          ,
          <string-name>
            <surname>Apr</surname>
          </string-name>
          .
          <year>2019</year>
          , doi: 10.5195/jmla.
          <year>2019</year>
          .
          <volume>650</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sigurdsson</surname>
          </string-name>
          , '
          <article-title>The future of arXiv and knowledge discovery in open science'</article-title>
          ,
          <source>in Proceedings of the First Workshop on Scholarly Document Processing, Online: Association for Computational Linguistics</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>9</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .sdp-
          <volume>1</volume>
          .2.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <article-title>European Organization For Nuclear Research</article-title>
          and OpenAIRE, 'Zenodo: Research. Shared.',
          <year>2013</year>
          , doi: 10.25495/7GXK-
          <fpage>RD71</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Maloney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sequeiera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Orris</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Beck</surname>
          </string-name>
          , '
          <article-title>Pubmed central', in The NCBI Handbook</article-title>
          , 2nd ed.,
          <year>2013</year>
          . [Online]. Available: https://www.ncbi.nlm.nih.gov/books/NBK153388/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>