<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Learning as an RDF Knowledge Graph</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Färber</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Lamprecht</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Karlsruhe Institute of Technology (KIT), Institute AIFB</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Scholarly Data, Open Science, Ontology Engineering</institution>
          ,
          <addr-line>Machine Learning</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we introduce Linked Papers With Code (LPWC), an RDF knowledge graph that provides comprehensive, current information about almost 400,000 machine learning publications. This includes the tasks addressed, the datasets utilized, the methods implemented, and the evaluations conducted, along with their results. Compared to its non-RDF-based counterpart Papers With Code, LPWC not only translates the latest advancements in machine learning into RDF format, but also enables novel ways for scientific impact quantification and scholarly key content recommendation. LPWC is openly accessible at https://linkedpaperswithcode.com and is licensed under CC-BY-SA 4.0. As a knowledge graph in the Linked Open Data cloud, we ofer LPWC in multiple formats, from RDF dump files to a</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org
enabling LPWC to be readily applied in machine learning applications.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>0000-0001-5458-8645 (M. Färber); 0000-0002-9098-5389 (D. Lamprecht)</p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
CEUR
Workshop
Proceedings
htp:/ceur-ws.org CEUR Workshop Proceedings (CEUR-WS.org)</p>
      <p>ISN1613-073
domains. By incorporating FAIR principles that focus on the availability and reuse of research
data and artifacts, we expect LPWC to improve the discoverability and applicability of machine
learning research results. We make the code used for knowledge graph creation and embedding
generation available online (https://github.com/davidlamprecht/linkedpaperswithcode). In the
following, we present LPWC in detail.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Linked Papers with Code</title>
      <p>Linked Papers With Code Ontology. First, we develop an ontology that adheres to the best
practices of ontology engineering and incorporates as much existing vocabulary as possible.
Given that the PWC data dump is sourced directly from the PWC website, thus lacks a
standardized schema and comprises diverse JSON objects, it was infeasible to directly model it within an
OWL/RDF framework. Consequently, we construct a novel semantic schema to model the data.
An overview of the entity types, object properties, and data type properties can be found in
Figure 1. The LPWC ontology encompasses 13 entity types and 47 relationship types. In addition
to the ontology, which is available as an OWL file, we provide a VoID file, following the Linked
Open Data good practices to describe our linked dataset.</p>
      <p>Linked Papers With Code Knowledge Graph. PWC provides access to its data via a
user-friendly, human-readable website. In addition, it ofers daily JSON data dumps. 1 However,
there are several aspects that currently make using the data dificult: 1. There is a lack of
semantic interoperability. Entities, such as authors or AI models, are represented as strings
without unique IDs. This prevents efective linking of data and creation of knowledge graphs.
2. Due to the complexity of the data, modeling in JSON format proves dificult, especially when
processing or querying the data. This issue becomes particularly apparent with evaluation
tables, which are nested within a JSON structure with up to 19 levels in depth. This results
in significant data redundancy within the file. In contrast, a graph representation provides a
more intuitive and manageable way of modeling. 3. The data, originally designed for a human
readable interface, uses markdown for natural language descriptions of entities, which may not
be optimal when being processed by NLP methods or displaying it outside of the website.</p>
      <p>Data Transformation. To overcome these limitations, we convert the JSON files from the
PWC data dump into an RDF knowledge graph based on the developed ontology. This requires
major changes in the data formatting and data modeling. In the transformation process we,
among other steps, (1) assign unique HTTP URIs to all entities, (2) convert all markdown test to
plain text and (3) link the entities to other scholarly data sources in the LOD cloud.</p>
      <p>
        Author Name Disambiguation. The disambiguation of author names given as strings is
a crucial step on top of the pure data transformation. Specifically, we develop an eficient
two-step method to link the 1,471,006 authors in LPWC to entities in SemOpenAlex, which is a
massive RDF dataset modeling the academic landscape with its publications, authors, sources,
and institutions, via its public SPARQL endpoint [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We leverage LPWC author names and
paper titles for the disambiguation. The first step involves exact name matching and publication
title substring comparison.
1See https://github.com/paperswithcode/paperswithcode-data
Paper
Evaluation
Paper with Evaluations
Repository
Model
Dataset
Task
Method
Conference
# Instances
      </p>
      <p>F1-score
MRR</p>
      <p>Accuracy
0.2 0.4 0.6 0.8 1
Spearman Correlation</p>
      <p>ROUGE-1/ROUGE-2</p>
      <p>If no match is found, the second step employs a variant search of LPWC paper titles in
SemOpenAlex works, and author matching based on fuzzy similarity techniques. This process yields
947,709 links to SemOpenAlex entities. The remaining 523,297 author names are represented in
LPWC using the lpwc:authorName property.</p>
      <p>Creating owl:sameAs statements. We further link all conferences modeled in LPWC to DBLP.
Moreover, we successfully map 267,314 papers (71% of all papers in LPWC) to SemOpenAlex
works, utilizing variations of the LPWC paper titles. Lastly, we are able to create 158 mappings
(2% of all datasets) between datasets modeled in LPWC and datasets modeled in Wikidata.</p>
      <p>Key Statistics. Our knowledge graph’s SPARQL endpoint enables the direct computation
of interesting statistics. For instance, Table 1 shows the frequency of entities across entity types.
Additionally, Figure 2 illustrates how to compare conferences (here: NAACL, EMNLP, ACL)
based on the used evaluation metrics of their papers.</p>
      <p>
        Knowledge Graph Embeddings. To enable additional use cases, we compute knowledge
graph embeddings for LPWC. Embeddings have proven to be valuable as implicit knowledge
representations in various scenarios. We train the embeddings based on state-of-the-art
embedding techniques such as TransE, DistMult, ComplEx, and RotatE [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. The training process
involves a maximum of 900 epochs, implementing early stopping based on the mean rank
calculated on the validation sets at intervals of 300 epochs. Among the evaluated techniques,
TransE shows the best results. Therefore, we provide the TransE-based embedding vectors
for all entities and relations online and all our evaluation results in our repository. Notably,
our provided embeddings are in line with state-of-the-art results on benchmark datasets with
similar characteristics in terms of the number of relations, triples, and entities [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>Use Case Examples. LPWC can enhance existing use cases while also enabling the
development of new ones. In the following, we highlight some potential use cases:
1. Machine Learning Data Analysis: LPWC is a novel scientific knowledge graph covering
the current field of machine learning. Complex analyses, such as comparing conferences
or detecting new research topics, become possible in this way.
2. Scholarly LOD Cloud Enrichment: LPWC is highly integrated with the LOD cloud and
connected to multiple data sources such as SemOpenAlex, Wikidata, and DBLP. This
enables eficient data integration and enhanced research data management in alignment
with the FAIR principles.
3. Academic Recommender Systems: Given the information overload in science, scientific
recommender systems are becoming increasingly important. LPWC and the provide
knowledge graph embeddings can be used directly to build state-of-the-art recommender
systems for key scientific content. With LPWC, these systems can recommend also items
such as datasets, methods, and conferences.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Conclusion</title>
      <p>In this paper, we presented Linked Papers with Code, the first RDF knowledge graph with detailed
information about the machine learning landscape, consisting of close to 8 million RDF triples.
We outlined the creation process of this dataset, discussed its characteristics, and examined
the procedure for training state-of-the-art knowledge graph embeddings. In future work, we
aim to leverage the extensive interconnectivity between LPWC and SemOpenAlex to facilitate
large-scale key content extraction from publications.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Haris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stocker</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. D'Souza</surname>
            ,
            <given-names>K. E.</given-names>
          </string-name>
          <string-name>
            <surname>Farfar</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Vogt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Prinz</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Wiens</surname>
            ,
            <given-names>M. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Jaradeh</surname>
          </string-name>
          ,
          <article-title>Improving Access to Scientific Literature with Knowledge Graphs</article-title>
          ,
          <source>Bibliothek Forschung und Praxis</source>
          <volume>44</volume>
          (
          <year>2020</year>
          )
          <fpage>516</fpage>
          -
          <lpage>529</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Färber</surname>
          </string-name>
          ,
          <article-title>The Microsoft Academic Knowledge Graph: A Linked Data Source with 8 Billion Triples of Scholarly Data</article-title>
          ,
          <source>in: Proceedings of the 18th International Semantic Web Conference, ISWC'19</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>129</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Färber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lamprecht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Krause</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Aung</surname>
          </string-name>
          , P. Haase,
          <source>SemOpenAlex: The Scientific Landscape in 26 Billion RDF Triples, in: Proceedings of the 22nd International Semantic Web Conference, ISWC'23</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Zhang, Knowledge Graph Embedding via Graph Attenuated Attention Networks</article-title>
          ,
          <source>IEEE access 8</source>
          (
          <year>2019</year>
          )
          <fpage>5212</fpage>
          -
          <lpage>5224</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Demir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-C. N.</given-names>
            <surname>Ngomo</surname>
          </string-name>
          ,
          <article-title>Convolutional Complex Knowledge Graph Embeddings</article-title>
          ,
          <source>in: Proceedings of the 18th Extended Semantic Web Conference, ESWC'21</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>409</fpage>
          -
          <lpage>424</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>