<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Approach for Automatic Construction of an Algorithmic Knowledge Graph from Textual Resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jyotima Patel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Biswanath Dutta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DRTC, Indian Statistical Institute</institution>
          ,
          <addr-line>Bangalore</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Library and Information Science, Calcutta University</institution>
          ,
          <addr-line>Kolkata</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>There is enormous growth in various fields of research. This development is accompanied by new problems. To solve these problems eficiently and in an optimized manner, algorithms are created and described by researchers in the scientific literature. Scientific algorithms are vital for understanding and reusing existing work in numerous domains. However, algorithms are generally challenging to find. Also, the comparison among similar algorithms is dificult because of the disconnected documentation. Information about algorithms is mostly present in websites, code comments, and so on. There is an absence of structured metadata to portray algorithms. As a result, sometimes redundant or similar algorithms are published, and the researchers build them from scratch instead of reusing or expanding upon the already existing algorithm. In this paper, we introduce an approach for automatically developing a knowledge graph (KG) for algorithmic problems from unstructured data. Because it captures information more clearly and extensively, an algorithm KG will give additional context and explainability to the algorithm metadata.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Knowledge Graph</kwd>
        <kwd>Algorithm</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>Algorithm knowledge graph</kwd>
        <kwd>Automatic approach</kwd>
        <kwd>Metadata extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Scientific knowledge is distributed and maintained through academic literature published in,
for example, conference proceedings, journal articles, and workshop proceedings. It is a
challenging task for researchers to keep track of innovations like proposals for new frameworks,
algorithms, and software with such an increased number of publications. Algorithms (where an
algorithm is a step-by-step strategy to tackle any issue) are published in areas ranging from
Mathematics to Geo-sciences and Computer Science [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Algorithms in the scientific literature
are expressed as flowcharts, pseudo-codes, or computer programs. As algorithms are written
in a step-by-step manner, understanding a scientific problem becomes simple. Studying an
algorithm also helps to know the processes and techniques used to solve a scientific problem.
In the academic community, algorithms are often used to communicate the goal, technique,
and procedures taken towards solving a problem, whereas practical people more often look for
the implementation of the algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. There is no system available where the researchers
can find all the major algorithms, though there are initiatives like The Stony Brook Algorithm
Repository1 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Algowiki2 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and Wikidata3. The main limitations of these repositories are
that they provide a minimal description of the algorithms, they also lack a proper search facility
and browsing for algorithms in these repositories is a tedious and time-consuming process (the
repositories are further elaborated on in section 2). Because of these issues, many important
algorithmic works often go unnoticed. Also, as stated above, since the algorithms are described
minimally, searching for them is a challenge. The algorithm metadata in most cases is either
missing or provided in a very minimal way. The extraction of the metadata is a challenging
task because the information is usually located within the scholarly or other textual resources,
such as the registries and repositories, like Algowiki and Stony Brook Algorithm Repository.
Gathering the metadata manually from these resources is a tedious and time-consuming process.
One of the main focuses of the current study is to provide an approach for extracting the
algorithmic metadata from textual resources. As a first step, we consider the Stony Brook
Algorithm Repository as a potential source for extracting the metadata. Further, we transform
these metadata automatically into a knowledge graph (KG) (a manifestation of an intelligent
web of Data informed by an ontology[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][5]). The transformation of algorithm metadata into
a KG helps in representing the algorithms and their relations with high dependability, logic,
and reusability. Apart from providing better search and retrieval, the KG identifies related
information and helps in the comparison of algorithms. The KG can also be used as a referential
model for the automatic extraction of information from the scientific literature.
      </p>
      <sec id="sec-1-1">
        <title>The primary contributions of this work are:</title>
        <p>• a methodology for the automatic creation of a knowledge graph for algorithmic problems.
• an approach for the automatic extraction of algorithmic problems and associated data
from textual resources.</p>
        <p>• generation of a knowledge graph using the extracted data.</p>
        <p>The rest of the paper is organized as follows: Section 2 discusses the existing algorithm
repositories and provides a comparison between them; section 3 gives an overview of the KG
development approach and the steps involved in it; section 4 discusses the data extraction i.e
the steps involved in the extraction of the data and the data processing; section 5 discusses the
KG creation process from the processed data and also the tools used; section 6 provides some
SPARQL queries to validate the KG; section 7 discusses the relevant related works and Section 8
concludes the paper and provides future research directions.</p>
      </sec>
      <sec id="sec-1-2">
        <title>1https://algorist.com/algorist.html 2https://wiki.algo.is/ 3https://www.wikidata.org/</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Existing Algorithm Repositories</title>
      <p>
        Algorithm repositories are storage areas for algorithms, where algorithms are often stored with
minimal metadata. We found three dedicated algorithm repositories on the web. The Stony
Brook Algorithm Repository provides an exhaustive collection of 75 fundamental algorithmic
problems along with their implementations. The algorithms in this repository are categorized
into two groups: by language and by problem [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Each algorithm has its own separate web
page containing information like input, output, input description, implementations, related
problems, and so on.
      </p>
      <p>
        Algowiki is an online encyclopedia of algorithms. It contains a list of algorithms on a wiki page
arranged in alphabetical order. Algowiki consists of algorithms from the field of mathematics
with some features and properties. It is a wiki dedicated to competitive programming [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The
algorithms in the Stony Brook algorithm repository are held in one web Server, whereas in
Algowiki once you click on the desired algorithm, the site is redirected to either a Wikipedia page
or a website that holds the description of the algorithm with minimal metadata. In Algowiki,
only the URL of each algorithm is stored [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Wikidata is a knowledge base that acts as central storage for structured data of its Wikimedia
sisters [6]. Initiatives are taken by Wikidata to model algorithms where they use generic
properties to describe algorithms and then interlink the various other algorithms as sub-classes
or instances. The Wikidata repository consists of items, each having a label and description
[6]. When compared with the Stony Brook repository, Wikidata provides a more generalised
description. For instance, the Convex Hull problem is described with metadata elements, such
as title, description, sub-class, and main category. The metadata like input, output, and related
problems are not provided. Wikidata has modeled the algorithms in a broad manner. The
metadata elements present in Wikidata are in a structured format. Wikidata provides query
service through the SPARQL endpoint4.</p>
      <p>Taking note of the data extraction for the KG creation, Wikidata provides a seamless service to
extract data in many formats like csv, json, xml, and many more. Wikidata provides structured
data by just executing SPARQL queries that do not fit the scope of our current work. Whereas,
the Stony Brook Algorithm repository challenges us to extract the unstructured data from its
website and process it to the set format. Algowiki does not meet our requirements as it describes
the algorithms minimally. It is more like a registry and not a repository as the information
is not held in the Algowiki server but spread across various third party websites. Whereas
in Stony Brook Repository apart from algorithmic problems various associated entities like
implementations and related problems are also described. Hence, in this work, our emphasis lies
on the Stony Brook Algorithm Repository for data extraction. Figure 1 shows a representational
algorithmic description from the repository.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Knowledge Graph Development Approach</title>
      <p>In this work, we build a KG for describing algorithmic problems and associated entities, such as
software implementations, recommended resources, persons, and so forth. We present here a
general approach toward the KG creation for algorithms and their relations. Figure 2 shows the
important steps that we follow.</p>
      <p>Step 0- Selection of data source. The initial step is to identify the objective data source for
data extraction, In section 2 comparisons are drawn among the accessible Algorithm Repositories.
The Stony Brook Algorithm Repository is recognized as the potential source to extract the data
for generating the KG.</p>
      <p>Step 1- Data extraction. The Stony Brook Algorithm Repository has 75 separate web pages
for its algorithms containing information like input, output, problem statement, description,
related problems, implementation, rating, and so on. There are 10 metadata elements about
each algorithm that interest us for extraction. Using an HTML parser and pattern matching the
metadata present in each webpage is extracted and stored in a convenient data structure. In
this work Python3, Dictionary and List are used as our preferred data structure.</p>
      <p>Step 2- Storing the data in tabular form. In the above step, the data is stored in a Python
variable but it needs to be exported to use it outside the Python environment. A Python package
called Pandas is used to export it in tabular form.</p>
      <p>Step 3- Data processing. The exported data file needs to be further processed in such a way
that KG transformation can be done easily. In the present data file, there are metadata elements
that have one to many relationships and are represented in a single column. These need to be
separated into multiple columns. Further, there are metadata elements that are combined, for
instance for an algorithm problem recommended books information column contains the book
title along with the author’s information as a single entity, which needs to be segregated into
multiple columns.</p>
      <p>Step 4- Knowledge Graph creation. The above data is taken as input and utilizing the
Algorithm Metadata Vocabulary (AMV) [7] and MappingMasterDSL [8] the data is transformed
into a KG. The details in regard to the KG development are discussed in section 5.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Data Extraction and Processing</title>
      <p>As discussed in section 2, for the current study, we selected Stony Brooks Algorithm Repository
as our target repository to extract data. The Repository categorizes the algorithmic problems into
two main classes: problems and language. On its home page (“https://algorist.com/algorist.html”),
all the algorithmic problems are listed. Each problem in the list points to an HTML page where
a detailed depiction of that algorithmic problem is present. Figure 3 shows a list of algorithmic
problems. Information about each algorithmic problem, e.g., Convex Hull, String Matching, and
Text Compression is to be extracted. The algorithmic problems are further grouped under seven
broad problem types, for example Combinatorial Problems, Graph Problems, and Computational
Geometry. Figure 3 shows a few of them.</p>
      <p>The process of extraction of information from the website is completely automated. To
automate the extraction and process the information in the required format the following
Python libraries were used: BeautifulSoup, Selenium, Urllib, Regex, and Pandas. Among these
BeautifulSoup along with Urllib is used to extract the relevant text from HTML pages. Webdriver
from the Selenium library is used for navigating to diferent HTML pages. Lastly, Pandas and
Regex are used to tabulate and clean the extracted data.</p>
      <sec id="sec-4-1">
        <title>4.1. Data Extraction</title>
        <p>In this section, the steps for extracting information from the Stony Brook Algorithm repository
are discussed. The pseudo-code for data extraction is given as Algorithm 1 (the source code is
available on GitHub 5).</p>
        <p>Due to the space problem, only two examples are provided. The extracted values as shown in
example 1 are appended to their corresponding keys in the dictionary created in step 3.</p>
        <p>Step 0- Importing the Python libraries. Started with importing the Python libraries, regex,
urllib, bs4, selenium and Pandas.</p>
        <p>Step 1- Links stored in a list. The selenium web driver is used to open the URL “https://algorist.
com/algorist.html” and all the links present on the homepage (as shown above in Figure 3 each
algorithmic problem is a link to an HTML page) are stored in a list (a Python data structure).</p>
        <p>Step 2- Filtering the URLs of algorithmic problems. In the above step all the URLs
present on the homepage are stored, but our interest lies in the URLs of each algorithmic
problem. Hence, an empty list ‘AlgorithmicProblem’ is created and all the URLs starting with
‘https://algorist.com/problems/’ are stored since the URL for all algorithmic problems starts
with the mentioned pattern. Regex is used to filter out the desired URLs.</p>
        <p>Step 3- Dictionary created for storing metadata. By visual inspection of the website it is
clear that each algorithmic problem has ten metadata elements. Hence, a Python dictionary
is created with keys: problem, problem_type, input_image, output_image, input_decription,
problem_statement, description, implementations, recommended_books and related_problems.
Initially all the keys are assigned an empty list.</p>
        <p>Step 4- Data population. This step uses links that were stored in step 1 to populate the
value in the dictionary corresponding to key problem_type. Regex is used to filter out the links
that begin with ‘https://www.algorist.com/sections’ and store it in a list. Since, all the broad
categories, e.g., Data Structures, Numerical Problems, and Combinatorial Problems, have URLs
starting with ‘https://www.algorist.com/sections’.</p>
        <p>Step 5- Metadata extraction. In this step, the metadata of each algorithmic problem is
extracted. BeautifulSoup ‘soup’(it contains HTML code in a hierarchical manner which is
easy to access) object is created which contains the HTML script of each algorithmic problem.
This is achieved by iterating over the list ‘AlgorithmicProblem’. For instance, the title of the
algorithmic problem is available in the ‘h1’ tag of the HTML page. Using the soup.select function
it is accessed, there are multiple implementations present and each has its ‘name’, ‘url’, and
‘rating’. This information is extracted as shown in example 1.</p>
        <p>Example 1. name1 | url1 | rating1 | implementation1_language_1 \n name2 | url2 | rating.</p>
        <sec id="sec-4-1-1">
          <title>5https://github.com/biswanathdutta/amv</title>
          <p>Step 6- Exporting the output as csv. A Pandas dataframe is created from the main dictionary
that holds all the information and this data frame is exported as a csv file.</p>
          <p>Algorithm for information extraction</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Data Processing</title>
        <p>The exported data in csv (as discussed above), needs to be processed and converted into a form
that fits with the Algorithm Metadata Vocabulary (AMV) model (discussed in section 5).The data
will be utilized to create a KG and in a KG each entity is represented in a unique node [9] [? ].In
the KG, a node can be any object, place, or person and the edge defines the relationship between
nodes [10]. In the exported data, some data points are merged together and do not represent a
unique entity. Hence, the processing is required for some columns e.g. implementation and
related problems.</p>
        <p>The columns like implementation, related problems, and recommended books have more
than one element. Figure 4 shows the columns implementation, recommended books, related
problems, and so on. As visible in the Figure 4 there is information that needs to be in diferent
columns. For instance, the implementation column has multiple entities. Firstly, each individual
implementation needs to be separated. Each individual implementation also has its name, rating,
and url, this information also needs to be split. Similar processing is required for the columns
recommended_books, related_problems, and implementation_in _languages.</p>
        <p>The details of processing the data are as follows:
Step 0- Importing the libraries and loading the data. Started by importing the Python
libraries, numpy, Pandas and regex and the csv file is loaded which contains the extracted data.</p>
        <p>Step 1- Processing the implementations. In this step, the focus is on the implementation
column, the multiple implementations were combined using ’\n’. The same is used to split each
implementation to get the maximum number of implementations present and create that many
columns with sufix (eg. Implementation_1, Implementation_2). Further, each implementation
is appended in separate columns that were created. Each implementations has its title, rating,
and url which are combined using ’|’. The same is used to split them and add them to separate
columns.</p>
        <p>Step 2- Processing the related problems. In this step, the focus is on related problem
column. There are multiple related problems combined with ‘/n’ and each problem has its name
and url combined using ‘|’. The similar approach as the previous step is followed to process this
column.</p>
        <p>Step 3- Processing the recommended books. In this step, the focus is on the recommended
book column, the book name contains the title of the book along with author names. The book
title and author’s name are separated by ‘by’ string. For books with multiple authors, each
author is separated by ‘and’ and ‘,’. Following the similar approach as above the books are split
into multiple columns and the related information like author name, book url are split from
each book.</p>
        <p>Step 4- Export the processed data. The final data frame is exported into a .xlsx file.</p>
        <p>Figure 5 shows the recommended_books column before and after processing. As visible in
Figure 5 after extraction, each recommended_book column has its title, url, and authors. After
processing, each recommended_book column is split and their authors, title and URLs are also
present in diferent columns.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Knowledge Graph Creation</title>
      <p>The processed data was received in a .xlsx file with 75 rows and 163 columns. This data is
taken as input for the KG development. In the current study, Algorithm Metadata Vocabulary
(AMV6) is used as a schema for the KG. MappingMaster ( 2) is used to transform the data
into a KG. AMV is a metadata vocabulary for describing algorithms, algorithm problems, and
related entities, like software code. The vocabulary is available as an OWL ontology. It can
be directly used by anyone interested to create and publish algorithm metadata as a KG, or to
provide metadata service through the SPARQL endpoint [? ]. We have used it as a schema for
the production of the algorithmic KG. MappingMaster (an open-source Java library to transform
the content of spreadsheets to OWL axioms) has GUI support and is available as a plugin called
Cellfie for Protege Desktop [ 11]. MappingMaster has domain-specific language (DSL) [ 8]. For
the current work, we used the Cellfie plugin and also the DSL language for developing mapping
rules for transforming the metadata available in a spreadsheet (as mentioned above) into KG. A
snippet of the developed mapping rules expressed in ( 2) DSL language is shown in Table 1.</p>
      <p>Table 1 provides a snippet of the mapping rule for the algorithm problems and their
corresponding types, identifier, representational depiction of input and output, input description for
the problem, problem description, and implementation details.</p>
      <p>The produced KG consists of 1494 individuals and 9706 axioms in addition to 59 classes, 44
object properties, and 50 data properties coming from the AMV ontology. The KG is available
as an RDF dump and can be downloaded from GitHub7.</p>
      <sec id="sec-5-1">
        <title>7https://github.com/biswanathdutta/amv</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. SPARQL Queries</title>
      <p>To validate the produced KG, we have conducted several SPARQL queries. The queries are
centered around the algorithm problems and their e.g. relations, implementation, and related
information like implementation language, platform, loop types, and data structure. Some of
the queries are: retrieve all the problems related to the sorting problem, with their
corresponding type and their implementation details like problem statement, implementation URI and
implementation language (Q1), retrieve the implementations of Eulerian Cycle problem in C++
programming language (Q2) and retrieve the related algorithms for text compression problem
along with their looping structure (Q3). Table 2 provides the SPARQL representation for the
ifrst query(Q1). Figure 7 displays the query result in a graph produced using a browser based
application Gruf (https://allegrograph.com/products/gruf/). The successful execution and the
retrieval of desired results for the various queries centered on algorithm problems proves the
eficacy of KG.</p>
      <sec id="sec-6-1">
        <title>PREFIX dct: &lt;http://purl.org/dc/terms/&gt; PREFIX amv: &lt;https://w3id.org/amv#&gt; PREFIX rdfs:&lt;http://www.w3.org/2000/01/rdfschema#&gt;</title>
        <p>SELECT DISTINCT ?problem ?type ?prob_desc ?y
?impl_uri ?impl_language
WHERE amv:Sorting dct:relation ?problem.
?problem a ?type ; amv:problemDescription
?prob_desc ; amv:hasImplementation ?y .
?y amv:inProgrammingLanguage
?impl_language; dct:identifier ?impl_uri.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Related Works</title>
      <p>Addressing all encompassing and factual knowledge utilizing RDF and Linked Data is quite
feasible [12]. There are industrial innovations like Thomson Reuters KG Feed for the Financial
Services market [12] and BBC’s KG for their operations with content [9]. Google’s KG to make
web searches more intelligent and augment the results with information relevant to the query
[13]. In academia, much work is focused on addressing the bibliographic metadata, while
machine readable portrayal of scientific information in academic writing has not received much
consideration [12]. There are few methodologies that focus on scientific literature. The Artificial
Intelligence Knowledge Graph(AI-KG) is a large-scale automatically generated KG that depicts
research entities. It uses deep learning techniques to extricate elements and relations from
scientific text [14].</p>
      <p>
        The Open Research Knowledge Graph (ORKG) [12] contributes toward representing scholarly
knowledge semantically with KGs. ORKG not only contains the bibliographic metadata like
authors, references, but also contains semantic depiction of scholarly literature like problem
statement, approach and implementation. Both AI-KG and ORKG utilize deep learning
techniques for extraction from the academic literature but KG of scientific software metadata focuses
on external code repositories, readme files and documentation of software [ 15]. It focuses on
metadata categories like description, installation instructions, execution, and citation for
extraction. Machine Learning techniques were employed to gather the data whereas in our work
we used pattern matching to gather data. For KG development a list of programme items is
scrapped from a target software registry (e.g., Zenodo). Then, for each item, its version data is
obtained, extract all code repository links, and download the complete text of its readme file.
SOMEF parses the readme file, and the findings are integrated and aggregated into a knowledge
Graph [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. OKG-Soft is also one efort towards creating a KG for scientific software metadata in
a machine readable manner. OKG-Soft includes an ontology designed to describe software and
the specific data formats it uses and publish software metadata as an open KG, linked to other
Web of Data object [16]. The capture of metadata is based on their previous work OntoSoft [17].
      </p>
      <p>The one ongoing efort towards KG is OpenAIRE, which considers many scientific artifacts like
research literature, research software, and research data. It also includes metadata records about
organizations involved in the research life-cycle, such as universities, research organizations, and
funders [18]. Graph4code is another work that focuses on the program code, the metadata in this
is extracted by the code documentation, forum discussions and then mapped into a knowledge
graph, this work employs extensive use of named graphs in RDF to make the knowledge graph
extensible [19]. The work focuses on and captures the semantics of Python codes whereas
our work focuses on algorithms that can be implemented in diferent programming languages.
Despite covering such diverse metadata records from various academic literature and external
sources, the idea of depicting the metadata of algorithms and related entities and their relations
is not considered previously. The current work may be considered a pioneer in this regard.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusion</title>
      <p>Algorithms published in the scientific literature are crucial to comprehend and reuse to better
understand the information. There is a surge in the number of research publications made
available each year. This creates a necessity to make algorithm searches more efective and
personalized. Algorithms should be treated as independent digital objects like research data,
research articles and in recent times software and ontologies [20]. In this work, we have
presented a novel approach for automatically creating a KG by extracting the data of 75 diferent
algorithmic problems along with their relations and a methodology from the textual resource
like Stony Brook algorithm repository. The work explores new metadata categories related to
algorithmic problems to better comprehend problems and reuse. Such categories are related
problems (a similar algorithmic problem) and recommended books (to better understand the
problem and get detailed explanations). In continuation to the current work, we aim to make
a comparative analysis of the designed KG approach with state of the art related approaches
as used in the creation of knowledge graphs, such as AI-KG and ORKG. The present KG was
developed primarily based on a single resource i.e., the Stony Brook algorithm repository. In
the future, we aim to extend the KG by extracting the information from several other sources,
such as scientific literature, repositories (e.g., GitHub, NIST Dictionary of Algorithms and Data
Structures [21]), and online discussion forums. For this purpose, we aim to focus on developing
a more generic and robust framework of information extraction.
[5] M. DeBellis, B. Dutta, The covid-19 codo development process: an agile approach to
knowledge graph development, in: B. Villazón-Terrazas, F. Ortiz-Rodríguez, S. Tiwari, A. Goyal,
M. Jabbar (Eds.), Knowledge Graphs and Semantic Web, volume 1459 of Communications
in Computer and Information Science, Springer International Publishing, Cham, 2021, p.
153–168. doi:10.1007/978-3-030-91305-2_12.
[6] Wikidata, 2012. URL: https://www.wikidata.org/wiki/Wikidata:Main_Page.
[7] B. Dutta, J. Patel, AMV: A representational model and Metadata Vocabulary for describing
and maintaing Algorithms, Journal of Information Science (2022).
[8] M. J. O’Connor, C. Halaschek-Wiener, M. A. Musen, Mapping master: A flexible approach
for mapping spreadsheets to owl, in: P. F. Patel-Schneider, Y. Pan, P. Hitzler, P. Mika,
L. Zhang, J. Z. Pan, I. Horrocks, B. Glimm (Eds.), The Semantic Web – ISWC 2010, Springer
Berlin Heidelberg, 2010, p. 194–208. doi:10.1007/978-3-642-17749-1_13.
[9] T. Petkova, The knowledge graph and the enterprise, 2018. URL: https://www.ontotext.</p>
      <p>com/blog/the-knowledge-graph-and-the-enterprise/.
[10] What is a knowledge graph, 2021. URL: https://www.ibm.com/cloud/learn/
knowledge-graph.
[11] M. A. Musen, The protégé project: a look back and a look forward, AI Matters 1 (2015)
4–12. doi:10.1145/2757001.2757003.
[12] S. Auer, M. Stocker, L. Vogt, G. Fraumann, A. Garatzogianni, Orkg: Facilitating the transfer
of research results with the open research knowledge graph, Research Ideas and Outcomes
7 (2021) e68513. doi:10.3897/rio.7.e68513.
[13] C. Harding, Semantic data platforms come of age, 2021. URL: https://www.linkedin.com/
pulse/semantic-data-platforms-come-age-chris-harding/.
[14] J. Tsay, A. Braz, M. Hirzel, A. Shinnar, T. Mummert, Aimmx: Artificial intelligence model
metadata extractor, in: Proceedings of the 17th International Conference on Mining
Software Repositories, ACM, 2020, p. 81–92. URL: https://dl.acm.org/doi/10.1145/3379597.
3387448. doi:10.1145/3379597.3387448.
[15] A. Mao, D. Garijo, S. Fakhraei, Somef: A framework for capturing scientific software
metadata from its documentation, in: 2019 IEEE International Conference on Big Data
(Big Data), 2019, pp. 3032–3037. doi:10.1109/BigData47090.2019.9006447.
[16] D. Garijo, M. Osorio, D. Khider, V. Ratnakar, Y. Gil, Okg-soft: An open knowledge graph
with machine readable scientific software metadata, in: 2019 15th International Conference
on eScience (eScience), 2019, pp. 349–358. doi:10.1109/eScience.2019.00046.
[17] Y. Gil, V. Ratnakar, D. Garijo, Ontosoft: Capturing scientific software metadata, in:
Proceedings of the 8th International Conference on Knowledge Capture, K-CAP 2015,
Association for Computing Machinery, New York, USA, 2015, p. 1–4. URL: https://doi.org/
10.1145/2815833.2816955. doi:10.1145/2815833.2816955.
[18] A.-L. Lamprecht, L. Castro, M. Kuzak, C. Martinez Ortiz, R. Arcila, E. Pico, V. Dominguez
Del Angel, S. Sandt, J. Ison, P. Martinez, P. McQuilton, A. Valencia, J. Harrow, F.
Psomopoulos, J. Gelpi, N. Chue Hong, C. Goble, S. Capella-Gutierrez, Towards fair principles for
research software, Data Science (2019).
[19] K. Srinivas, I. Abdelaziz, J. Dolby, J. McCusker, Graph4code: A machine interpretable
knowledge graph for code (2020).
[20] B. Dutta, A. Toulet, V. Emonet, C. Jonquet, New generation metadata vocabulary for
ontology description and publication, in: E. Garoufallou, S. Virkus, R. Siatri, D. Koutsomiha
(Eds.), Metadata and Semantic Research, Springer International Publishing, Cham, 2017, p.
173–185.
[21] P. E. Black, Dictionary of algorithms and data structures, 1998. URL: http://www.nist.gov/
dads.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kelley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garijo</surname>
          </string-name>
          ,
          <article-title>A framework for creating knowledge graphs of scientific software metadata</article-title>
          ,
          <source>Quantitative Science Studies</source>
          <volume>2</volume>
          (
          <year>2021</year>
          )
          <fpage>1423</fpage>
          -
          <lpage>1446</lpage>
          . URL: https://doi.org/10.1162/ qss_a_00167. doi:
          <volume>10</volume>
          .1162/qss_a_
          <fpage>00167</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Skiena</surname>
          </string-name>
          ,
          <article-title>Who is interested in algorithms and why? lessons from the stony brook algorithms repository</article-title>
          ,
          <source>ACM SIGACT News</source>
          <volume>30</volume>
          (
          <year>1999</year>
          )
          <fpage>65</fpage>
          -
          <lpage>74</lpage>
          . doi:
          <volume>10</volume>
          .1145/333623.333627.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Voevodin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Antonov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dongarra</surname>
          </string-name>
          ,
          <article-title>Why is it hard to describe properties of algorithms?</article-title>
          ,
          <source>Procedia Computer Science</source>
          <volume>101</volume>
          (
          <year>2016</year>
          )
          <fpage>4</fpage>
          -
          <lpage>7</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.procs.
          <year>2016</year>
          .
          <volume>11</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K. U.</given-names>
            <surname>Idehen</surname>
          </string-name>
          ,
          <article-title>Linked data, ontologies, and knowledge graphs, 2020</article-title>
          . URL: https://www. linkedin.com/pulse/linked
          <article-title>-data-ontologies-knowledge-graphs-kingsley-uyi-idehen/.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>