<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Knowledge Graph Based Services in Accounting Use Cases</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Schulze</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michelle Pelzer</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Markus Schröder</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Jilek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heiko Maus</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Dengel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department, Technische Universität Kaiserslautern</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Smart Data &amp; Knowledge Services Department, Deutsches Forschungszentrum für Künstliche Intelligenz GmbH</institution>
          ,
          <addr-line>DFKI</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>envia Mitteldeutsche Energie AG</institution>
          ,
          <addr-line>Chemnitz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>To assist knowledge work in accounting use cases such as bookkeeping, this paper presents a pipeline for constructing and enriching an accounting knowledge graph from heterogeneous accounting resources. To show the feasibility of the approach, we applied the pipeline in a multi-group energy provider by employing real company data. A set of prototypical knowledge services was realized with the accounting knowledge graph as the basis, for example, the suggestion of similar accounting cases to the accountant. For training decision trees to predict accounts, our results suggest that using semantically enriched data from the knowledge graph leads to better results compared to not using semantically enriched data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The work of accountants is characterized by searching relevant information which is distributed
in diferent company sources. One typical task is bookkeeping where besides verifying an invoice,
decisions are made about the set of accounts and VAT rates that need to be selected for a particular
accounting case. For solving such tasks, accountants typically tap into several data sources,
for example, chart of accounts, accounting manuals and handbooks, or accounting policies.
In multi-group companies, accounting departments are usually responsible for all subsidiary
groups which means that they have to consider such documents for each group separately.
Further common sources are personal notes, referenced documents on an invoice, historical
accounting cases as references, additional attachments, or the knowledge of a colleague.</p>
      <p>
        Tapping into this many data sources can lead to cases where the time needed to find the
relevant information is high, especially for non-standard cases. Jain and Woodcock [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] predict
that also in future, 21 % of tasks that deal with invoice processing are not suited for automation.
Therefore, we conclude that knowledge workers will be still important for such tasks. With our
research, we aim to assist knowledge workers in such settings with an "Information Butler"
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] paradigm emphasizing the importance of context in which a knowledge worker is situated
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Also, by following this idea and adopting it to accounting scenarios, our goal is to embed
knowledge graph-based services directly into the work context of an accountant, for example,
when the accountant opens up an invoice for processing. Because we cooperated with a
multigroup energy provider, we were able to base this research on real company data. There, the
accounting department was interviewed to derive requirements from a knowledge worker’s view:
a) a suitable approach should cover food and beverage scenarios in particular, such as eating
at restaurants or catering, because such cases are, compared to other scenarios, particularly
error-prone, b) historic accounting cases should be suggested as an assistance for using them
as references; c) accounts and VAT rates that may fit to the current work context should be
displayed with detailed information; and d) suitable co-workers that may have experience in
the current accounting case should be suggested.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Use cases around the Financial Industry Business Ontology (FIBO) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] deal with integrating
ifnancial information and providing knowledge services on top of it. However, reasonably, such
services often have a strong and useful analytical rationale to help with decision making. Work
and knowledge graph based services in the context of public procurement [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] aim to assist the
knowledge worker. Compared to accounting, the procurement domain is more upstream but it
may be useful to interlink such data for downstream tasks.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Approach</title>
      <p>Because the envisioned knowledge services require upfront queries over information that resides
in the diferent introduced data sources, which have diferent formats as well such as PDF or
CSV, and because the work context of an accountant is important but nowhere modeled and
displayed in the task process explicitly, we decided to first construct an accounting knowledge
graph and then to implement the knowledge services on top of it.</p>
      <sec id="sec-3-1">
        <title>3.1. Accounting Knowledge Graph Construction and Vocabulary</title>
        <p>Figure 1 depicts the pipeline of constructing and enriching an accounting knowledge graph. The
data basis was a CSV file containing 65k processes from the year 2019. It contained information
such as the invoice issuer, invoice issuer number, booking area, the service description, or the
amount of money. According to requirement a), we extracted at first all processes that deal with
food and beverage scenarios. For this, two approaches were followed: On the one hand, with
a keyword list, we mimicked the approach an accountant would use by searching the service
descriptions of invoices for keywords such as "secretariat service" or "restaurant". On the other
hand, we filtered processes which were booked on standard food and beverage accounts where
food and beverage cases are booked exclusively. By merging the results, 1267 food and beverage
processes have been detected which represent the first input for the knowledge graph.</p>
        <p>
          To construct the first version of the knowledge graph (KG v1 in Figure 1), another CSV file
containing a set of booking statements for each process has been considered. Such booking
statements contain, among others, an account, a VAT rate identifier and the accordingly split
amount of money. Technology-wise, we used RML [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] with rmlmapper1 in version 5.0.0.
Ontology-wise, we used the P2P-O ontology [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] as a basis where vocabulary is provided for
describing electronic invoices, such as the invoice issuer (seller). To account for the need to
semantically describe and connect booking statements and the content of accounting handbooks,
we developed a bookkeeping ontology which aims to extend P2P-O. To the best of our knowledge,
there is no ontology available covering such bookkeeping vocabulary, and thus fitting our exact
requirements2. However, we could reuse vocabulary from the FIBO Ontology [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], such as for the
ledger account and account identifier. The bookkeeping-ontology is documented with WIDOCO
[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and published at https://purl.org/p2p-o/bkk. It further has a business friendly license to
enable reuse. For validation, we checked the consistency and iteratively evaluated the ontology
with OOPS! (OntOlogy Pitfall Scanner!) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Semantic Enrichment</title>
        <p>For the second version of the knowledge graph (KG v2 in Figure 1), the type of invoice issuer
was refined as well as the service description. By means of keyword matching, triples to the
invoice issuer have been automatically added which state whether the issuer is, for example,
a restaurant, a hotel, or a caterer. In the same way, triples have been added to the service
description to specify whether the invoice is, for example, about a breakfast, a business lunch, or
a meeting service. For a full list of enrichments, we would like to refer to the ontology page. To
reach KG v3 in Figure 1, we incorporated information about accounts which resides in multiple
PDF handbooks for each company group separately. Accordingly, per account, information
about the account ID, short- and long descriptions as well as the provenance in form of the
accounting handbook were incorporated into the knowledge graph. Because the accounting
handbooks were all structured identically starting with the account identifier per account, and
by exploiting this structure with regular expressions, the handbook text could be processed
automatically. To enable the suggestions of colleagues (Figure 1, KG v4), we firstly included all
1https://github.com/RMLio/rmlmapper-java
2By searching in ontology repositories such as https://lov.linkeddata.es/dataset/lov/
accountants as person resources with their contact data. For linking accountants to historical
processes, historic process protocols have been analyzed to find out which accountants were
involved in a process (also by employing regular expressions).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Knowledge Graph Based Services and Prediction of Accounts</title>
      <p>
        Because the goal was to embed knowledge services directly into the workspace of an accountant,
information is shown in form of a sidebar on the desktop. The screenshot in Figure 2 shows an
excerpt of the suggestions of similar accounts. In this way, an accountant can look into such
processes to get clues for processing the current accounting case. For calculating the similarity
between processes, we leveraged the semantically enriched information about types of invoice
issuer and services. An insight from interviews was that the similarity indicators needed to be
weighted diferently so that the booking area is more important than the invoice issuer and the
invoice issuer is more important than the type of service. Therefore, initially, a ratio of 4:2:1:1
was adopted (where the additional "1" represents the type of invoice issuer). Because during
enrichment accountants have been linked to historic processes, the similarity matrix could also
be used for suggesting colleagues. For account and VAT rate prediction, two approaches have
been followed: First, based on the similarity matrix, we suggest the accounts and VAT rates
of similar processes. Second, Table 1 summaries results for learning decision trees by using
diferent degrees of semantically enriched data: using non-enriched data and using semantically
enriched data with (I) the type of the invoice issuer, (II ) the type of the service, and (III) with the
combination of both. WEKA [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] was used because it supports categorical data [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and results
are obtained by employing 10-fold cross validation. First results in Table 1 suggest that adding
and in particular combining a few features obtained from a semantically enriched knowledge
graph can lead to better results.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this paper, we presented a pipeline for constructing and enriching an accounting knowledge
graph for bookkeeping use cases. To enable reuse, we further proposed required bookkeeping
vocabulary. Besides showing the feasibility of realizing knowledge services on top of such an
accounting knowledge graph, first results indicate that using semantically enriched data leads
to better results when learning decision trees to predict accounts. One lane for future work is
the ofering of explanations in form of derived rules.</p>
      <p>Acknowledgements: This work was funded by BMBF project SensAI (grantno. 01IW20007).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Woodcock</surname>
          </string-name>
          ,
          <article-title>A road map for digitizing source-to-</article-title>
          <string-name>
            <surname>pay</surname>
          </string-name>
          ,
          <year>2017</year>
          . URL: https://www.mckinsey.
          <article-title>com/business-functions/operations/our-insights/ a-road-map-for-digitizing-source-to-</article-title>
          <string-name>
            <surname>pay</surname>
          </string-name>
          ,
          <source>Last accessed 12 Jul</source>
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dengel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Maus</surname>
          </string-name>
          , Personalisierte Wissensdienste:
          <article-title>Das Unternehmen denkt mit, IM+io Fachmagazin 3 (</article-title>
          <year>2018</year>
          )
          <fpage>46</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Jilek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schröder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schwarz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Maus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dengel</surname>
          </string-name>
          ,
          <article-title>Context spaces as the cornerstone of a near-transparent and self-reorganizing semantic desktop, in: The Semantic Web: ESWC 2018 Satellite Events</article-title>
          , volume
          <volume>11155</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bennett</surname>
          </string-name>
          ,
          <article-title>The financial industry business ontology: Best practice for big data</article-title>
          ,
          <source>Journal of Banking Regulation</source>
          <volume>14</volume>
          (
          <year>2013</year>
          )
          <fpage>255</fpage>
          -
          <lpage>268</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Soylu</surname>
          </string-name>
          , Ó. Corcho,
          <string-name>
            <given-names>B.</given-names>
            <surname>Elvesaeter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Badenes-Olmedo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Blount</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. Y.</given-names>
            <surname>Martínez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kovacic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Posinkovic</surname>
          </string-name>
          , I. Makgill,
          <string-name>
            <given-names>C.</given-names>
            <surname>Taggart</surname>
          </string-name>
          , E. Simperl,
          <string-name>
            <given-names>T. C.</given-names>
            <surname>Lech</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roman</surname>
          </string-name>
          ,
          <article-title>Theybuyforyou platform and knowledge graph: Expanding horizons in public procurement with open linked data</article-title>
          ,
          <source>Semantic Web</source>
          <volume>13</volume>
          (
          <year>2022</year>
          )
          <fpage>265</fpage>
          -
          <lpage>291</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dimou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Sande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Colpaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Verborgh</surname>
          </string-name>
          , E. Mannens, R. V. de Walle,
          <article-title>Rml: A generic language for integrated rdf mappings of heterogeneous data</article-title>
          ,
          <source>in: Proc. of the Workshop on Linked Data on the Web</source>
          , volume
          <volume>1184</volume>
          <source>of CEUR Workshop Proc., CEUR-WS.org</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schulze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schröder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jilek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Albers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Maus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dengel</surname>
          </string-name>
          ,
          <article-title>P2P-O: A purchase-to-pay ontology for enabling semantic invoices</article-title>
          ,
          <source>in: The Semantic Web - 18th International Conference, ESWC</source>
          <year>2021</year>
          , volume
          <volume>12731</volume>
          <source>of LNCS</source>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>647</fpage>
          -
          <lpage>663</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Garijo</surname>
          </string-name>
          ,
          <article-title>Widoco: a wizard for documenting ontologies</article-title>
          , in: International Semantic Web Conference, Springer, Cham,
          <year>2017</year>
          , pp.
          <fpage>94</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Poveda-Villalón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gómez-Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Suárez-Figueroa</surname>
          </string-name>
          ,
          <article-title>Oops! (ontology pitfall scanner!): An on-line tool for ontology evaluation</article-title>
          ,
          <source>International Journal on Semantic Web and Information Systems (IJSWIS) 10</source>
          (
          <year>2014</year>
          )
          <fpage>7</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          ,
          <article-title>The WEKA workbench</article-title>
          .,
          <source>in: Online Appendix for "Data Mining: Practical Machine Learning Tools and Techniques"</source>
          ,
          <source>Fourth Edition</source>
          , Morgan Kaufmann,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>