<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Neurosymbolic Approach to Fraud Detection on Financial Data in the Public Administration.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michele Vitale</string-name>
          <email>michele.vitale@unical.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Neurosymbolic AI, Fraud Detection, Knowledge Graph, Public Administration</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Matematics and Computer Science, DEMACS, University of Calabria Rende</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <fpage>9</fpage>
      <lpage>13</lpage>
      <abstract>
        <p>This doctoral research, in collaboration with Invitalia, the Italian National Development Agency, focuses on using a neurosymbolic approach to detect fraud in public administration. The main goal is to create a decision support system that can both identify past fraudulent activities and proactively flag suspicious requests for human review. The project addresses the challenges of limited labeled data and the need for interpretability in AI models, which is crucial in the public investment field. The initial phase involved creating a data lake from public funding documents using Large Language Models (LLMs). This data will be used to construct a knowledge graph to identify fraudulent patterns. The research proposes using neurosymbolic approaches to overcome existing challenges. Symbolic AI methods will be used for transparent dataset labeling by incorporating expert domain knowledge and legal regulations. Additionally, a neural ensemble approach is proposed where individual classifiers for specific fraud indicators are built, and their outputs are combined using a symbolic program. The author is seeking guidance on integrating these neurosymbolic paradigms into their research.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>My research work focuses mainly on Fraud Detection. The project is carried out in cooperation with</p>
      <p>CEUR
Workshop</p>
      <p>ISSN1613-0073</p>
    </sec>
    <sec id="sec-2">
      <title>2. Project Overview</title>
      <p>The project goal is to build a complete decision support system that can help domain experts, mainly
with economic and legal backgrounds, to analyze and identify frauds and to report potential fraudulent
intents in new fund requests. A simple schema can be seen in Figure 1.</p>
      <p>To achieve this, the initial step is to build a data collection from all instances of funding requests and
revocations directed to Invitalia, the Italian National Development Agency. A revocation occurs when,
following thorough checks, Invitalia identifies irregularities in an entity’s acquisition of funds, after
that entity had previously received financing. Revocations can stem from numerous reasons, including
but not limited to: non-compliance with conditions, failure to submit crucial documentation, deviation
from the spending plan, non-payment of installments for loans, or ongoing criminal proceedings. Thus,
it is important to note that a revocation do not necessarily imply that the beneficiary has fraudulent
purposes: a variety of diferent causes might lead the agency to retire the funds, such as unmet deadlines
or loss of eligibility for funding. Furthermore, not every fraud has been identified because the manual
verification process is long and complex. Therefore, a fraud might not have an associated revocation
document, but only an approval and access to funds. It is clear that the main step is the identification of
a fraud, based on the data that have been collected in the internal protocol and, thus, in the data lake
previously built.</p>
      <p>The next step involves creating a knowledge graph from the extracted data, integrating it with financial
data collections from paid databases. Based on this knowledge graph, a deep learning model needs
to be built to identify recurring patterns in organizations prone to fraud. Indeed, these organizations
are often identifiable by various components, such as a network of interconnected people or entities,
or a recurring pattern in transaction amounts. Another strategy could involve calculating similarity
metrics with previously recognized frauds. This module will be the primary focus of my contribution
and, consequently, the core of my future research.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Data and Knowledge Extraction</title>
      <p>The first part of the project, which is already complete, involved developing a knowledge extraction
pipeline. Its goal was to create a data lake containing all relevant information from public funding
approval and revocation documents granted to third parties. We achieved this by processing a large set
of PDF documents resulting from the period 2018 - 2025 protocol registrations, and using LLMs (Large
Language Models) for querying the files content.</p>
      <p>
        A general overview of the architecture can be seen in Figure 2. The module implements a pipeline flow,
in which both old and new documents are processed in the I/O Documents Handler object, that via a
set of API calls to the Data Query Server can retrieve the predefined set of information that are relevant
to the construction of the Data Lake. The Document Loader takes advantage of the caching mechanism
built in the project to reduce the number of tokens used for each processed document, thus to reduce
the cost of the infrastructure. After many diferent strategies that have been tested, I ended up choosing
the one that has a main instruction prompt that defines the overall context, the background, and the
general instructions that the agent must keep clear and one single and straightforward prompt for each
of the fields that we want to populate in the final tabular structure in which the files are stored. This
drastically reduced the LLM misinterpretation of data and requests, achieving an high level of data
quality metrics, verified manually on a consistent set of about 100 documents. The prompt engineering
also includes zero-shot, few-shot techniques, with refinements for each single piece of data that must
be extracted[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The data lake will be expanded with external data from paid databases as sources and represented on a
knowledge graph, built tracking all the relations among the documents, the involved entities (people,
business), and the presented projects.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Anomaly and Fraud Detection</title>
      <p>Anomaly detection and fraud detection are closely related fields within data science and machine
learning, both aiming to identify unusual patterns or behaviors that deviate significantly from what is
considered ”normal”. Although anomaly detection is a broader concept, fraud detection is a specific
application of anomaly detection in the context of criminal deception.</p>
      <p>
        Anomaly detection, also known as outlier detection, is the process of identifying data points, events, or
observations that do not conform to an expected pattern or other items in a dataset. These deviations
are often indicative of some underlying problem or rare event that requires further investigation.
Fraud detection is a specialized application of anomaly detection aimed at identifying and preventing
deceptive activities undertaken to gain illegal financial or other benefits. It focuses on recognizing
patterns, anomalies, or suspicious activities in transactional or behavioral data that indicate fraudulent
intent [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Neurosymbolic Approaches</title>
      <p>In the neural approach to fraud detection, a significant challenge is the lack of an empirical, readily
interpretable method to identify fraud. Given the inherent complexity of the domain, building a neural
classifier from scratch can be prohibitively expensive and ultimately inefective. This is because neural
networks often act as ”black boxes,” making it dificult to understand why a particular decision was
made, which is crucial in regulated fields like financial fraud detection. Furthermore, they typically
require vast amounts of labeled data, that not only are not present in the domain, but also are often
scarce and imbalanced for rare events like fraud.</p>
      <p>Two Neurosymbolic approaches are presented in the following sections, with the goal of my application
for the Doctoral Consortium being to have the important chance to explore other possible Neurosymbolic
approaches to fulfill the task and also find some new research streams to expand in the field.</p>
      <sec id="sec-5-1">
        <title>5.1. Dataset Construction and Labeling</title>
        <p>In this context, symbolic approaches can contribute significantly to solving the task of labeling the data
already collected in the knowledge graph to build a labeled dataset. Unlike neural networks, symbolic
AI methods ofer transparency and explainability. They allow us to explicitly define and reason about
the patterns and relationships indicative of fraudulent activity. This means that we can encode domain
expertise, legal regulations, and known fraud indicators directly into the system. Furthermore, it can
give an explanation on the model’s behaviour, starting from the domain knowledge given by experts.
For instance, an example of domain-expert description might be:</p>
        <p>If the financing originates from more than two funds and the aggregate amount is greater
than €20,000, then it is a potential fraud.</p>
        <p>
          Starting from the description, some simple rules with the Answer Set Programming formalism [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] might
be inferred, as follows.
        </p>
        <p>at_least_3_funds :- #count{A : fund(A,I)} &gt; 2.
total_over_20k :- #sum{I,A : fund(A,I)} &gt; 19999.</p>
        <p>potential_fraud :- at_least_3_funds, total_over_20k.</p>
        <p>The predicate potential_fraud can also be expressed in a probabilistic manner, with weighted atoms
based on the seriousness of its semantic meaning in the domain.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Neural Ensemble</title>
        <p>
          The task of fraud detection is large and complex, including several subtasks that naturally emerge from
decomposing the initial domain. A fraud is a dynamic entity, potentially composed of diverse elements
such as collusion among a group of individuals, specific patterns within transaction amounts, or a
project proposal that includes certain scopes or elements. Each one of these subtask has an impact on
the final classification, asserting whether a fund application is a fraud or not. A neural model might not
be able to leverage the importance of each of the sub-components that, put together, represent a fraud.
An interesting approach to this problem could involve identifying each potential fraud indicator, with
the help of domain experts. Instead of building a single and monolithic model that tries to ingest every
information in the dataset, we could construct a series of individual classifiers for each specific subtask.
The results could then be combined at the end using a symbolic program that considers the output of
each model, in an approach very similar to semantic loss [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Such an ensemble might be more capable
to adapt to the patterns of the data, with each task having a precise weight defined from the practical
knowledge of the domain given by experts. A simple schema of this idea can be seen in Figure 3.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This paper presented my doctoral research on fraud detection within public funds, specifically focusing
on the decision support system being developed for Invitalia, the Italian National Development Agency.</p>
      <p>We highlighted the project’s dual aim: to both identify past fraudulent activities and proactively flag
suspicious requests for human review, recognizing the complex, multi-faceted nature of fraud.</p>
      <p>My research begins with building a rich knowledge graph from Invitalia’s funding data, augmented
by external financial databases. This will constitute the dataset for a deep learning model designed to
pinpoint recurring fraudulent patterns, a module central to my future research contribution.</p>
      <p>A key challenge lies in the interpretability of AI models and the scarcity of labeled data in this sensitive
domain. I am particularly enthusiastic about exploring neurosymbolic AI approaches to address these
issues. Symbolic methods can provide transparent means for dataset labeling by incorporating expert
domain knowledge, while a neural ensemble approach could ofer greater adaptability and explainability
by combining task-specific classifiers.</p>
      <p>My motivation for attending this Doctoral Consortium is to gain crucial insights and guidance on
these neurosymbolic avenues. As my project is still in its early stages, the consortium’s interdisciplinary
discussions and expert feedback will be invaluable in shaping my research direction and efectively
integrating neurosymbolic paradigms into my work on Fraud Detection.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>The author has not employed any Generative AI tools.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Boonstra</surname>
          </string-name>
          , Prompt engineering,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Carneiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Veloso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ventura</surname>
          </string-name>
          , G. Palumbo,
          <string-name>
            <given-names>J.</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <article-title>Network Analysis for Fraud Detection in Portuguese Public Procurement</article-title>
          ,
          <year>2020</year>
          , pp.
          <fpage>390</fpage>
          -
          <lpage>401</lpage>
          . doi:
          <volume>10</volume>
          .1007/978- 3-
          <fpage>030</fpage>
          - 62365- 4_
          <fpage>37</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Brewka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Eiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Truszczyński</surname>
          </string-name>
          ,
          <article-title>Answer set programming at a glance</article-title>
          ,
          <source>Commun. ACM</source>
          <volume>54</volume>
          (
          <year>2011</year>
          )
          <fpage>92</fpage>
          -
          <lpage>103</lpage>
          . URL: https://doi.org/10.1145/2043174.2043195. doi:
          <volume>10</volume>
          .1145/2043174.2043195.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , T. Friedman,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>V. den Broeck, A semantic loss function for deep learning with symbolic knowledge, 2018</article-title>
          . URL: https://arxiv.org/abs/1711.11157. arXiv:
          <volume>1711</volume>
          .
          <fpage>11157</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>