<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>International Conference on Neural In- weight initialization method for sigmoidal feedfor-
formation Processing Systems, NIPS'</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1938-7228</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1016/j.neucom.2016.05</article-id>
      <title-group>
        <article-title>Knowledge Injection via ML-based Initialization of Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lars Hofmann</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Bartelt</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heiner Stuckenschmidt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Mannheim, Chair of Artificial Intelligence</institution>
          ,
          <addr-line>B6, 26, 68131 Mannheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Mannheim, Institute for Enterprise Systems</institution>
          ,
          <addr-line>L15, 1-6, 68131 Mannheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>207</volume>
      <issue>2016</issue>
      <fpage>676</fpage>
      <lpage>683</lpage>
      <abstract>
        <p>Despite the success of artificial neural networks (ANNs) for various complex tasks, their performance and training duration heavily rely on several factors. In many application domains these requirements, such as high data volume and quality, are not satisfied. To tackle this issue, diferent ways to inject existing domain knowledge into the ANN generation provided promising results. However, the initialization of ANNs is mostly overlooked in this paradigm and remains an important scientific challenge. In this paper, we present a machine learning framework enabling an ANN to perform a semantic mapping from a well-defined, symbolic representation of domain knowledge to weights and biases of an ANN in a specified architecture.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Knowledge Injection</kwd>
        <kwd>Neural Networks</kwd>
        <kwd>Initialization</kwd>
        <kwd>Machine Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>at local minima or saddle points [8].</p>
      <p>Consequently, there is a great variety of approaches in
Despite the substantial achievements of artificial neural this field focusing on diferent elements within the ANN
networks (ANNs) driven by high generalization capa- generation process. One prominent category targets the
bilities, flexibility, and robustness, the training duration learning process by adding domain-specific constraints
and performance still highly depend on several factors, or loss terms to the cost function, such as [9], [10], [11]
such as the network architecture, the loss function, the and [12]. However, there is little to no research on
initialinitialization method, and most importantly on the avail- izing the weights and biases of an ANN based on domain
able training data. However, in many real-world appli- knowledge. Such knowledge can act as a pointer towards
cations, e.g., in safety-critical systems, there are various a promising starting point in the optimization landscape.
issues regarding data collection and generation. In these The resulting “warm start” of the learning process can
scenarios, the capabilities of exclusively data-oriented reduce the required training time as well as improve the
approaches to train an ANN are limited. overall performance. This efect shall be exploited
efi</p>
      <p>To tackle these domain-specific challenges, the concept ciently by the framework presented in this paper.
of integrating or injecting existing domain knowledge With this goal in mind, the existing collection of
netinto the generation process of ANNs becomes increas- work initialization techniques were analyzed. Aguirre
ingly attractive in research and practice, indicated by sev- and Fuentes [13] define three groups. “Data-independent”
eral survey papers for machine learning (e.g., [1, 2, 3, 4]) methods are based on randomly drawing samples of
difas well as deep learning in particular (e.g., [5, 6, 7]). ferent distributions, e.g., LeCun [14], Xavier [15] and He
Furthermore, this paradigm bares the potential to miti- [16]. “Data-dependent” approaches, such as WIPE [17],
gate general weaknesses of ANNs, like slow convergence LSUV [18] and MIWI [19], additionally take statistical
speed, high data demands and the risk of getting stuck properties of the available training data into account.
Approaches within the third group, like [20, 21, 22, 23, 24],
apply the concept of “pre-training“. Their goal is to learn
an ANN on a related problem (with suficient availability
of high-quality data) and use it as an initialization for the
primary task. Consequently, they are not limited to the
actual training data, and thus to some extent
independent to task-specific data issues. Although, none of these
approaches explicitly considers domain knowledge, they
could be adapted to pre-train an ANN on synthetic data
encoding domain knowledge; also proposed by Karpatne
et al. [4]. These data can be generated, for instance, by
simulations or querying a domain model. But this comes
KINN@CIKM’21: Proceedings of CIKM Workshop on Knowledge
Injection in Neural Networks, November 1, 2021, Online Virtual Event
" hofmann@es.uni-mannheim.de (L. Hofmann);
bartelt@es.uni-mannheim.de (C. Bartelt);
heiner@informatik.uni-mannheim.de (H. Stuckenschmidt)
~ https://www.uni-mannheim.de/ines/ueber-uns/
wissenschaftliche-mitarbeiter/lars-hofmann (L. Hofmann);
https://www.uni-mannheim.de/ines/ueber-uns/
wissenschaftliche-mitarbeiter/dr-christian-bartelt (C. Bartelt);
https://www.uni-mannheim.de/dws/people/professors/
prof-dr-heiner-stuckenschmidt/ (H. Stuckenschmidt)</p>
      <p>0000-0002-9667-0310 (L. Hofmann); 0000-0003-0426-6714
(C. Bartelt); 0000-0002-0209-3859 (H. Stuckenschmidt)</p>
      <p>© 2021 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org)
with several disadvantages. First, the data samples must 2.1.  -Net Generation
be eficiently generated to fully represent the domain
knowledge, and second, the ANN must be able to learn Before a given task-specific function approximation, also
the contained knowledge. On top of this potential loss referred to as domain model, can be injected, a suitable
of information, this entire process must be repeated ev-  -Net must be generated once with the proposed ML
framework. This can be done completely with synthetic
data. The overall objective is to maximize the
transformation fidelity between the input function and the predicted
ANN. A schematic overview of the framework consisting
of three main steps is shown in Figure 1.
ery time the domain knowledge or the target network
architecture changes.</p>
      <p>Replacing this indirect, data-based knowledge transfer
from an already existing domain model into an ANN with
a direct mapping or transformation can potentially solve
these problems. Although they do not relate to
knowledge injection, several authors engineered explicit map- 2.1.1. Algebra Selection and  -Function
ping algorithms for diferent representations. A promi- Generation
nent example are Decision Trees (DTs), because they At first, a diverse set of functions Λ in the same algebra
also have a graph-based structure. Early work was al- as the given task-specific domain model is created. If
ready performed in the 1990s (e.g., [25, 26, 27, 28, 29]), such a domain model is not already defined, a well-suited
but this topic is recently becoming more attention again algebra for representing the existing domain knowledge
(e.g., [30, 31, 32, 33]). Nevertheless, there are two ma- needs to be selected and the model generated. This can
jor shortcomings if applied to knowledge injection with be done implicitly by pre-training or explicitly by an
initialization. On the one hand, they are model-specific expert. However, each function   ∈ Λ must operate on
and hard to engineer, which makes them impractical the same solution space defined by the overall task to be
considering the diversity of knowledge representations. solved, for instance, a binary classification or regression
On the other hand, they cannot map to arbitrary ANN problem. In addition, each function   requires a set of
architectures, which may restrict an ANN’s ability to dis-  representative examples   as
cover new characteristics in the subsequent optimization.</p>
      <p>Tthheisexbperceosmseedskinnocrwelaesdingeglayncdrtihtiecaelnatisrethteasgkacpombeptlwexeietny {︁  = {(, , , )}=1}︁|Λ=|1 ,
widens.</p>
      <p>Instead of engineering such mappings by hand, this where , = (,1 , ,2 , . . . , , ) denotes one data
paper introduces a machine learning (ML) framework ca- point of dimensionality  and , =  (, ) is the
pable of training an ANN to become a semantic mapping result after applying the function   to , . How to
from a well-defined, algebraic representation of domain generate Λ and set  depends on the given context.
knowledge to a network’s weights and biases. We call
such a mapping “Transformation Network” or  -Net. 2.1.2. Data Preparation for  -Net Training
This data-driven framework can be applied to various Before the  -Net training, the data and  -functions must
model algebras, such as DTs or polynomials, with only be put in the correct shape, i.e., numeric vectors for ANNs.
slight adaptions. Thereby, it tackles the challenge of vari- Therefore, an encoding method, denoted as  , is
ability in domain knowledge representations by transfer- required. Similarly, a decoding method  enables the
ring the complex mapping generation from humans to translation of the returned network weights and biases
machines. Furthermore, an arbitrary network structure to an executable ANN  . The dataset required for the
as  -Net output can be selected, which achieves an inde-  -Net training is defined as
pendence between the complexity of the domain model
and the target ANN.
 := {( ( ),   )}|Λ=|1 ,</p>
    </sec>
    <sec id="sec-2">
      <title>2. Framework and Approach</title>
      <p>In this section, we give a brief introduction on (1) how
the proposed framework trains an ANN to become a
semantic mapping ( -Net) from the internals of a given
algebraic model to a network’s weights and biases, and
(2) how to utilize its capabilities for knowledge injection
via ANN initialization.
where each example is a tuple of the encoded function
  and its representative samples   . For clarification,
these samples are not the target output of the  -Net, but
are required for the loss calculation. This is described in
the next step.
2.1.3.  -Net Training
After the preparations, the  -Net is trained. Its weights
and biases are adjusted based on the backpropagated
prediction error over the training dataset  . This error
 1, 2, … , |Λ| ∈ Λ</p>
      <p>|Λ|

 =1  =1
   = ൛(  , ,   , )ൟ
 ∶= ൛(
 (  ),    )ൟ|Λ|
function generation, data preparation and  -Net training.
measures how diferent each input function   is com- it for ANN initialization. Therefore, we conducted
expared to the currently predicted ANN counterpart  . To
quantify this diference, a traditional task-specific
measure (), e.g., categorical cross-entropy, is applied to
the true values  =  () and the predictions  ()
given the vector of all input samples . The overall
optimization goal can be formally described as
∈{1,2,...,|Λ|}
minimize (,  ()).</p>
      <sec id="sec-2-1">
        <title>Thus, the  -Net training aims to maximize the transfor</title>
        <p>mation fidelity. By that, we want to enable the  -Net to
generalize to previously unseen  -functions, making it a
capable mapping for this family of functions.</p>
        <p>Execution
2.2. Knowledge Injection via  -Net</p>
      </sec>
      <sec id="sec-2-2">
        <title>After the one-time efort of generating a suitable  -Net, it</title>
        <p>is able to instantly initialize ANNs for all possible domain
models within the trained function algebra. Therefore,
we just need to pass the encoded representation to the
 -Net and let it predict the initial weights and biases.
In the current state, one  -Net maps to ANNs with a
pre-defined specification, i.e., architecture and activation
functions. To achieve a high fidelity, it must be assumed
that ANNs with this specification are capable of
accurately approximating the input functions. However, if
changes to the network specification are required, only
the  -Net training must be repeated with an adapted
output layer and/or  -Decoding.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Evaluation</title>
      <p>In this section, we want to briefly show that our
framework shows promising results in practice in terms of the
 -Net mapping fidelity as well as the efects of applying
periments on polynomials as symbolic domain models to
support solving random regression problems including
two variables. The  -Net was trained on 10,000
polynomials with orders between 0 and 8 to find the closest ANN
approximations with one hidden layer of 225 neurons.</p>
      <p>Without extensive hyperparameter optimization, the
tified by the coeficient of determination (
 -Net could achieve on average a mapping fidelity
quan2) of 0.77
(± 0.26) over representative samples on a set of 2,500 test
polynomials. Despite the noticeable distance to a perfect
mapping (2 = 1), it significantly proves the learning
capability of the proposed ML framework.</p>
      <p>To investigate the impact of injecting knowledge by
utilizing  -Nets on a given ANN task, the training
duration and prediction performance were analyzed. A total
of 2,500 synthetic regression problems were randomly
created and then two ANNs were trained on each
problem; one lets the  -Net predict the initial weights and
biases based on a polynomial approximation, and the
second one applies the Xavier uniform initializer [15] as
a benchmark. Early stopping was used to indicate
convergence during the optimization. Besides the diferent
initialization, all other factors and parameters remained
the same.</p>
      <sec id="sec-3-1">
        <title>By applying the  -Net for initialization, in 91% of</title>
        <p>the regression test cases the prediction performance
increased and 96% required less epochs to converge, i.e.,
hitting the early stopping criterion, compared to the naive
benchmark. More specifically, the training duration could
be reduced on average by 64%. In addition, the resulting</p>
        <sec id="sec-3-1-1">
          <title>ANN performance in terms of the mean absolute error</title>
          <p>(MAE) showed an average increase of 2.7%. Figure 2
illustrates these two benefits in more detail.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>This condensed evaluation demonstrates the benefits of the proposed framework and emphasizes the potential for knowledge injection into ANNs. A more sophisticated evaluation is currently a work in progress.</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>In this paper, we have introduced a novel approach of
knowledge injection into ANNs by utilizing existing
domain models for initialization. Therefore, a semantic
mapping from the domain model’s internals to weights
and biases is applied. Instead of engineering such an
explicit mapping by hand, we designed a machine learning
framework capable of training an ANN to perform this
transformation. We call such a transformation network
 -Net. Besides the reduction of manual efort, it has
the big advantage of decoupling the complexity of the
domain model and ANN space.</p>
      <p>Based on promising initial experiments, we
hypothesize that this framework can generate  -Nets with
suficient fidelity by appropriately addressing the following
aspects: (1) its network specification (e.g., architecture
and activation functions), (2) the learning behavior (e.g.,
loss function and optimizer), (3) the training data
generation (e.g., diversity of domain models), and (4) the
numeric encoding of the domain model algebra.
Learning Systems Practical, volume Making
Learning Systems Practical of Computational Learning
Theory and Natural Learning Systems, MIT Press,
Cambridge, MA, USA, 1997, pp. 3–15.
[29] R. Setiono, W. K. Leow, On mapping
decision trees and neural networks,
KnowledgeBased Systems 12 (1999) 95–99. URL: https://doi.
org/10.1016/S0950-7051(99)00009-X. doi:10.1016/
S0950-7051(99)00009-X.
[30] R. Balestriero, Neural Decision Trees,
arXiv:1702.07360 [cs, stat] (2017). URL:
http://arxiv.org/abs/1702.07360.
[31] S. Wang, C. Aggarwal, H. Liu, Using a
Random Forest to Inspire a Neural Network
and Improving on It, in: Proceedings of
the 2017 SIAM International Conference on
Data Mining (SDM), SIAM, Houston, Texas,
USA, 2017, pp. 1–9. URL: https://epubs.siam.org/
doi/abs/10.1137/1.9781611974973.1. doi:10.1137/
1.9781611974973.1.
[32] G. Biau, E. Scornet, J. Welbl, Neural Random
Forests, Sankhya A 81 (2019) 347–386. URL: https://
doi.org/10.1007/s13171-018-0133-y. doi:10.1007/
s13171-018-0133-y.
[33] K. D. Humbird, J. L. Peterson, R. G. Mcclarren,
Deep Neural Network Initialization With
Decision Trees, IEEE Transactions on Neural
Networks and Learning Systems 30 (2019) 1286–1295.
doi:10.1109/TNNLS.2018.2869694, conference
Name: IEEE Transactions on Neural Networks and
Learning Systems.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>