<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>OVERLAY</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Minimal Rules from Decision Forests: a Systematic Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giovanni Pagliarini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Paradiso</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Perrotta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guido Sciavicco</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Ferrara</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>6</volume>
      <fpage>28</fpage>
      <lpage>29</lpage>
      <abstract>
        <p>Extracting logical rules from ensembles of symbolic learning models, and especially from ensembles of decision trees, is a very well-known discipline, and several methods and algorithms have been proposed for its solution. However, the existing approaches are characterized by being purely statistical. In this paper, we discuss the problem of systematically extracting minimal logical rules from ensembles of trees from both a theoretical and an algorithmic point of view.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;minimal logical rules</kwd>
        <kwd>ensembles of symbolic models</kwd>
        <kwd>forests of trees</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In sharp opposition to the proliferation of machine learning models used to approach all range of
typical artificial intelligence problems, from classification, to regression, to reinforcement learning, the
universal concept linked to the interpretation and the explanation of such models is that of rule. As
a matter of fact, it can be said that all diferent statistical learning models are diferent methods for
implicit or explicit rule extraction from data.</p>
      <p>
        Unsurprisingly, the idea of expressing the behaviour of a data-driven artificial intelligent agent in
terms of rules is extremely pervasive. On the one side, learning models are usually separated into
symbolic ones, such as decision trees or linear regressions, sub-symbolic ones, such as neural networks,
and mixed ones, such as ensembles of trees. On the other side, methods for explaining learning models
are classified into global ones, that are focused on the model as a whole, and local ones, focused on
the behaviour of a model on a specific instance. Numerous literature surveys on explainable artificial
intelligence and interpretable machine learning have been conducted (e.g., see [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]).
      </p>
      <p>
        With symbolic learning methods we represent data and relationships using symbolic structures.
Decision trees [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] are a classic example of symbolic learning models, where the tree branches typically
represent formulas of propositional logic. Despite their efectiveness, decision trees sufer from a limited
ability to generalize to new data. To address such an issue, ensembles of independent decision trees,
known as decision forests, are commonly used to improve the generalization ability of single trees.
Decision forests are symbolic in nature, but they include a functional component to amalgamate the
output of single trees, and can be therefore classified as mixed symbolic/sub-symbolic techniques; the
most famous algorithm for decision forest learning, namely random forest [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], produces decision forests
for classification/regression in which the aggregation function is simple majority.
      </p>
      <p>
        The development of global explanation methods for decision forests is paramount, and several
solutions have been proposed in the past, including Partial Dependence Plots [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Accumulated Local
Efects Plots [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], global surrogate models [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and SHapley Additive exPlanations (SHAP) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In terms
of rule extraction from decision forests, the relevant methods include the celebrated Simplified Tree
Ensemble Learner (STEL) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] (recently extended to the modal case in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]), and several heuristic
techniques [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ]. Two common traits characterize this plethora of proposals: first, global rules from
decision forests are extracted in statistical form, and, second, these methods are seldom integrated into
widely used machine learning portfolios and frameworks.
      </p>
      <p>
        We initiate a systematic study of purely logical global rules extraction from decision forests. The
language of standard decision forest is propositional; propositions range from simple assertions of the
type  ◁▷ , where  is an attribute,  is a constant, and ◁▷ is a comparison operator (e.g., fever is over
38 degrees), to complex (yet atomic) sentences, of the type  (1, . . . , ) ◁▷ , where 1, . . . ,  are
attributes and  is an (arbitrary) function applied to them (e.g., the averaged vibration of sensors  and
 is below 100 Hertz — examples of such propositions emerge, among others, in oblique decision trees
and forests [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]). In any case, atomic propositions form a simple theory  (e.g., fever is over 38 degrees
implies fever is over 37 degrees); this is ignored during the learning phase, which generates a certain
amount of redundancy in learnt models. In the simple and representative case of decision forests for
classification, that is, learning a class  from a dataset ℐ, given a decision forest  learnt from ℐ we
define a strong class rule for it as an object of the type  ⇔ , where  is a class and  is a propositional
formula, such that  classifies an instance  ∈ ℐ as  if and only if  satisfies  . Obviously, it is to be
expected that useful class rules have relatively short antecedent, and since the latter is a propositional
formula, it can be minimized. Minimization of propositional formulas is a very well-known problem. In
the most common case the size of a formula  , denoted by | |, is defined as the number of its symbols,
and the minimization problem asks, given a propositional formula  : which is a formula  ′, equivalent
to  (denoted  ′ ≡  ), and minimal in size? In its decision version, the problem becomes, given  and a
number : does there exists a formula  ′, such that | ′| ≤  and  ′ ≡  ? This problem is Σ 2-complete,
and there exist a number of approaches for it [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>In this paper we ask the question, given a decision forest  and a class : which is a strong class
rule for  whose antecedent  is minimal in size? In its decision version, given a number , it becomes:
is there a strong class rule for  whose antecedent  is such that | | ≤ ? Obtaining small class
rules difers from pure logical minimization in two key aspects. First, minimization of propositional
formula is usually not intended modulo a theory  , but, in general, classic minimization methods and
techniques can be adapted to this case. Second, in our case minimization does not need to preserve
logical equivalence, but only equivalence modulo the set of instances ℐ on which the original forest
was learnt. We shall see that this problem can be very hard in terms of computational complexity, and
it makes sense to consider other (ideally, simpler) versions of it. We define the concept of right weak
class rule, (resp., left weak class rule) that is, a rule of the type  ⇒  (resp.,  ⇐ ) such that if  ∈ ℐ
satisfies  (resp.,  classifies  as ) then  classifies  as , (resp., that if  ∈ ℐ satisfies  ), and we ask
the question, given a decision forest  and a class : which is a right (resp., left) weak class rule for 
whose antecedent  is minimal (resp., non-trivially maximal) in size? Or, in its decision version, given
also a number : is there a right (resp., left) weak class rule for  whose antecedent (resp., non-trivial
antecedent)  is such that | | ≤  (resp., | | ≥ )?</p>
    </sec>
    <sec id="sec-2">
      <title>2. Decision Trees, Decision Forests, and Class Formulas</title>
      <p>Definition 1. A dataset is a set of  instances ℐ = {1, . . . , }, each one of which is described by the
values of  attributes  = {1, . . . , }.</p>
      <p>Without lack of generality we assume that the value of each attribute in an instance is a real number.
Several problems are usually associated with datasets; in the case of supervised learning, each instance
is also associated to a label (or class)  ∈ ℒ, and a dataset is termed labelled. Given a labelled dataset ℐ,
supervised classification consists of synthesizing an algorithm (a classifier ) that is able to classify the
instances of an unlabelled dataset  whose instances are defined on the same set of attributes.</p>
      <p>In the symbolic context, instances are seen as logical models. To help this interpretation one takes
into consideration that datasets are naturally associated to a logical vocabulary  of propositional
Definition 2.
ℒ) is a tuple
letters, from which formulas are built. In the most general case, we have

=</p>
      <p>{( (1, . . . , ) ◁▷ ) |  ∈ ℱ ,  ∈ R, ◁▷ ∈ {&lt;, ≤ , =, ≥ , &gt;}},
where ℱ is a set of suitable feature extraction functions. To a dataset ℐ, we associate its vocabulary 
and a (possibly empty) theory  , that is, a set of propositional formulas of the type  →  , where
,  ∈  , that expresses semantic constraints between propositional letters (e.g.,  &gt; 5 implies
 &gt; 4). In the following, we write  |=  to denote that a propositional formula  is satisfied by  .</p>
      <p>Let ℒ be a set of classes and  a finite set of propositional letters. Then, a decision tree (on
 = ⟨, , , ⟩,
where ⟨, ⟩ is a full binary directed tree,  is a leaf-labelling function that assigns a class from ℒ to
each leaf node in  , and  is an edge-labelling function that assigns a decision from {, ¬ |  ∈  }
to each edge in , in such a way that two siblings always have opposite decisions. A decision forest
 = { 1, . . . ,  } (on ℒ) is a set of  decision trees (on ℒ).</p>
      <p>A decision tree/forest is learnt from a dataset ℐ based on its vocabulary  (the language of the tree/forest).
An instance is classified by a tree by progressively checking the truth value of each proposition on
a path (and we denote with  ( ) the class assigned to an instance  by  ), and by a forest ( ( ))
by systematically querying each tree individually and then aggregating their decisions; among other
possibilities, a typical aggregation function is simple majority, which is assumed here.
Definition 3. Given a decision tree  on ℒ and a class  ∈ ℒ, then: () given and a path  in  from
the root to a leaf labeled with  (an L-path), the conjunction of all decisions on  is called L-path tree
formula, and it is denoted by  , and () the disjunction of all -path formulas in  is called L-class tree
formula ( ). Given a decision forest  on ℒ with  trees, and a class  ∈ ℒ, then: () given a collection
 1 , . . . ,   ∈  and one -path   (1 ≤  ≤ ) per tree   , the conjunction of all -path tree formulas
 1 , . . . ,  1 is called partial L-path forest formula ( 1 ,...,  ); () if  &gt; /2, then  1 ,...,  is called
L-path forest formula; and () the disjunction of all possible -path forest formulas is called L-class
forest formula (  ).</p>
    </sec>
    <sec id="sec-3">
      <title>3. Rules from Decision Forests</title>
      <p>Definition 4. Given a decision forest  , learnt from a dataset ℐ, on ℒ, and a class  ∈ ℒ, a strong
-class rule is an object of the type  ⇔ , where  (the antecedent) is a propositional formula in the
language of  , such that, for every  ∈ ℐ,  |=  if and only if  ( ) = , a right weak -class rule is an
object of the type  ⇒  such that  |=  implies  ( ) = , and a left weak -class rule is an object of
the type  ⇒  such that  ( ) =  implies  |=  .</p>
      <p>Given a decision forest  on ℒ, a class  ∈ ℒ, and the -class forest formula   ,   ⇔  is a (trivial)
strong -class rule. In practical terms, it will likely be very redundant, due to the fact that decision trees
and forest are general learnt via sub-optimal learning algorithms (recall that the problem of extracting
a minimal decision tree is NP-hard, and sub-optimal, polynomial algorithms are commonly used for
learning), and the fact that the theory  underlying the language is ignored during learning. Thus,
given a decision forest  on ℒ, learnt from a dataset ℐ, and a class  ∈ ℒ, we are interested in finding
a minimal strong -class rule, that is, a strong -class rule  ⇔  such that, for every strong class rule
 ′ ⇔  so that  ≡ ℐ  ′ (i.e., so that  and  ′ are equivalent modulo  at least with respect to the
instances in ℐ), it is the case that | | ≤ |  ′|.</p>
      <p>One way to assess the complexity of the problem of finding minimal strong rules is to study its
decision version, that is: given a decision forest  , learnt from a given dataset ℐ, on ℒ, a class  ∈ ℒ,
and a number , is there a strong -class rule  ⇔  such that | | ≤ ? Given that the size of the input
of this problem is the number of symbols of  (denoted by | |) plus the number of instances in ℐ (|ℐ|)
and the size of the representation of  (||), we have the following result.</p>
      <p>Theorem 1. Given a decision forest  , learnt from a dataset ℐ, on ℒ, a class  ∈ ℒ, a theory  , and a
number , the problem of establishing if there exists a strong -class rule  ⇔  such that | | ≤  is in
NEXPTIME.</p>
      <p>Proof[sketch]. An (DNF) antecedent  of size less than or equal to  can be guessed. Then, for each
instance  ∈ ℐ,  is checked against both  and  : if  () =  (resp., not ) and  |=  (resp.,  ̸|=  ),
then  is marked. If all instances in ℐ end up being marked, then  ⇒  is a strong -class rule. The
complexity of this process is polynomial in | | and |ℐ|, but exponential in ||, and the value of the
minimal  for which a strong -class rule exist may be exponential in | |. □
Since extracting a minimal strong class rule may turn out to be impractical, we turn our attention
to weak rules. Given a decision forest  , learnt on a dataset ℐ, on ℒ, a class  ∈ ℒ, and a -path
forest formula  1 ,...,  ,  1 ,...,  ⇒  is a (trivial) right weak -class rule whose antecedent may
exhibit the same kind of redundancy and can be minimized as in the previous case. Similarly, ⊤ ⇐ 
is a (trivial) left weak -class rule whose antecedent can be maximized; in this case, however, one is
interested in non-trivial maximal antecedents.</p>
      <p>Theorem 2. Given a decision forest  on ℒ, a class  ∈ ℒ, a theory  , and a number , the problem of
establishing if there exists a right weak -class rule  ⇒  such that | | ≤  is in NP, and the problem of
establishing if there exists a non-trivial left weak -class rule  ⇒  such that | | ≥  is in NEXPTIME.
Proof[sketch]. As for right weak -class rules, a (term) antecedent  of size less than or equal to  can
be guessed. Then, for each instance  ∈ ℐ,  is checked against both  and  : if  ̸|=  , or  |=  and
 () = , then  is marked. If all instances in ℐ end up being marked, then  ⇒  is a weak -class
rule. The complexity of this process is polynomial in | |, |ℐ|, but exponential in ||; however, the value
of the minimal  for which a right weak -class rule exist is polynomial in | |. As for left weak -class
rules, a (DNF) antecedent  of size grater than or equal to  can be guessed. Then, we first check that 
is non-trivial, that is, there are no repeated literals in any term, no term is unsatisfiable, and no terms
implies any other term. Then, for each instance  ∈ ℐ,  is checked against both  and  : if  () ̸= ,
or  () =  and  |=  , then  is marked. If all instances in ℐ end up being marked, then  ⇒  is a
weak -class rule. The complexity of this process is polynomial in | | and |ℐ|, but exponential in ||,
and the value of the minimal  for which a left weak -class rule exist may be exponential in | |. □</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>We started a systematic study of logical methods for rule extraction from decision forests, a well-known
classification model. Extracting rules from decision forests is a well-known problem, but existing
solutions are statistical and data-driven. In our work, we apply known logical algorithms to rule
extraction, contributing to bridging the gap between logic and machine learning.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>We acknowledge the support of the FIRD project Methodological Developments in Modal Symbolic
Geometric Learning, funded by the University of Ferrara, and the INDAM-GNCS project Symbolic
and Numerical Analysis of Cyberphysical Systems (code CUP_E53C23001670001), funded by INDAM;
Giovanni Pagliarini and Guido Sciavicco are GNCS-INdAM members. Moreover, this research has also
been funded by the Italian Ministry of University and Research through PNRR - M4C2 - Investimento
1.3 (Decreto Direttoriale MUR n. 341 del 15/03/2022), Partenariato Esteso PE00000013 - "FAIR - Future
Artificial Intelligence Research" - Spoke 8 "Pervasive AI", funded by the European Union under the
NextGeneration EU programme".</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Pereira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Cardoso</surname>
          </string-name>
          ,
          <source>Machine Learning Interpretability: A Survey on Methods and Metrics, Electronics</source>
          <volume>8</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <article-title>Techniques for interpretable machine learning</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>63</volume>
          (
          <year>2020</year>
          )
          <fpage>68</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monreale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruggieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Turini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Giannotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pedreschi</surname>
          </string-name>
          ,
          <article-title>A Survey of Methods for Explaining Black Box Models</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>51</volume>
          (
          <year>2019</year>
          )
          <volume>93</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>93</lpage>
          :
          <fpage>42</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Quinlan</surname>
          </string-name>
          ,
          <article-title>Induction of Decision Trees, Machine Learning 1 (</article-title>
          <year>1986</year>
          )
          <fpage>81</fpage>
          -
          <lpage>106</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Breiman</surname>
          </string-name>
          , Random forests,
          <source>Machine Learning</source>
          <volume>45</volume>
          (
          <year>2001</year>
          )
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <article-title>Greedy function approximation: A gradient boosting machine</article-title>
          .,
          <source>The Annals of Statistics</source>
          <volume>29</volume>
          (
          <year>2001</year>
          )
          <fpage>1189</fpage>
          -
          <lpage>1232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Apley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Visualizing the efects of predictor variables in black box supervised learning models</article-title>
          ,
          <source>Journal of the Royal Statistical Society Series B</source>
          <volume>82</volume>
          (
          <year>2020</year>
          )
          <fpage>1059</fpage>
          -
          <lpage>1086</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Craven</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shavlik</surname>
          </string-name>
          ,
          <article-title>Extracting Tree-Structured Representations of Trained Networks</article-title>
          ,
          <source>in: Proceedings of the 8th Advances in Neural Information Processing Systems (NIPS)</source>
          ,
          <year>1995</year>
          , pp.
          <fpage>24</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lundberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>A unified approach to interpreting model predictions</article-title>
          ,
          <source>in: Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>4768</fpage>
          -
          <lpage>4777</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <article-title>Interpreting tree ensembles with inTrees</article-title>
          ,
          <source>International Journal of Data Science and Analytics</source>
          <volume>7</volume>
          (
          <year>2019</year>
          )
          <fpage>277</fpage>
          -
          <lpage>287</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ghiotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Manzella</surname>
          </string-name>
          , G. Pagliarini, G. Sciavicco,
          <string-name>
            <surname>I. Stan</surname>
          </string-name>
          ,
          <article-title>Evolutionary explainable rule extraction from (modal) random forests</article-title>
          ,
          <source>in: Proc. of the 26th European Conference on Artificial Intelligence (ECAI) and the 12th Conference on Prestigious Applications of Intelligent Systems (PAIS)</source>
          , volume
          <volume>372</volume>
          <source>of Frontiers in Artificial Intelligence and Applications</source>
          , IOS Press,
          <year>2023</year>
          , pp.
          <fpage>827</fpage>
          -
          <lpage>834</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mashayekhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gras</surname>
          </string-name>
          ,
          <article-title>Rule extraction from decision trees ensembles: New algorithms based on heuristic search and sparse group lasso methods</article-title>
          ,
          <source>International Journal of Information Technology &amp; Decision Making</source>
          <volume>16</volume>
          (
          <year>2017</year>
          )
          <fpage>1707</fpage>
          -
          <lpage>1727</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bénard</surname>
          </string-name>
          , G. Biau,
          <string-name>
            <given-names>S. Da</given-names>
            <surname>Veiga</surname>
          </string-name>
          , E. Scornet,
          <article-title>Interpretable random forests via rule extraction</article-title>
          ,
          <source>in: Proc. of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS)</source>
          , volume
          <volume>130</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>937</fpage>
          -
          <lpage>945</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>B.</given-names>
            <surname>Menze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kelm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Splitthof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Koethe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hamprecht</surname>
          </string-name>
          ,
          <article-title>On oblique random forests</article-title>
          ,
          <source>in: Proc. of the European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          , Springer,
          <year>2011</year>
          , pp.
          <fpage>453</fpage>
          -
          <lpage>469</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Umans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Villa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sangiovanni-Vincentelli</surname>
          </string-name>
          ,
          <article-title>Complexity of two-level logic minimization</article-title>
          ,
          <source>IEEE Transactions on Compututer Aided Design of Integrated Circuits Systems</source>
          <volume>25</volume>
          (
          <year>2006</year>
          )
          <fpage>1230</fpage>
          -
          <lpage>1246</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>