<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extracting the Common Structure of Compounds to Induce Plant Immunity Activation using ILP</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Atsushi matsumoto</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katsutoshi Kanamori</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kazuyuki Kuchitsu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hayato Ohwada</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. Department of Applied Biological Science, Faculty of Science and Technology, Tokyo University of Science</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>. Department of Industrial Administration, Faculty of Science and Technology, Tokyo University of Science</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>While recent studies have referred to plant immunity activators, it is difficult to find a compound to use for the immunity activation of plants. In this study, we seek to determine compounds that enable plant immunity activity using ILP. With the proposed method, it is possible to predict compounds that induce plant immunity activity, based on the structural features of the compounds. The predicted structure rule also includes structures of known plant immunity activators. However, further investigation is needed regarding the relationship between plant immunity and structure rules.</p>
      </abstract>
      <kwd-group>
        <kwd>ILP</kwd>
        <kwd>Machine learning</kwd>
        <kwd>Plant immunity activation</kwd>
        <kwd>Virtual screening</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Virtual screening is an important approach in the drug discovery process.
Especially, machine learning has recently received broad attention. This paper picks up two
method, Support Vector Machine (SVM) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and Inductive Logic Programming
(ILP). Both method are often used in drug discovery field [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        On the other hand, decreased production of agricultural crops due to pathogenic
bacteria and pests is a serious problem that has not yet been solved. To address this
problem, grower have made a deal with fungicides and pesticides, however, it is
difficult to act selectively on the target (e.g., pests and pathogens). There is a
possibility that the cause of health damage in humans and destruction of biota. In
addition, long-term use of the same drug may cause the emergence of resistant bacteria;
thus, the effect of the drug gradually decreases. In recent years, plant immunity
activators have attracted attention, based on the idea of increasing the immunity of
the plant rather than directly killing pathogens and pests. However, only three
types of plant immunity activator are currently marketed in Japan (Fig. 1). In
addition, the mechanism of plant-immunity activation is still largely unknown [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
The development of plant immunity activators has been slow, due to the time
required and the high cost of screening candidate compounds. Cause of this problem
is the kind of candidate compounds is enormous and each of the compounds were
reacted to the cells to confirm the effect of immunity activation.
      </p>
      <p>
        In this study, we predict compounds that induce plant-immunity activation using
ILP to study compound structures. ILP can be used to determine relationship
patterns between data; therefore, it is suitable to represent the structure of compounds.
Additionally, we obtained the structure of the predicted compound as a rule, which
is one of the excellent points of ILP. A recent study that was conducted to predict
the structure of compounds using ILP exhibited high performance [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In those
cases, the target of compound bonds was known. However, in the present study,
the target of compound bonds is not known. Additionally, we also tried SVM for
comparison with ILP. SVM also exhibited high performance [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. PLANT IMMUNITY</title>
      <p>
        Plant immunity is a defense system to protect plants from various enemies. A
plant-immunity activator is a drug that activates plant immunity. The Kuchitu
group constructed a screening system to find a candidate using the amount of ROS
(reactive oxygen species) generation as an index [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Experiment results
indicated that if the ROS value is high, the compound is likely to be a plant-immunity
activator.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. DATASET</title>
      <p>In the present study, the datasets are experiment data about the plant immunity
activator in Arabidopsis thaliana, compiled by the Kuchitu group. This dataset includes
10000 compounds. Positive examples are 271 high-ROS compounds, and negative
examples are the other 9729 compounds. However, negative examples were reduced
to 813 compounds by random sampling for two reasons. First, imbalanced data
deteriorates learning accuracy. Second, if there are many compounds, calculation takes a
long time. Therefore, 1084 compounds were used in this study.
4. METHOD</p>
      <p>This chapter describes our method. We had two approaches. Fig. 2 shows the
overview of our method.</p>
      <sec id="sec-3-1">
        <title>4.1 ILP Approach</title>
        <p>
          With the ILP approach, structural features and some numerical features of the
compound were used as background knowledge. In this study, we used GKS [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which
is an ILP system. We defined seven predicates to represent the features of the
compounds. In parentheses, there are argument of predicates.
・atom (compound_name, atom_id, element)
        </p>
        <p>Types of atoms present in the compound
・bond (compound_name, atom_id, atom_id, bondtype)</p>
        <p>Bonding state between atoms and bond type in the compound
・Num_AromaticRings (compound_name,Num_AromaticRing)</p>
        <p>The number of aromatic rings in the compound
・Num_Rings (compound_name, Num_Ring)</p>
        <p>The number of rings in the compound
・LogP98 (compound_name, value)</p>
        <p>Lipid solubility of the compound
・LogD (compound_name, value)</p>
        <p>Indication of a change in lipid solubility by a change in Ph value
・ring (compound_name,ring_id,atom_id,ringsize,ringtype)</p>
        <p>Type of ring structure that is composed of each atom. It can represent the
connection of the ring structure and other structures by using this predicate.
By selecting several predicates as background knowledge, we can obtain the structure
of the compound as a learning result (Table 1). Background knowledge is a set of
atomic formulas of each predicate. Atom and bond are always necessary. The reason
why selecting LogP98 and LogD is result of importance calculation using the average
Gini coefficient.</p>
        <p>Predicate
atom,bond
atom,bond,Num_AromaticRings
atom,bond,Num_AromaticRings,Num_rings
atom,bond,ALogP98
atom,bond,Num_AromaticRings,Num_rings,ALogP98,LogD
atom,bond,Num_AromaticRings,Num_rings,LogD
atom,bond,LogD,ring
atom,bond,ring
Mode declaration as input is shown in Fig. 3. A rule selected if it was covered more
than 10 positive examples and less than 10 negative examples.</p>
        <p>@dock,+molecular
@atom,+molecular,+atomid,#atomtype
@atom,+molecular,-atomid,#atomtype
@bond,+molecular,+atomid,+atomid,#bondtype
@bond,+molecular,-atomid,+atomid,#bondtype
@bond,+molecular,+atomid,-atomid,#bondtype
@bond,+molecular,-atomid,-atomid,#bondtype
@Num_Rings,+molecular,#Num_Ring
@Num_AromaticRings,+molecular,#Num_AromaticRing
@LogD,+molecular,#value
@ALogP98,+molecular,#value
@ring,+molecular,+ringid,+atomid,#ringsize,#ringtype
@ring,+molecular,-ringid,+atomid,#ringsize,#ringtype
@ring,+molecular,+ringid,-atomid,#ringsize,#ringtype
@ring,+molecular,-ringid,-atomid,#ringsize,#ringtype</p>
        <p>Fig. 3. Mode declaration</p>
      </sec>
      <sec id="sec-3-2">
        <title>4.2 SVM Approach</title>
        <p>We also tried SVM for comparison with ILP, using 77 features for learning (Table 2).
Detail information is shown in Appendix A.
Cost parameters and gamma parameters were determined using a grid search for 20
split from 0.0001 to 10,000. The kernel used RBF.</p>
      </sec>
      <sec id="sec-3-3">
        <title>4.3 Evaluation</title>
        <p>Ten-fold cross-validation was used in both approaches. True Positive (tp) , False
Negative (fn) , True Negative (tn) , False Positive (fp) , Accuracy , Precision , Recall
and F value were used for Evaluation. Especially, this paper focuses on tp and F
value.
5.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>RESULTS</title>
      <p>Although SVM F values slightly exceeded those of ILP, ILP tp values greatly
exceeded those of SVM. For virtual screening, it is very important to reduce the positive
example of misclassification. Results of this study indicate that structural features of
the compounds are useful in predicting immunity activation.</p>
      <p>Using the ring structure as background knowledge yielded better results than not
using ring structure. Therefore, the ring structure is considered an important factor in
plant immunity activation.</p>
      <p>When analyzing rules using ILP, comparison of known plant immunity activators
indicated that Rule 2 was true for all three compounds. For rule showing a structure
that is different from the known plant immunity activator, there is a need for further
investigation.</p>
      <p>In this study, it was possible to predict the partial structure that exists in all
compounds of known plant-immunity activators. In addition, the rule that is unknown the
relationship between immunity activity has been predicted. In order to improve
prediction accuracy, it is essential to improve background knowledge in the future.</p>
      <p>Table 6 shows feature list in SVM approach. Feature name depends on Discovery
Studio.</p>
      <p>Category Feature name
C HBA_Count</p>
      <p>HBD_Count
NPlusO_Count
Num_AromaticBonds
Num_AromaticRings
Num_AtomClasses
Num_Atoms
Num_Bonds
Num_BridgeBonds
Num_BridgeHeadAtoms
Num_ChainAssemblies
Num_Chains
Num_ExplicitAtoms
Num_ExplicitBonds
Num_ExplicitHydrogens
Num_H_Acceptors
Num_H_Acceptors_Lipinski
Num_H_Donors
Num_H_Donors_Lipinski
Num_Hydrogens
Num_NegativeAtoms
Num_PositiveAtoms
Num_RingAssemblies
Num_RingBonds
Num_Rings
Num_Rings3
Num_Rings5
Num_Rings6
Num_Rings7
Num_Rings8
Num_RotatableBonds
Num_SpiroAtoms
Num_StereoAtoms
Num_StereoBonds
Num_TerminalRotomers
Num_TrueStereoAtoms
Num_UnknownPseudoStereoAtoms
Num_UnknownTrueStereoAtoms
Organic_Count</p>
      <sec id="sec-4-1">
        <title>C: Related to structure A: Related to AlogP E: Related to energy O: Other</title>
      </sec>
      <sec id="sec-4-2">
        <title>W: Related to size or weight</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Appendix: B</title>
      <p>Positive Negative
20
10
7
6
2</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. V.
          <article-title>Vapnik，The Nature of Statistical learning Theory</article-title>
          . Spring-Verlag，NY， USA，
          <fpage>1995</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Tadasuke</given-names>
            <surname>Ito</surname>
          </string-name>
          ，
          <article-title>Hayato Ohwada and Shin Aoki，Combining two machine learning methods for predicting protein-ligand docking using structure and physiochemical properties，</article-title>
          <source>Proc. of the 7th International Conference on Bioinformatics and Computational Biology</source>
          ， pp.
          <fpage>19</fpage>
          -
          <lpage>24</lpage>
          ，March 2015
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>A.</given-names>
            <surname>Srinivasan，S.H.Muggleton，R.D.King</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.J.E.</given-names>
            <surname>Sternberg</surname>
          </string-name>
          <article-title>，Mutagenesis ILP experiments in a non-determinate biological domain，</article-title>
          <source>Proceedings of the Fourth International Inductive Logic Programming Workshop，1994</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Yoshiteru</given-names>
            <surname>Noutoshi</surname>
          </string-name>
          ，
          <article-title>Masateru Okazaki，Tatsuya Kida，Yuta Nishina， Yoshihiko Morishita，Takumi Ogawa，Hideyuki Suzuki，Daisuke Shibata， Yusuke Jikumaru，Atsushi Hamada，Yuji Kamiya，Ken Shirasu，Novel Plant Immune-Priming Compounds Identified via High-Throughput Chemical Screening Target Salicylic Acid Glucosyltransferases in Arabidopsis．The Plant Cell</article-title>
          ，vol.
          <volume>24</volume>
          :
          <fpage>3795</fpage>
          -
          <lpage>3804</lpage>
          ，
          <fpage>2012</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jose C A Santos， Houssam Nassif， David Page ， Stephen H Muggleton， Michael J E Sternberg</surname>
          </string-name>
          <article-title>，Automated identification of protein-ligand interaction features using Inductive Logic Programming:a hexose binding case study ． Santos st al．</article-title>
          <source>BMC Bioinformatics</source>
          <year>2012</year>
          ，
          <volume>13</volume>
          :
          <fpage>162</fpage>
          ，
          <fpage>2012</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>T</given-names>
            <surname>Higashi</surname>
          </string-name>
          ，
          <article-title>T Kurusu，S Hasegawa，</article-title>
          K Kuchitsu，
          <article-title>Dynamic intracellular reorganization of cytoskeletons and the vacuole in defense responses and hypersensitive cell death in plants．</article-title>
          <source>Journal of Plant Research，Volume 124，Issue</source>
          <volume>3</volume>
          ，
          <fpage>pp315</fpage>
          -
          <lpage>324</lpage>
          ，
          <fpage>2011</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Hayato</given-names>
            <surname>Ohwada</surname>
          </string-name>
          ，
          <article-title>Hiroyuki Nishiyama，Fumio Mizoguchi，Concurrent execution of optimal hypoyhesis search for inverse entailment．</article-title>
          <source>Lecture Notes in Artificial Intelligence，</source>
          Spring-Verlag，No.
          <year>1866</year>
          ，Vol.
          <volume>4</volume>
          ，pp.
          <fpage>165</fpage>
          -
          <lpage>173</lpage>
          ，
          <year>2000</year>
          21
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>