<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Adversarial Consistent Learning on Partial Domain Adaptation of PlantCLEF 2020 Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Youshan Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brian D. Davison</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lehigh University, Computer Science and Engineering</institution>
          ,
          <addr-line>Bethlehem, PA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Domain adaptation is one of the most crucial techniques to mitigate the domain shift problem, which exists when transferring knowledge from an abundant labeled sourced domain to a target domain with few or no labels. Partial domain adaptation addresses the scenario when target categories are only a subset of source categories. In this paper, to enable the e cient representation of cross-domain plant images, we rst extract deep features from pre-trained models and then develop adversarial consistent learning (ACL) in a uni ed deep architecture for partial domain adaptation. It consists of source domain classi cation loss, adversarial learning loss, and feature consistency loss. Adversarial learning loss can maintain domain-invariant features between the source and target domains. Moreover, feature consistency loss can preserve the ne-grained feature transition between two domains. We also nd the shared categories of two domains via down-weighting the irrelevant categories in the source domain. Experimental results demonstrate that training features from NASNetLarge model with proposed ACL architecture yields promising results on the PlantCLEF 2020 Challenge.</p>
      </abstract>
      <kwd-group>
        <kwd>Adversarial learning</kwd>
        <kwd>Partial domain adaptation</kwd>
        <kwd>Plant identi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Automated plant identi cation is important in recognizing plant species. The
availability of massive labeled training data is a prerequisite of machine learning
models. Unfortunately, such a requirement cannot be met in the plant identi
cation problem since we have sparse labels for real-world plant images. Therefore,
we propose to transfer knowledge from an existing auxiliary labeled herbarium
domain to the eld photo domain with limited or no labels. However, due to
the phenomenon of data bias or domain shift [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], classi cation models do not
generalize well from an existing herbarium domain to a novel eld photo domain.
      </p>
      <p>
        Domain adaptation (DA) has been proposed to leverage knowledge from an
abundant labeled source domain to learn an e ective predictor for the target
domain with few or no labels, while mitigating the domain shift problem [
        <xref ref-type="bibr" rid="ref16 ref17 ref19 ref20">16, 17,
19,20</xref>
        ]. In this paper, we focus on unsupervised domain adaptation (UDA), where
the target domain has no labels. Since we have fewer classes in the eld photo
domain, and the classes of the eld photo domain is a subset of the classes of
the source herbarium domain, we investigate partial domain adaptation (PDA)
for the PlantCLEF 2020 Challenge.
      </p>
      <p>
        Recently, deep neural network methods have been widely used in the domain
adaptation problem. Notably, adversarial learning shows its power in embedding
in deep neural networks to learn feature representations to minimize the
discrepancy between the source and target domains [
        <xref ref-type="bibr" rid="ref14 ref9">9, 14</xref>
        ]. Inspired by the generative
adversarial network (GAN) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], adversarial learning also contains a feature
extractor and a domain discriminator. The domain discriminator can distinguish
the source domain from the target domain, while the feature extractor can learn
domain-invariant representations to fool the domain discriminator [
        <xref ref-type="bibr" rid="ref10 ref18 ref9">9,10,18</xref>
        ]. The
target domain risk (the error of the target domain) is expected to be minimized
via minimax optimization. Cao et al. presented adversarial learning for PDA,
which alleviates negative transfer by reducing the outlier of source classes for
training the source classi er and domain labels, while positive transfer is
improved via matching the feature distributions in the shared label space [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Similarly, the example transfer network is proposed to jointly learn domain-invariant
representations and a progressive weighting method to examine the
transferability of source examples. The model can improve positive transfer by relevant
examples and mitigate negative transfer by identifying irrelevant examples [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Although many methods are proposed for partial domain adaptation, they
still su er from two challenges: (1) the models are evaluated on small datasets,
while it has lower transferability to the large-scale dataset, and (2) the feature
consistency of two domains is inappropriately ignored.</p>
      <p>To address the aforementioned challenges, we aggregate three di erent loss
functions in one framework: source domain classi cation loss, adversarial
learning loss, and feature consistency loss to reduce the discrepancy of the two
domains. Moreover, our model is evaluated on a large-scale plant identi cation
dataset to improve the estimate of the generalization ability of our model.</p>
      <p>Our contributions are three-fold:
1. We propose a novel adversarial consistent learning network (ACL) for PDA,
to adversarially minimize the domain discrepancy of the source and target
domains and maintain domain-invariant features;
2. The proposed adversarial learning loss and feature consistency loss can
distinguish the target domain from the source domain, and preserve the
negrained feature transition between the two domains;
3. We impose shared category selection to lter out the irrelevant categories
in the source domain. By down-weighting the irrelevant categories in the
source domain, we can reduce negative transfer from the source domain to
the target domain.</p>
      <p>Experimental results show that ACL achieves higher classi cation accuracy than
several baseline methods and yields promising results on the PlantCLEF 2020
Challenge.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Dataset</title>
      <p>
        PlantCLEF 2020 is a large-scale dataset of the PlantCLEF 2020 task [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
organized in the context of the LifeCLEF 2020 challenge [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Fig. 1 shows some
challenging images in this dataset. The herbarium domain contains 320,750
images in 997 species, and the number of images in di erent species are
unbalanced. This dataset consists of herbarium sheets whereas the test set will be
composed of eld pictures. The validation set consists of two domains
herbarium photo associations and photos. Herbarium photo associations domain
includes 1,816 images from 244 species. This domain contains both herbarium
sheets and eld pictures for a subset of species, which enables learning a mapping
between the herbarium sheets domain and the eld pictures domain. Another
photo domain has 4,482 images from 375 species and images are from plant
pictures in the eld, which is similar to the test dataset. The test dataset contains
3,186 unlabeled images. Due to the signi cant di erence between herbarium and
real photos, it is extremely di cult to identify the correct class.
      </p>
      <p>We exclude the classId of \108335" in the photo domain since the major
classes are from the herbarium domain. In addition, herbarium domain does not
contain the \108335" category. Therefore, eight images are excluded in the photo
domain. The statistics of the PlantCLEF 2020 dataset are listed in Tab. 1.
and @@LCCoonn ) from class label classi er, domain label predictor and feature
consistency regressor. The ACL model consists of three di erent loss functions (source
classi cation loss LS , adversarial domain loss LA, and feature consistency loss
LCon). The feature extractor G in the shared layers is used for both classi er f
and domain discriminator D (The blue dash lines are the backward gradients,
and GRL stands for gradient reversal layer). Layers visualization of architecture
is shown in Fig. 3.
3
3.1</p>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <sec id="sec-3-1">
        <title>Motivation</title>
        <p>
          Previous partial domain adaptation methods [
          <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
          ] evaluated their models based
on a small dataset (e.g., O ce 31), while their models have lower generalizability
to large-scale datasets. In addition, feature consistency of both source and target
domains is not well addressed in the PDA.
        </p>
        <p>In this paper, we present our approach: adversarial consistent learning (ACL)
on partial domain adaptation. It can align the feature distribution of the source
and target domains in the shared categories and guarantee feature consistency
across the two domains. Importantly, ACL identi es irrelevant source categories
via down-weighting class importance automatically. Evaluation on the large-scale
PlantCLEF 2020 challenge dataset shows a high generalizability of our model.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Problem and notation</title>
        <p>For unsupervised domain adaptation, given a source domain DS = fXS i; YS igiN=S1
of NS labeled samples across the set of categories CS and a target domain DT =
fXT j gjN=T1 of NT samples without any labels (YT is unknown) across the set of
categories CT . For partial domain adaptation, the number of categories in CT is
less than the number of categories in CS , and CT $ CS . The samples XS and XT
obey the marginal distributions of PS and PT . The conditional distributions of
two domains are denoted as QS and QT . Due to the discrepancy of two domains,
the distributions are assumed to be di erent, i.e., PS 6= PT and QS 6= QT . Our
ultimate goal is to learn a classi er f under a feature extractor G, which selects
shared categories between two domains, and ensures lower generalization error
in the target domain.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Deep features extraction</title>
        <p>
          To circumvent the large computation resource requirement of training
largescale PlantCLEF 2020 challenge datasets, we instead focus on deep features
from pre-trained models. Based on Zhang and Davison [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], the deep features
are extracted from the last fully connected layer of the pretrained model via
. One represented feature vector has the size of 1 1000 and corresponds to
one plant image. Therefore, the source domain and the target domain can be
represented by (XS ) 2 RNS 1000 and (XT ) 2 RNT 1000, respectively.
3.4
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Source classi er</title>
        <p>The task in the source domain is trained using the typical cross-entropy loss in
following equation:</p>
        <p>LS (f (G( (XS ))); YS ) =</p>
        <p>NS i=1 c=1
1 XNS XCS YS iclog(f (G( (XS i))));
(1)
where YS ic 2 [0; 1]CS is the probability of each class of ground truth for the
ith element of S, f is the classi er in Fig. 2, and f (G( (XS i))) is the predicted
probability.
3.5</p>
      </sec>
      <sec id="sec-3-5">
        <title>Adversarial domain loss</title>
        <p>In general adversarial learning, the system learns a mapping from the source
domain to the target domain. Given the feature representation of feature extractor
G, we can learn a discriminator D, which can distinguish the two domains using
the following loss function:</p>
        <p>LA(GXS XT ; G( (XS )); G( (XT ))) =
1 XNS log(1
NS i=1
However, Eq. 2 only guarantees source domain data will be close to the target
data (GXS XT ), and it does not ensure that the target data will be close to the
source data. We hence introduce another mapping from the target domain to
the source domain GXT XS in Eq. 3 and train it with the same adversarial loss
as in GXS XT as shown in Eq. 2.</p>
        <p>LA(GXT XS ; G( (XS )); G( (XT )))
(3)
For GXS XT , the source domain has the label of 0 and the target domain has the
label of 1, which is corresponding to Domain Label 1 in Fig. 2. Meanwhile, for
GXT XS , 1 is the new label for the source domain and and 0 is the new label for
target domain, which is corresponding to Domain Label 2 in Fig. 2. Therefore,
we de ne the adversarial learning loss as:</p>
        <p>LA(G( (XS )); G( (XT ))) = LA(GXS XT ; G( (XS )); G( (XT )))
+ LA(GXT XS ; G( (XS )); G( (XT ))):
3.6</p>
      </sec>
      <sec id="sec-3-6">
        <title>Feature consistency loss</title>
        <p>To encourage the source domain and target domain information to be preserved
during adversarial learning, we propose a feature consistency loss in our model.
Details of the feature reconstruction layers are shown in Fig. 2; the reconstructed
layers are right behind the feature extractor G in the shared layers, and they
aim to reconstruct the extracted features and maintain the invariant features
during the conversion process. The feature consistency loss is de ned as:
LCon(GXS XT ; GXT XS ; G( (XS )); G( (XT )))
= Exs G( (XS))[`(GXT XS (GXS XT (xs))
+ Ext G( (XT ))[`(GXS XT (GXT XS (xt))
xs)]
xt)];
where ` is the mean squared error loss function, which calculates the di erence
between true features and the reconstructed features.
3.7</p>
      </sec>
      <sec id="sec-3-7">
        <title>Shared categories selection</title>
        <p>In PDA, the set of target domain labels is a subset of the source domain labels,
i.e., CT $ CS . In the PlantCLEF challenge, the size of irrelevant label set (CS
CT ) is far larger than the size of CT (jCS CT j &gt;&gt; jCT j). If we use all elements
of the source domain distribution to match the target domain distribution, it
will cause negative transfer since the target domain will also be forced to match
the irrelevant labels (CS CT ). Therefore, it is important to identify the shared
categories between source and target domains.</p>
        <p>To address the aforementioned challenge, we re-weight the source domain
label set via reducing the irrelevant label set. During the training, we can get
^
the predicted probability of the target domain: YT j = f (G( (XT j ))), which
gives a probability of each source label in CS . As we know, the set of irrelevant
source labels and target label set are disjoint, and the target data are signi cantly
dissimilar to the source data in the irrelevant label set. Therefore, the probability
(4)
(5)
of irrelevant categories should be su ciently small and can be ignored. We then
de ned the weight vector as:</p>
        <p>W =</p>
        <p>NT j=1
1 XNT Y^T j ;
(6)
where W is a jCS j-dimensional weight vector. The irrelevant categories (CS CT )
will have a much smaller weight than the shared categories. We then assign the
weight as 0 if its element Wc is less than a su ciently small number (e.g.,
10e 9). By reducing the weight of irrelevant categories, the shared categories
can be emphasized and negative transfer will be mitigated. The weight vector W
is applied in both the source classi er and domain discriminator over the source
domain data as shown in the following objective function.
3.8</p>
      </sec>
      <sec id="sec-3-8">
        <title>Overall objective</title>
        <p>We combine the three aforementioned loss functions to formalize our objective
function:</p>
        <p>L(XS ; XT ; YS ; GXS XT ; GXT XS )
= LS (f (W(G( (XS )))); YS ) +
+</p>
        <p>LA(W(G( (XS ))); G( (XT )))
(7)</p>
        <p>LCon(GXS XT ; GXT XS ; W(G( (XS ))); G( (XT )));
where and are tradeo parameters between di erent loss functions. Our
model ultimately minimizes the di erence during the transition from the source
domain to target domain and from the target domain to the source domain.
Meanwhile, it maximizes the ability to distinguish the two domains.
3.9</p>
      </sec>
      <sec id="sec-3-9">
        <title>Gradients of shared layers</title>
        <p>The shared layers consist of the feature extractor G and the feature
reconstruction layers. In G, there are two dense layers, a \Relu" activation layer, and a
dropout layer. The numbers of units of the dense layer are 1000 and 997,
respectively. The rate of the Dropout layer is 0.5. The feature reconstruction layers
have a \Relu" activation layer, a dropout layer and a dense layer with the
number of units of 1000. The shared layers are jointly optimized by both the source
classi cation loss, adversarial domain loss and feature consistency loss.</p>
        <p>Let FE ( ; E ) be the output of the shared encoder with parameters of E . In
addition, let FS ( ; S ) be the output of class label classi er with parameters of
S , FA( ; A) be the output of domain label predictor with parameters of A, and
FCon( ; Con) be the output of feature consistency regressor with parameters of
Con. Therefore, the shared layers are optimized by these three gradients. The
parameters in the shared layers are updated in the following equation:
S</p>
        <p>S</p>
        <p>
          E
E
where is the learning rate and
layer (GRL) in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
3.10
        </p>
      </sec>
      <sec id="sec-3-10">
        <title>Theoretical Analysis</title>
        <p>is the adaptation factor from gradient reversal
(10)
(11)
We now formalize the error bound of our model. ACL model is trained with both
the labeled source domain and the unlabeled target domain. The error bound of
the source domain and the target domain ( (h)) in our model is then formally
written as:
(h) =</p>
        <p>XS (h; YS ) +</p>
        <p>XT (h; Y^T );
(9)
where Y^T is the predicted label of target domain. The term XS (h; YS ) =
Ex XS [jh(x) YS j] and XT (h; Y^T ) = Ex XT [jh(x) Y^T j] denote the expected
risk over the source domain and the target domain with respect to the ground
truth labels and predicted labels, respectively (where j j is the L1 norm).</p>
        <p>
          During the training, we expect the error XT (h; Y^T )) to be close to XT (h; YT ),
which evaluates the classi er f with true target domain labels. The smaller the
di erence between these two errors, the better the model performs and more
discrepancies of the two domains are reduced. Existing domain adaptation
theory shows that the risk in the target domain can be minimized by bounding the
source risk and discrepancy between source and target domains [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]). Therefore,
the generalization error bound of our model is shown in the following Lemma.
        </p>
        <p>Lemma 1 Let h be a hypothesis in a class H. Then
(h) =</p>
        <p>XS (h; YS ) +</p>
        <p>XT (h; Y^T )
2 XS (h) + dH(DS ; DT ) + C;
where dH(DS ; DT ) = 2 suph;h02H j XS (h; h0) XT (h; h0)j is the H-divergence of
training and test data in the hypothesis space H: C = XS (h ; YS ) + XT (h ; YT )
is the adaptability to quantify the error in ideal hypothesis h space of training
and test data, which should be small and is the optimal hypothesis via
minimizing the joint error in Eq. 11.</p>
        <p>h = arg min XS (h; YS ) +</p>
        <p>XT (h; YT )</p>
        <p>In Lemma 1, the generalization boundary of our model consists of three
terms: training data error, data discrepancy dH(DS ; DT ), which is estimated by
the disagreement of hypothesis in the space H, and the adaptability C of the ideal
joint hypothesis. In ACL model, the rst term is measured by Eq. 1. The domain
discrepancy is assessed by adversarial learning loss and feature consistency loss.
Furthermore, ACL nds the ideal hypothesis and reduces the training error in
each iteration. Hence, our model can nd a minimal boundary for two domains.
In other words, ACL can implicitly minimize the target domain risk, domain
discrepancy, and the adaptability of true hypothesis h in terms of the hypothesis
space H:</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <sec id="sec-4-1">
        <title>Implementation details</title>
        <p>
          As aforementioned, the deep features are extracted from the last fully
connected layer [
          <xref ref-type="bibr" rid="ref15 ref17">15, 17</xref>
          ]. One represented feature vector has the size of 1 1000
and is corresponding to one plant image. Therefore, the feature
representation of domain herbarium (H) has the size of 320; 750 1000, domain
herbarium photo associations (A) has the size of 1; 816 1000, domain photo (P) has
the size of 4; 482 1000, and domain test (T) has the size of 3; 186 1000. In the
experiment, our task is to reduce the error in the target domain (real-world plant
images), i.e., photo domain or test domain. Our tasks will focus more on the
evaluation of the domain P and domain T. Since the herbarium photo associations
(A) is important to bridge the map between two domains, we hence include
the domain A in the training procedure to form a new source domain, which
consists of domain herbarium (H) and domain A. Domain H + A has the size
of 322; 566 1000. We then train the model based on these extracted feature
vectors. In Tab. 2, H P represents learning knowledge from domain H, which
is applied to domain P.
        </p>
        <p>The parameters of ACL are rst tuned based on the performance of the
domain P, while the model is trained with H + A domain. We then apply these
parameters to domain T and submit it to the challenge for the evaluation. Our
implementation is based on Keras. The parameters settings are = = 0:5,
= 0:31, learning rate: = 0:0001, batch size = 128, the number of iterations
is 1000 and the optimizer is Adam. The details of the layers are shown in Fig. 3.</p>
        <p>
          We also compare our results with two domain adaptation methods: DANN [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
and ADDA [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. In addition, we extracted features from four well-trained models
(ResNet50 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], InceptionV3 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], Inception-Resnet-V2 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], NASNetLarge [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]),
which is trained based on large-scale ImageNet datasets. We then feed these
di erent extracted features into the shared layers and optimize the objective
function in Eq. 7.
The performance of the photo domain is shown in Tab. 2. We report the
accuracy of the whole photo domain (Acc = PjN=T1(Y^T j == YT j )=NT 100),
where Y^T is the predicted label for the target domain. We can observe that the
extracted features from NASNetLarge with our ACL architecture achieves the
highest performance across all three tasks. We observe that two domain
adaptation methods have relatively lower performance in all three tasks. One reason is
that these two methods have weak feature extractors, and they do not exclude
the irrelevant categories in the source domain, which might cause the negative
transfer. Moreover, with the increasing of the ImageNet model, we can extract
better features from plant images, which lead to the high performance of the
NASNetLarge-ACL model. In addition, we conduct an ablation study in which
we train the best NASNetLarge-ACL model without the shared categories
selection (NASNetLarge-ACL W). The results from all three tasks are lower
than NASNetLarge-ACL model, which indicates the shared categories selection
is useful in our model. These experiments demonstrate the e ciency of the ACL
model in nding the invariant-features of two domains.
        </p>
        <p>In the nal stage of the PlantCLEF 2020 Challenge, our solutions are
evaluated by the organizers using the test domain data. As shown in Tab. 3, our
method achieved mean reciprocal rank (MRR) of 0.032 in the whole test domain,
and MRR of 0.016 in the subset of the test domain, and our method places 4th
in the contest.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>There are two compelling advantages of the ACL model. First, we consider the
adversarial consistent learning paradigm, which maintains the domain-invariant
features from the source domain to the target domain and vice versa. Secondly,
we reduce the weight of irrelevant categories in the source domain, which
eliminates the negative transfer during the training. Although the performance of our
model is better than several baseline methods, the highest accuracy of the photo
domain is less than 10%, which illustrates that the transfer learning ability in
the real world image is lower. One underlying reason is that PlantCLEF 2020
Challenge has di cult datasets|that there are signi cant di erences between
herbarium domain and photo domain, as shown in Fig. 1. Another reason is
caused by the weakness of our model since we only train deep features instead
of raw images to reduce the computational requirements; some features might
be ignored during the training. The performance of the ACL model could be
improved if we train the architecture with raw images.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper, we propose an adversarial consistent learning network on
partial domain adaptation termed (ACL) to overcome limitations in nding proper
shared categories and guaranteeing the feature consistency of two domains. Our
model is optimized via minimizing a three-component loss function. As each
component of our ACL model, explicit domain-invariant features are maintained
through such a cross-domain training scheme. Experimental results demonstrate
our proposed ACL model yields promising results on the PlantCLEF 2020
Challenge.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ben-David</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blitzer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crammer</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kulesza</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vaughan</surname>
            ,
            <given-names>J.W.:</given-names>
          </string-name>
          <article-title>A theory of learning from di erent domains</article-title>
          .
          <source>Machine Learning</source>
          <volume>79</volume>
          (
          <issue>1-2</issue>
          ),
          <volume>151</volume>
          {
          <fpage>175</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Partial adversarial domain adaptation</article-title>
          .
          <source>In: Proceedings of the European Conference on Computer Vision (ECCV)</source>
          . pp.
          <volume>135</volume>
          {
          <issue>150</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Domain adversarial reinforcement learning for partial domain adaptation</article-title>
          . arXiv preprint arXiv:
          <year>1905</year>
          .
          <volume>04094</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ghifary</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleijn</surname>
            ,
            <given-names>W.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , M.:
          <article-title>Domain adaptive neural networks for object recognition</article-title>
          .
          <source>In: Paci c Rim international conference on arti cial intelligence</source>
          . pp.
          <volume>898</volume>
          {
          <fpage>904</fpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Goeau, H.,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Overview of lifeclef plant identi cation task 2020</article-title>
          . In: CLEF working notes
          <year>2020</year>
          ,
          <article-title>CLEF: Conference and Labs of the Evaluation Forum</article-title>
          , Sep.
          <year>2020</year>
          , Thessaloniki,
          <string-name>
            <surname>Greece.</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Goodfellow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pouget-Abadie</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mirza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warde-Farley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ozair</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courville</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Generative adversarial nets</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>2672</volume>
          {
          <issue>2680</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          . pp.
          <volume>770</volume>
          {
          <issue>778</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Goeau, H.,
          <string-name>
            <surname>Kahl</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deneu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Servajean</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cole</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Picek</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Ruiz De Castan~eda, R., e, Lorieul,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Botella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Glotin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Champ</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Vellinga</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.P.</surname>
          </string-name>
          , Stoter,
          <string-name>
            <given-names>F.R.</given-names>
            ,
            <surname>Dorso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          , Muller, H.:
          <article-title>Overview of lifeclef 2020: a systemoriented evaluation of automated species identi cation and species distribution prediction</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          <year>2020</year>
          ,
          <article-title>CLEF: Conference and Labs of the Evaluation Forum</article-title>
          , Sep.
          <year>2020</year>
          , Thessaloniki,
          <string-name>
            <surname>Greece.</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Transferable adversarial training: A general approach to adapting deep classi ers</article-title>
          .
          <source>In: International Conference on Machine Learning</source>
          . pp.
          <volume>4013</volume>
          {
          <issue>4022</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Conditional adversarial domain adaptation</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          . pp.
          <volume>1647</volume>
          {
          <issue>1657</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          , et al.:
          <article-title>A survey on transfer learning</article-title>
          .
          <source>IEEE Transactions on knowledge and data engineering</source>
          <volume>22</volume>
          (
          <issue>10</issue>
          ),
          <volume>1345</volume>
          {
          <fpage>1359</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , Io e, S.,
          <string-name>
            <surname>Vanhoucke</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alemi</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>Inception-v4, inception-resnet and the impact of residual connections on learning</article-title>
          .
          <source>In: Thirty-First AAAI Conference on Arti cial Intelligence</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanhoucke</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , Io e, S.,
          <string-name>
            <surname>Shlens</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wojna</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Rethinking the inception architecture for computer vision</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>2818</volume>
          {
          <issue>2826</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tzeng</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            <given-names>man</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Saenko</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Adversarial discriminative domain adaptation</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>7167</volume>
          {
          <issue>7176</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allem</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Unger</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz</surname>
          </string-name>
          , T.B.:
          <article-title>Automated identi cation of hookahs (waterpipes) on instagram: an application in feature extraction using convolutional neural network and support vector machine classi cation</article-title>
          .
          <source>Journal of Medical Internet Research</source>
          <volume>20</volume>
          (
          <issue>11</issue>
          ),
          <year>e10513</year>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          :
          <article-title>Modi ed distribution alignment for domain adaptation with pre-trained inception resnet</article-title>
          . arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>02322</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          :
          <article-title>Impact of imagenet model selection on domain adaptation</article-title>
          .
          <source>In: Proceedings of the IEEE Winter Conference on Applications of Computer Vision Workshops</source>
          . pp.
          <volume>173</volume>
          {
          <issue>182</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Domain-symmetric networks for adversarial domain adaptation</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>5031</volume>
          {
          <issue>5040</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          :
          <article-title>Transductive learning via improved geodesic sampling</article-title>
          .
          <source>In: Proceedings of the 30th British Machine Vision Conference</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          :
          <article-title>Domain adaptation for object recognition using subspace sampling demons</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Zoph</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasudevan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shlens</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          :
          <article-title>Learning transferable architectures for scalable image recognition</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>8697</volume>
          {
          <issue>8710</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>