<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Reduction by Selecting a Hierarchical Order of Implicit Author Demographic Characterizations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chung-Chi Chen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hen-Hsen Huang</string-name>
          <email>hhhuang@iis.sinica.edu.tw</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hsin-Hsi Chen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Japan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Information Engineering, National Taiwan University</institution>
          ,
          <country country="TW">Taiwan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>In: M. Litvak, I.Rabaev, R. Campos, A. Jorge, A. Jatowt (eds.): Proceedings of the IACT'23 Workshop</institution>
          ,
          <addr-line>Taipei</addr-line>
          ,
          <country country="TW">Taiwan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Information Science</institution>
          ,
          <addr-line>Academia Sinica</addr-line>
          ,
          <country country="TW">Taiwan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>its eficacy. This paper focuses on the selection of hierarchical orders in multi-task architectures, a significant challenge in developing neural network architectures. We propose a systematic methodology based on the statistical results of the Apriori algorithm to arrange the order of co-training tasks. Our findings demonstrate that this approach can provide near-optimal performance, significantly reducing the exploration times in multi-task scenarios. The models developed using this methodology surpass state-of-the-art performances in flu vaccination intent prediction and music review sentiment analysis tasks, demonstrating Hierarchical order, demographic characterization, Exploration reduction,</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The development of neural network architectures frequently necessitates a significant degree
of trial-and-error, in addition to substantial computational time. State-of-the-art models across
various tasks often require thousands of GPU days for training [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The environmental and
computational expenses associated with such complex models present serious challenges [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Therefore, developing guidelines to enhance performance and reduce iterative testing becomes
a critical area of exploration.
      </p>
      <p>
        A key area of computational intensity arises when probing the hierarchical order of a
hierarchical multi-task architecture. Given four candidate tasks within an architecture, to obtain
optimal performance, we would need to experiment with all 24 possible hierarchical orders.
Previous works have demonstrated the eficacy of multi-task architectures in natural language
processing (NLP) tasks [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
        ]. However, the literature remains sparse in providing insights
on choosing the hierarchical order for these architectures. This paper attempts to bridge this
gap, ofering a systematic analysis for selecting an optimal hierarchical order for multi-task
architectures.
      </p>
      <p>
        The crux of multi-task architectures lies in sharing learned embeddings and information
across tasks. We propose an approach based on the Apriori algorithm’s statistical results [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
to organize the order of the co-training tasks. Our findings suggest that the hierarchical
order proposed by our approach achieves near-optimal performance, significantly reducing the
number of required exploration iterations.
      </p>
      <p>
        The contributions of this paper are three-fold:
1. We highlight a crucial intersection between sustainable NLP and multi-task learning.
2. We propose an eficient method for arranging the hierarchical order in multi-task
architecture, ofering near-optimal performance with fewer explorations.
3. The models developed using our methodology outperform state-of-the-art performances
in flu vaccination intent prediction [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and music review sentiment analysis [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] tasks.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        The human learning process often involves sharing information or experience across tasks,
an idea reflected in multi-task learning architecture. This concept has seen success in diverse
applications, such as computer vision [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and NLP [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. However, decisions regarding what
and how to share remain open questions. A comprehensive survey of multi-task learning is
provided by Zhang and Yang [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The field typically bifurcates into hard-sharing and
softsharing methods, as overviewed by Ruder [13]. In both cases, a majority of previous works
have focused on information sharing within the encoder [14, 15, 16]. This paper instead ofers a
guideline for information sharing by designing a hierarchical architecture, particularly focusing
on the learning order selection. Bidirectional Encoder Representations from Transformers
(BERT) have revolutionized the NLP field [ 17]. Researchers have leveraged BERT and other
pre-trained text-encoders to set new standards on several NLP tasks [18, 19, 20]. This work
uses BERT as an encoder and delves further into the issue of hierarchical order selection.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets</title>
      <p>Our experiments utilize two publicly available datasets. Each dataset contains four labels for a
single input sample, therefore, we treat our task setting as a four-label classification problem.</p>
      <p>
        The first dataset, denoted as Twitter FV, is sourced from Twitter [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The corresponding
task involves predicting whether the author of a tweet has already received a flu vaccination
or intends to do so. The second dataset, referred to as Amazon Sentiment, is collected from
      </p>
      <p>
        Amazon music reviews [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The task for this dataset is to label given reviews as either positive
or negative. Huang and Paul [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] identified a correlation between the writer’s demographic
factors (region, gender, and age) and the tasks of flu vaccination intent and sentiment analysis.
Accordingly, we incorporate these demographic features from their dataset as labels for auxiliary
tasks. The statistics for both the Twitter FV and Amazon Sentiment datasets are detailed in
Table 1 and Table 2, respectively.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <sec id="sec-4-1">
        <title>4.1. Models</title>
        <p>For individual task training performance testing, we adopt BERT-Large [17], a 24-layer
Transformer [21]. We compare the standard multi-task architecture depicted in Figure 1 (a) with the
hierarchical multi-task architecture shown in Figure 1 (b). We preprocess the input text using
WordPiece [22] to obtain the input embedding,  . Post BERT-Large encoding, we acquire the
token embedding  ∈ ℝ 1024. Following the classification task setup in previous work [ 17], we
utilize the first token embedding of an input instance (  1) to represent the encoded
information. Subsequently, a one-layer perceptron is adopted for decoding and prediction. The Adam
optimizer [23] is used for stochastic optimization, employing the cross-entropy loss function.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Hierarchical Order Selection Approach</title>
        <p>
          In a hierarchical multi-task architecture, the optimal task order selection remains an open
question. Exhaustively exploring all hierarchical orders to achieve the best performance is an
obvious yet highly ineficient approach. This paper proposes a method based on the Apriori
algorithm [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] for selecting the hierarchical order.
        </p>
        <p>The Apriori algorithm is typically employed for association rule learning, with market basket
analysis being a common application. Using this algorithm, given an item in a customer’s basket
(for instance, a bottle of milk), we can calculate the probability of another item (like cereal) also
being included in the basket, based on previous transaction statistics. In our method, we treat
each label as an individual item and compute the  
of a given label towards other labels as
defined in Equation 1.</p>
        <p>(
 ,   ) =
(
(</p>
        <p>∩   )
 ) × (
)

Here,  (⋅)</p>
        <p>denotes the frequency of the given label set in the dataset, and  denotes the
label set. Table 3 presents the Lift between diferent labels in each dataset.</p>
        <p>To estimate the informativeness of each auxiliary task, we further calculate the
informativeness score ( - 
) using Equation 2.</p>
        <p>-( 
 |  ) =</p>
        <p>∑=1 ∣  (
  ,   ) − 1 ∣


In this equation,  denotes the number of labels in  
 . The principle behind the  - 
that whether the target label has a positive (&gt; 1) or negative (&lt; 1) correlation to   , the further
it is from 1, the more information the target label provides. Table 4 displays the  - 
for each
dataset.</p>
        <p>We recommend arranging the auxiliary tasks in an ascending order of  - 
. The proposed
approach’s suggested hierarchical orders for both datasets are listed in Table 5. Consequently,
given the order of auxiliary tasks, we only need to explore four hierarchical orders instead of
probing all possible 24 combinations.
(1)
(2)
is</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <p>
        Huang and Paul [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] divide the dataset into training and test sets randomly but do not specify
the indices for splitting. Given Gorman and Bedrick’s findings [ 24], single ”standard split”
results may not be reliable. Thus, we employ five-fold cross-validation to gauge each model’s
performance. To ensure reproducibility, the splitting indices are provided in the supplementary
materials.1 We report the average macro-F1 score. Since Huang and Paul [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] use the weighted
F1-score on the Twitter FV and Amazon Sentiment datasets, we also report results using this
metric.
      </p>
      <sec id="sec-5-1">
        <title>5.1. Comparison with Baselines</title>
        <p>
          In this section, we juxtapose the performance of the best hierarchical model against other
baseline models. For the Twitter FV and Amazon Sentiment datasets, we utilize NUFA+w [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
as a baseline, a BiLSTM-based multi-task architecture. We also employ single-task BERT as
1http://explorationreduction.nlpfin.com/
a robust baseline for all datasets. Table 6 and Table 7 present the experimental results, also
detailing the hierarchical order. The parenthesized information denotes the hierarchical order
from Task 1 to Task 4. We observe that the hierarchical architecture consistently outperforms
across all datasets.
        </p>
        <p>Interestingly, the vanilla multi-task architecture’s performance trails behind that of
singletask BERT in both the Twitter FV and Amazon Sentiment datasets. This finding suggests that
when the encoder is merely fine-tuned without shared information between the task-specific
components, some task-specific information may not be efectively learned by the models.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Performance with Suggested Orders</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <sec id="sec-6-1">
        <title>6.1. Correlation between the Tasks</title>
        <p>The performance of the hierarchical multi-task architecture is found to be on par with
singletask BERT according to Table 6. However, Table 7 shows a significant diference between the
performances of the hierarchical multi-task architecture and single-task BERT. We provide an
in-depth analysis of this phenomenon in this section.</p>
        <p>
          We execute ordinary least squares regression (OLS) on the task performances in both datasets.
The OLS inputs consist of experimental results from all hierarchical orders, resulting in 24
distinct experimental results. Table 9 presents the statistics. We find that the performance
of flu vaccination intent detection is not significantly correlated with the performance of all
auxiliary tasks. This indicates that while demographic information proves useful in
BiLSTMbased architectures [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], it may not be beneficial for the flu vaccination intent detection task in
a BERT hierarchical multi-task architecture. Conversely, demographic information remains
useful for music review sentiment analysis, as performance improvements in auxiliary tasks
also enhance the sentiment analysis task’s performance.
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Limitations</title>
        <p>While our study ofers initial insights into optimizing hierarchical order selection in multi-task
architectures, it does have certain limitations. First, we did not analyze all multi-task setting
datasets due to the sheer volume of possibilities. Another potential limitation is our omission of
GPU cost calculations in our study. However, the inferred reduction in carbon dioxide emissions
stemming from decreased exploratory iterations is an important point to note. Assuming
identical datasets and models, our proposed method implies that nearly optimal performance
can be achieved with only one-sixth of the carbon dioxide emissions associated with exhaustive
exploration. Lastly, the performance correlation between tasks observed in our study could
be context-specific. As demonstrated, the demographic information was useful in the context
of music review sentiment analysis but not in the flu vaccination intent detection task. This
ifnding suggests that not all information may be universally useful across diferent tasks in a
hierarchical multi-task learning setup.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions</title>
      <p>This paper presented a systematic analysis for selecting an optimal hierarchical order for
multitask architectures, addressing a gap in the literature. We proposed an approach based on the
Apriori algorithm to organize task order and demonstrated that the resulting hierarchical order
achieves near-optimal performance while considerably reducing the number of exploration
iterations.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Acknowledgments</title>
      <p>This research is supported by National Science and Technology Council, Taiwan, under grants
110-2221-E-002-128-MY3, 110-2634-F-002-050-, and 111-2634-F-002-023-. The work of
ChungChi Chen was supported in part by JSPS KAKENHI Grant Number 23K16956.
[13] S. Ruder, An overview of multi-task learning in deep neural networks, arXiv preprint
arXiv:1706.05098 (2017).
[14] S. Maharjan, J. Arevalo, M. Montes, F. A. González, T. Solorio, A multi-task approach
to predict likability of books, in: Proceedings of the 15th Conference of the European
Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, 2017,
pp. 1217–1227.
[15] R. Masumura, Y. Shinohara, R. Higashinaka, Y. Aono, Adversarial training for multi-task
and multi-lingual joint modeling of utterance intent classification, in: Proceedings of the
2018 Conference on Empirical Methods in Natural Language Processing, 2018, pp. 633–639.
[16] M. S. Akhtar, D. Chauhan, D. Ghosal, S. Poria, A. Ekbal, P. Bhattacharyya, Multi-task
learning for multi-modal emotion recognition and sentiment analysis, in: Proceedings of
the 2019 Conference of the North American Chapter of the Association for Computational
Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp.
370–379.
[17] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional
transformers for language understanding, in: NAACL, 2019, pp. 4171–4186.
[18] C. Alberti, D. Andor, E. Pitler, J. Devlin, M. Collins, Synthetic QA corpora generation with
roundtrip consistency, in: ACL, 2019, pp. 6168–6173.
[19] K. Clark, M.-T. Luong, U. Khandelwal, C. D. Manning, Q. V. Le, BAM! born-again multi-task
networks for natural language understanding, in: ACL, 2019, pp. 5931–5937.
[20] J. Straková, M. Straka, J. Hajic, Neural architectures for nested NER through linearization,
in: ACL, 2019, pp. 5326–5331.
[21] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I.
Polosukhin, Attention is all you need, in: Advances in neural information processing systems,
2017, pp. 5998–6008.
[22] Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao,
K. Macherey, et al., Google’s neural machine translation system: Bridging the gap between
human and machine translation, arXiv preprint arXiv:1609.08144 (2016).
[23] D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, arXiv preprint
arXiv:1412.6980 (2014).
[24] K. Gorman, S. Bedrick, We need to talk about standard splits, in: Proceedings of the 57th
Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 2019.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <surname>DARTS</surname>
          </string-name>
          :
          <article-title>Diferentiable architecture search</article-title>
          , in: International Conference on Learning Representations,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Strubell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ganesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <article-title>Energy and policy considerations for deep learning in NLP, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</article-title>
          , Florence, Italy,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>Gao, Multi-task deep neural networks for natural language understanding</article-title>
          ,
          <source>in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4487</fpage>
          -
          <lpage>4496</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Masumura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shinohara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Higashinaka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Aono</surname>
          </string-name>
          ,
          <article-title>Adversarial training for multi-task and multi-lingual joint modeling of utterance intent classification</article-title>
          , in: EMNLP, Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>633</fpage>
          -
          <lpage>639</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruder</surname>
          </string-name>
          ,
          <article-title>A hierarchical multi-task approach for learning embeddings from semantic tasks</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>6949</fpage>
          -
          <lpage>6956</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hashimoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tsuruoka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <article-title>A joint many-task model: Growing a neural network for multiple NLP tasks</article-title>
          , in: EMNLP, Copenhagen, Denmark,
          <year>2017</year>
          , pp.
          <fpage>1923</fpage>
          -
          <lpage>1933</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Srikant</surname>
          </string-name>
          , et al.,
          <article-title>Fast algorithms for mining association rules</article-title>
          ,
          <source>in: Proc. 20th int. conf. very large data bases, VLDB</source>
          , volume
          <volume>1215</volume>
          ,
          <year>1994</year>
          , pp.
          <fpage>487</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Paul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ryzhkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Quinn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Broniatowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dredze</surname>
          </string-name>
          ,
          <article-title>Examining patterns of influenza vaccination in social media</article-title>
          ,
          <source>in: Workshops at the ThirtyFirst AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Paul</surname>
          </string-name>
          ,
          <article-title>Neural user factor adaptation for text classification: Learning to generalize across author demographics</article-title>
          ,
          <source>in: Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (* SEM</source>
          <year>2019</year>
          ),
          <year>2019</year>
          , pp.
          <fpage>136</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Johns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Davison</surname>
          </string-name>
          ,
          <article-title>End-to-end multi-task learning with attention</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1871</fpage>
          -
          <lpage>1880</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Luong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <article-title>Kaiser, Multi-task sequence to sequence learning</article-title>
          , in: Y. Bengio, Y. LeCun (Eds.),
          <source>4th International Conference on Learning Representations, ICLR</source>
          <year>2016</year>
          , San Juan, Puerto Rico, May 2-
          <issue>4</issue>
          ,
          <year>2016</year>
          , Conference Track Proceedings,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>A survey on multi-task learning</article-title>
          ,
          <source>arXiv preprint arXiv:1707.08114</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>