<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Knowledge Augmented Language Models for Causal Question Answering ?</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>SFI Centre for Research Training in Arti cial Intelligence, Data Science Institute, National University of Ireland - Galway</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>17</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>The task of causal question answering broadly involves reasoning about causal relations and causality over a provided premise. Causal question answering can be expressed across a variety of tasks including commonsense question answering, procedural reasoning, reading comprehension, and abductive reasoning. Transformer-based pretrained language models have shown great promise across many natural language processing (NLP) applications. However, these models are reliant on distributional knowledge learned during the pretraining process and are limited in their causal reasoning capabilities. Causal knowledge, often represented as cause-e ect triples in a knowledge graph, can be used to augment and improve the causal reasoning capabilities of language models. There is limited work exploring the e cacy of causal knowledge for question answering tasks. We consider the challenge of structuring causal knowledge in language models and developing a uni ed model that can solve a broad set of causal question answering tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>causal reasoning • causal question answering • language models • causal knowledge graphs</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Historically, research on causal reasoning in natural language processing (NLP)
has primarily focused on causal relation identi cation and extraction. Recently,
there has been an emerging interest in more complex applications of causal
reasoning, especially around question answering and several new benchmark tasks
were released. The task of causal question answering can be expressed across a
variety of tasks. Broadly these tasks involve reasoning about causality and causal
relations over a provided premise. Pretrained transformer-based language
models such as BERT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and RoBERTa [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] have been found to be generally e ective
for these tasks. However, the distributional knowledge contained in these models
is opaque and reliant on the quality and depth of scope of the pretraining corpus.
Additionally, it is unclear the extent to which language models support causal
reasoning. We hypothesize that causal knowledge can help language models
represent causality and identify causal relations necessary for downstream question
answering. Causal facts, often extracted from causal descriptions and expressed
as cause-e ect triples, can succinctly express causal knowledge. We consider the
challenge of structuring causal knowledge in language models and developing a
uni ed causal language model that can be e ective across all causal reasoning
tasks.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2 Importance</title>
      <p>
        Causal reasoning has a long and rich history rooted in philosophy,
psychology, and many other academic disciplines. Psychologists and philosophers have
posited that causal reasoning is critical to our mental models of reality and that
our knowledge is de ned through identifying causal chains over our observations
of the world [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In the context of natural language applications, causal reasoning
can allow us to produce new knowledge from disparate observations and explore
various hypotheses. For example, causal search engines in the clinical domain
aim to identify causal factors that can help develop new drugs and diagnose rare
medical conditions. Causal question answering systems can be used to better
understand the causes of observed events and explore counterfactuals. Language
models augmented with causal knowledge can be used to develop more
transparent and explainable AI systems. At inference time these model could produce
causal explanations of the predicted answer.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        There is limited historical work on causal question answering with external causal
knowledge. Hassanzadeh et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and Kayesh et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] consider the problem
of binary question answering which poses questions about causes and e ects as
yes/no questions. Extracted cause-e ect pairs are scored using a mixture of
cooccurrence statistics and cosine-similarity scores over BERT embeddings. These
scores are then evaluated against a threshold to answer yes/no for an input
question. Sharp et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and Xie and Mu [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] consider the task of answer
reranking for open-ended causal question answering. Both papers are evaluated
on a set of causal questions extracted from the Yahoo! Answers corpus, which
follow the patterns What causes ... and What is the result of .... [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Sharp et
al. present three distributional similarity models (adapted Skipgram,
monolingual alignment, and a convolutional neural network) to model the contextual
relationship between cause and e ect phrases. The answer choices are re-ranked
based on the cosine similarity between extracted cause and e ect vectors. Our
CausalSkipgram model for representing causal knowledge expands upon the
adapted Skipgram model presented by Sharp et al.
      </p>
      <p>
        Next, we summarize the current causal question answering benchmark datasets.
ROPES (reasoning over paragraph e ects in situations) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is a reading
comprehension dataset where the goal is to use causal relationships expressed in a
background passage to answer questions about a hypothetical premise. CosmosQA [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
is a multiple-choice reading comprehension challenge where the aim is to answer
questions concerning likely causes or e ects of events that require commonsense
knowledge outside of the provided context. COPA (Choice of Plausible
Alternatives) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a multiple-choice question answering task where the goal is to
identify which alternative is the likely cause or e ect of a provided premise. WIQA
(What if question answering) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is another multiple-choice task that aims to
reason about the magnitude e ects of perturbations to procedural descriptions
of events. aNLI (Abductive Natural Language Inference) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a multiple-choice
task where the goal is to identify which of provided hypotheses best explains a
provided context. Our preliminary work has primarily focused on the COPA and
WIQA datasets as they allow the most direct evaluation of causal knowledge in
the context of multiple-choice causal question answering.
      </p>
      <p>
        Finally, we summarize sources of causal knowledge. CauseNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is currently
the largest publicly available knowledge graph of claimed causal facts. It
contains over 11 million relations and 12 million concepts that were extracted from
Wikipedia and ClueWeb [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. ConceptNet [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], a public knowledge graph,
consists of 36 relations and includes a causes relation. The ATOMIC knowledge [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
graph focuses on knowledge for commonsense inference. ATOMIC is organized
around if-then relations that primarily describe the relations and interactions
around human-centric activities. We use CauseNet as the primary source for
causal knowledge in our experiments.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Research Questions</title>
      <p>RQ1 Does incorporating structured causal knowledge into language
models improve performance on causal question answering tasks?</p>
      <p>Current model-based solutions have converged on ne-tuning pretrained
language models on task-speci c datasets. These approaches rely on the
transferability of distributional knowledge learned during the pretraining process. To
the best of our knowledge, there is no empirical research that demonstrates the
e cacy of external causal knowledge in the context of causal question answering.
Our work aims to establish those baselines.</p>
      <p>RQ2 What is the most e ective way of representing causal
knowledge in language models for causal question answering tasks?</p>
      <p>RQ2.1 How can language models be augmented with external knowledge
from causal knowledge graphs for downstream causal question answering tasks?</p>
      <p>RQ2.2 How can causal knowledge be injected into the language model during
the pretraining process such that it is available as transferable distributional
knowledge for downstream causal question answering tasks?</p>
      <p>Augmenting language models with structured knowledge is an emerging area
of research. Our research aims to provide a methodology for representing causal
knowledge that used by language models in causal question answering tasks.
We consider the strategies of knowledge augmentations and knowledge
injection. The knowledge augmentation approach (RQ2.1 ) aims to train the language
model to consider external knowledge provided as input features during
prediction time. The knowledge injection approach (RQ2.2 ) aims to convert causal
knowledge into structured distributional knowledge during pretraining process
to produce causality-aware language models (CALM) which would ideally
support any downstream causal question answering task.</p>
      <p>RQ3 How do we evaluate the causal reasoning capabilities of
language models in the context of question answering?</p>
      <p>
        To date, most research on causal reasoning applications in NLP focuses on
task-speci c model implementations. There is no comprehensive de nition of
causal question answering nor a uni ed way to evaluate a language model's cause
reasoning capabilities. We hope to contextualize causal question answering as an
extension of fundamental NLP problems and produce a uni ed benchmark
(similar in spirit to GLUE (General Language Understanding Evaluation) benchmark
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]).
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Preliminary Results</title>
      <p>
        Our experiments explore the e cacy augmenting RoBERTa [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] with causal
knowledge for multiple-choice question answering on the COPA and WIQA
benchmark tasks. Causal facts are extracted from CauseNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and selected
based on the lexical overlap between the cause-e ect concepts and question text.
Additional details on knowledge selection and experiment results can be found
in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>COPA consists of a premise and two alternatives. The task is to identify
which alternative is most likely the cause or e ect of the provided premise.
Background commonsense causal knowledge is required to answer questions as
there is limited lexical overlap between the premise and alternatives. WIQA
consists of multiple-choice questions where the answer options (more, less, and
no effect) describe the magnitude e ect of a proposed perturbation to a
procedural event. Each question has an associated procedural description consisting
of a sequence of events. The question proposes a perturbation to a speci c event
and asks what impact that perturbation would have on another event in the
procedural description.</p>
      <p>
        Next, we present three strategies for representing causal knowledge to a
language model. The most direct way to incorporate causal information is to append
it to the end of the input text, which we call the InputAugmentation method.
Relevant causal tuples are converted into causal statements which follow the
pattern C causes E. CausalSkipgram adapts the skip-gram word embedding
approach [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to model causal pairs. The last method is CausalKGE, which
represents causal knowledge as a knowledge graph embedding. We adapt the TransE
model presented by Bordes et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. To model our causal tuples as a knowledge
graph, we add the explicit relation "cause-e ect" to each tuple. The modeling
goal of TransE is thus to predict an e ect E, given a cause C and "cause-e ect"
CR such that C + CR E. A causal triple is represented by a single vector
which is generated by mean pooling the head, tail, and relation vectors.
      </p>
      <p>To incorporate causal embeddings with RoBERTa, we propose the Causality
Enhanced RoBERTa neural architecture (Figure 1). This architecture is used with
both the CausalSkipgram and CausalKGE embeddings. The rst layer is the
causality-enhanced input layer which combines the pooled embedding output
of RoBERTa with the causal knowledge embeddings. For inputs where we were
able to extract causal facts, the causal embedding vector is generated by
concatenating and attening all relevant causal embeddings. Up to ve causal facts
are selected per input. The RoBERTa pooled output is then concatenated with
causal embeddings. This input is next passed into a FeedForward Network (FFN)
with a hidden layer and classi er
6</p>
    </sec>
    <sec id="sec-6">
      <title>Evaluation</title>
      <p>
        Table 1 provides the results of our experiments on the COPA test set and
the COPA-Balanced Hard set. Recent pretrained models such as BERT and
RoBERTa have seen improved performance on the COPA dataset. However,
Pride et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] found that these models exploited super cial cues such as the
token frequency in the correct answers. To mitigate this e ect, Pride et al.
expanded the development set to include mirror instances to balance the lexical
distribution between correct and incorrect answers. This new dataset, called
COPA-Balanced, also categorized the test set into easy and hard groups. The
easy group consists of 190 questions where RoBERTa-Large and BERT-Large
could answer correctly without the provided premise and the hard group is the
remaining 310 questions. We use the COPA-Balanced development set for
training and the hard category (which we will refer to as COPA-Balanced Hard) for
evaluation. For the COPA test set, we were able to extract causal information
from CauseNet for 32% of the questions. All three causal augmentation methods
outperform the RoBERTa baseline. The CausalKGE and Input Augmentation
have similar performance, improving accuracy on average by +6.0pp and +3.9pp
over the RoBERTa baseline on the COPA test set and COPA-Balanced Hard
set.
7
      </p>
    </sec>
    <sec id="sec-7">
      <title>Discussion and Future Work</title>
      <p>Our initial work validates (RQ1 ) by demonstrating the e cacy of causality
enhanced language models on the COPA and WIQA question answering
benchmarks. Further work will explore improving recall on causal fact selection from
CauseNet and more sophisticated techniques to reduce the selection of
irrelevant facts. We also plan to explore knowledge injections techniques described in
RQ2.2. We are investigating adapting the masked-language modeling objective
to predict masked causal concepts across sentence-level descriptions of causal
events.</p>
      <p>Broadly, our goal is to develop a uni ed causal knowledge-enhanced language
model that can be e ective across all causal reasoning tasks. To that extent
we need to be able to de ne and measure the causal reasoning capabilities of
the language model (RQ3 ). While COPA and WIQA are both multiple-choice
question answering tasks, the causal reasoning requirements are distinct for each
application. Yet we nd our simple augmentation strategies are e ective in both
cases. This raises interesting questions about how language models are using
causal knowledge and if our current tasks accurately represent causal reasoning.
We hope to create a more meaningful de nition of causal reasoning by exploring
the causal knowledge needs of existing NLP tasks and develop new probing
methods to better understand the causal reasoning capabilities of these language
models. We hope this work is a stepping stone towards the more ambitious goals
of general AI with causal reasoning capabilities. In order to make that jump,
language models need to be able reason across semantic knowledge found on the
web and in causal graphs.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work has been funded with the nancial support of the Science Foundation
Ireland Centre for Research Training in Arti cial Intelligence under Grant No.
18/CRT/6223 and is supervised by Dr. Paul Buitelaar and Dr. Mihael Arcan.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bhagavatula</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bras</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malaviya</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakaguchi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holtzman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rashkin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Downey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yih</surname>
            ,
            <given-names>S.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Abductive commonsense reasoning</article-title>
          .
          <source>CoRR</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Usunier</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Duran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yakhnenko</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Translating embeddings for modeling multi-relational data</article-title>
          . In: Burges,
          <string-name>
            <given-names>C.J.C.</given-names>
            ,
            <surname>Bottou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Welling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Ghahramani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Weinberger</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.Q</surname>
          </string-name>
          . (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          . vol.
          <volume>26</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dalal</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arcan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Enhancing multiple-choice question answering with causal knowledge. In: Proceedings of Deep Learning Inside Out (DeeLIO): The Second Workshop on Knowledge Extraction and Integration for Deep Learning Architectures. Association for Computational Linguistics (</article-title>
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          : BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In: NAACL 2019 Proceedings: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers)
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Goldman</surname>
            ,
            <given-names>A.I.:</given-names>
          </string-name>
          <article-title>A causal theory of knowing</article-title>
          .
          <source>The Journal of Philosophy</source>
          <volume>64</volume>
          (
          <issue>12</issue>
          ),
          <volume>357</volume>
          {
          <fpage>372</fpage>
          (
          <year>1967</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gordon</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozareva</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roemmele</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning</article-title>
          . In: SemEval@
          <string-name>
            <surname>NAACL-HLT</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hassanzadeh</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhattacharjya</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feblowitz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srinivas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrone</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sohrabi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Answering binary causal questions through large-scale text mining: An evaluation using cause-e ect pairs from human experts</article-title>
          .
          <source>In: Proceedings of the Twenty-Eighth International Joint Conference on Arti cial Intelligence</source>
          , IJCAI-19. International Joint Conferences on Arti cial Intelligence Organization
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Heindorf</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scholten</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wachsmuth</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Causenet:
          <article-title>Towards a causality graph extracted from the web</article-title>
          .
          <source>In: CIKM</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bras</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhagavatula</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Cosmos</surname>
            <given-names>QA</given-names>
          </string-name>
          <article-title>: machine reading comprehension with contextual commonsense reasoning</article-title>
          .
          <source>CoRR</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kavumba</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inoue</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heinzerling</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reisert</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inui</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>When choosing plausible alternatives, clever hans can be clever</article-title>
          .
          <source>In: Proceedings of the First Workshop on Commonsense Inference in Natural Language Processing</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kayesh</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saiful</surname>
            <given-names>Islam</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Anirban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Kayes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.S.M.</given-names>
            ,
            <surname>Watters</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Answering binary causal questions: A transfer learning based approach</article-title>
          . In: 2020
          <source>International Joint Conference on Neural Networks (IJCNN)</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tafjord</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gardner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Reasoning over paragraph e ects in situations</article-title>
          .
          <source>In: MRQA@EMNLP</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoyanov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Roberta: A robustly optimized BERT pretraining approach</article-title>
          .
          <source>CoRR</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient estimation of word representations in vector space</article-title>
          .
          <source>In: ICLR</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Rajagopal</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tandon</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalvi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          , E.:
          <article-title>What-if I ask you to explain: Explaining the e ects of perturbations in procedural text</article-title>
          .
          <source>In: Findings of the Association for Computational Linguistics: EMNLP</source>
          <year>2020</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Sap</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>LeBras</given-names>
            , R.,
            <surname>Allaway</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Bhagavatula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Lourie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Rashkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Roof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.A.</given-names>
            ,
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>Atomic: An atlas of machine commonsense for if-then reasoning (</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sharp</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Surdeanu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jansen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hammond</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Creating causal embeddings for question answering with minimal supervision</article-title>
          .
          <source>In: ACL 2016 Proceedings of Conference on Empirical Methods in Natural Language Processing</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Speer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Havasi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Conceptnet 5.5: An open multilingual graph of general knowledge</article-title>
          .
          <source>In: Proceedings of the Thirty-First AAAI Conference on Arti cial Intelligence</source>
          . AAAI Press (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Tandon</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mishra</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakaguchi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosselut</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Wiqa: A dataset for "what if</article-title>
          ...
          <article-title>" reasoning over procedural text</article-title>
          .
          <source>In: EMNLP/IJCNLP</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michael</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.R.:</given-names>
          </string-name>
          <article-title>GLUE: A multi-task benchmark and analysis platform for natural language understanding (</article-title>
          <year>2019</year>
          ),
          <source>in the Proceedings of ICLR.</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Distributed representation of words in cause and e ect spaces</article-title>
          .
          <source>In: AAAI</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>