<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting Themes within Complex Unstructured Texts: A Case Study on Safeguarding Reports</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksandra Edwards</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Rogers</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose Camacho-Collados</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Helene de Ribaupierre</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alun Preece</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Crime and Security Research Institute,Cardi University</institution>
          ,
          <addr-line>Cardi</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science and Informatics, Cardi University</institution>
          ,
          <addr-line>Cardi</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Text classi cation typically requires large amounts of labelled training data; however, the acquisition of high volumes of labelled datasets is often expensive or unfeasible, especially for highly-specialised domains for which both training data (documents) and access to subject-matter expertise (for labelling) is limited. Language models pre-trained on large amounts of text corpora provide state-of-the-art performance against most standard natural language processing (NLP) benchmarks, including text classi cation. However, their prevalence over more traditional linear classi ers and domain-based approaches have not been investigated fully. In this paper, we address the combination of state-of-the-art deep learning and classi cation methods and provide an insight into what combination of methods t the needs of small, domain-speci c, and terminologicallyrich corpora. We focus on a real-world scenario related to a collection of safeguarding reports comprising learning experiences and re ections on tackling serious incidents involving children and vulnerable adults. Our aim is to automatically identify the main themes in a safeguarding report using three main types of classi cation. Our results show that for a very small amount of data a simple linear classi er outperforms state-of-the-art language models. Further, we show that the performance of classi ers is more a ected by the size of the training data rather than the amount of context given.</p>
      </abstract>
      <kwd-group>
        <kwd>text classi cation small domain corpus language models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The performance of natural language processing (NLP) classi cation tasks is
heavily reliant on the amount of training data available [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. However, the
acquisition of high volumes of labelled data can be an expensive, time- and
resource-consuming process [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], especially when the text to be labelled is in a
highly-specialised domain where only scarce domain experts can perform the
manual labelling task [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Current pre-trained neural models such as BERT
(Bidirectional Encoder Representations from Transformers) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] proved to
provide state-of-the-art results in most standard NLP benchmarks [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], including
text classi cation. However, the applicability of these language models to very
small collections of highly specialised documents has not been fully explored or
compared to more traditional methods. A limitation to pre-trained models is
that there is still a need for task-speci c datasets for these models to perform
well in a speci c domain [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Therefore, adapting these large but generic models
to speci c domains and tasks has become the new standard approach for many
NLP problems [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. For instance, the authors of [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] provide a more extensive
research on whether it is still helpful to tailor a pretrained model to the domain
of a target task. However, this research is not focused on text classi cation and
does not compare neural models to other types of machine learning models.
Further, a recent research [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] on few-shot classi cation, analysed the role of
labeled and unlabeled data for classi ers by comparing a linear model such
as fastText coupled with domain-speci c embeddings against ne-tuned BERT
model using both domain-speci c and generic corpora. However, the authors
performed analysis using generic datasets assuming the presence of large amounts
of unlabeled data which can be used for ne-tuning models on domain data.
We build on this research [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] by comparing the performance of three types of
classi ers for a real-world scenario related to the safeguarding domain where
there is very limited amount of labeled and unlabeled dataset. Previous work on
performing NLP analysis on the safeguarding corpus emphasized the challenges
of extracting knowledge from the documents using o -the-shelf text analysis tools
due to the highly specialised lexical characteristics of the reports [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Further,
there is no existing knowledge resources which t the needs of the domain [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
which makes the use of semantic enrichment approaches for the domain di cult.
      </p>
      <p>We look at whether domain-trained embeddings are e ective even when
trained on very limited corpus. Our main contribution is that we conduct a
thorough analysis of what combination of embedding and language models and
classi cation approaches t the needs of a small domain-speci c and
terminologyrich corpus. Further, we also look at how deep learning approaches are a ected
by training dataset size versus the amount of context given.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Case study: Safeguarding reports</title>
      <p>
        The purpose of a safeguarding report is to identify and describe related events
that precede a serious safeguarding incident | for example, involving a child
or vulnerable adult | and to re ect on agencies' roles and the application of
best practices. Each report contains key information about learning experiences
and re ections on tackling serious incidents. The reports carry great potential to
improve multi-agency work and help develop better safeguarding practices and
strategies [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Analyzing and understanding safeguarding reports is crucial for
health and social care agencies; in particular, a key task is to identify common
themes across a set of reports. Traditionally, this is done in social science by a
process of manually annotating the reports with themes identi ed by
subjectmatter experts using a qualitative analysis tool such as NVivo. However, each
report is lengthy and complex, so manual annotation is a time-consuming and
potentially bias-prone process [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Furthermore, in our particular case, the
safeguarding collection is expected to grow signi cantly in the near future, with the
additional resourcing of 500 historical reports, making the manual annotation
of these additional documents unfeasible. Therefore, we aim to automate the
process of document annotation.
      </p>
      <p>Furthermore, in our particular case, the safeguarding collection is expected
to grow signi cantly in the near future, with the additional resourcing of 500
historical reports, making the manual annotation of these additional documents
unfeasible. Therefore, we aim to automate the process of document annotation.</p>
      <p>
        The thematic framework [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] used for performing document classi cation
resulted from collaborative work between multiple subject-matter experts. In this
context, a theme refers to the main topic of discussion related to safeguarding
incidents, speci cally relevant to domestic homicide and mental health homicide.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>The Dataset</title>
      <p>At the time of development the corpus consisted of 27 full safeguarding reports.
The annotations were carried out by a social science team following standard
methodology in the eld. They used a qualitative analysis tool (NVivo) to label
parts of documents with thematic annotations from 5 top-level themes according
to the thematic framework described in Section 2. The annotation was performed
by labelling di erent-length passages of the reports with themes from the thematic
framework. The majority of report contents were labeled except appendices. The
total number of sentences in the corpus was 3,421 (see Table 1) with unbalanced
distribution between the di erent themes where sentences can be associated with
multiple themes. We evaluated models performance using training, development
(both training and development were randomly sampled from the 27 reports) and
test sets. Both development and test sets were annotated at the passage-level.
The test set was extracted from safeguarding reports di erent from the original 27
documents. The test set contained a 100 randomly selected passages where each
passage consisted of 3 sentences. Due to the limited amount of reports available,
we built and evaluated classi er models on a sentence level (i.e., results presented
in Section 4.2). Thus, each sentence was assigned the label of the passage to
which it belongs. Further, we ensured that the train and development set do
not intersect by automatically selecting random non-overlapping partitions for
the two subsets. We also performed analysis at the passage-level, presented in
Section 5.</p>
    </sec>
    <sec id="sec-4">
      <title>Classi cation Experiments</title>
      <sec id="sec-4-1">
        <title>4.1 Methods</title>
        <p>In our experiments, we perform multi-label classi cation to identify main themes
within documents. We compare three classi ers | a simple count-based classi er,
a linear classi er based on word embeddings, and a state-of-the-art language
model. We perform experiments with pre-trained and corpus-trained embeddings
as well as di erent methods for building feature vectors. We use an n-gram feature
representation and a Naive Bayes classi er as our baseline.</p>
        <p>
          Our method consists of four overall steps, described below.
Step 1: Pre-processing We extracted terms from the corpus using FlexiTerm
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], an open-source software tool for automatic recognition of multi-word terms.
We used the sentences pre-processed with the terminology extraction step for
building sentence embeddings and for creating simple n-gram feature vectors.
Step 2: Feature Extraction (FE) We used fastText word embeddings [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
pretrained with subword information on Common Crawl. We also used fastText for
learning domain-speci c embeddings because it captures the meaning of rare
words better than other approaches. We used the skip-gram method for building
word embeddings with 300 dimensions.
        </p>
        <p>Step 3: Feature Integration (FI) We use several ways for combining the word
embeddings into reduced sentence representations: In the rst approach, we
average the embeddings of each word in a sentence along each dimension. In
the second approach, we assign TF-IDF weights to the words in a sentence, and
calculate the weighted average of the word embeddings along each dimension
(where the contribution of a word is proportional to its TF-IDF weight).</p>
        <p>
          Finally, we use Bidirectional Encoder Representations from Transformers
(BERT) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. A limitation of the word embedding model described above is that
it produces a single vector of a word despite the context in which it appears.
In contrast to the other embedding methods, BERT is designed to pre-train
deep bidirectional representations from unlabeled text by jointly conditioning on
both left and right context in all layers. This characteristic allows the model to
learn the context of a word based on all of its surroundings, thus it generates
more contextually-aware word representations. There are two steps in the BERT
framework: pre-training and ne-tuning. In this step of the methodology, we use
the base pre-trained BERT model, trained on the Books Corpus and English
Wikipedia, for extracting contextualized sentence embeddings. The ne-tuning
step consists of further training on the downstream tasks.
Step 4: Classi ers We perform classi cation on a sentence level where each
sentence had been assigned the theme of the passage the sentence belonged to.
Here, we take `ground truth' to be the annotations made by the social scientist
expert annotators who were involved in creating the thematic framework (see
Section2). For a baseline we use GNB classi er based on frequency-based features
available in Scikit-Learn library [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], since it is considered a strong baseline for
many text classi cation tasks[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. A potential problem with linear classi ers is
that they struggle with OOV words, ne-grained distinctions and unbalanced
datasets. The fastText classi er [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] addresses this problem by integrating a linear
model with a rank constraint, allowing sharing parameters among features and
classes. Further, we ne-tune BERT for the classi cation task using a sequence
classi er, a learning rate of 5e-5 and 4 epochs. In particular, we made use of
the BERT's Hugging Face default transformers implementation for classifying
sentences [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2 Results</title>
        <p>
          We evaluate the performance of the machine learning algorithms by using precision,
recall, and F1-measure metrics. The summary results are calculated using
microand macro- based measures. Early experiments using Word2Vec embeddings [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
and SVM classi er showed unsatisfactory performance compared to fastText
embeddings and GNB classi er. Thus, these results are omitted from Table 2.
        </p>
        <p>The results in Table 2 show that a simple terminology-based pre-processing
step leads to slight improvements over the baseline with micro F1 of 0.59 in
comparison to baseline micro-F1 of 0.57. Despite the small amount of data, we
found that corpus trained embedding provide a notable advantage over pre-trained
embeddings in the classi ers performance.
fastText classi er outperformed GNB model, especially when domain-based
embeddings were used. A non-verbatim example of a sentence where fastText
model, based on corpus-trained embeddings performs better than pre-trained
embedding models is: 'The police received information that the subject was selling
crack'. A potential reason for fastText to classify correctly this sentence versus
the classi ers using pre-trained embeddings is that the word `crack' has the
meaning of a `drug' in the reports. However, this is not the widely accepted
meaning for this word and thus it cannot be interpreted correctly by pre-trained
models. The GNB based on pre-trained BERT model outperforms the classi ers
based on pre-trained embeddings, however it does not lead to improvements over
the domain-based models. Fine-tuning BERT is the best performing classi er
with micro-F1 of 0.64 and macro-F1 of 0.59 which gives 0.5 improvement over
the baseline. The improvement in the results achieved by ne-tuning BERT
indicate the importance of adapting even the more context-aware pre-trained
language models to the speci c domain, especially when the domain contains
highly specialised language. Further, the poor performance of classi ers based
on pre-trained word models shows the lack of transferability of pre-trained
embeddings for a highly specialised domain such as the safeguarding reports.</p>
        <p>The three best-performing classi ers give similar average results between the
dev and test set (see Table 3). Further, models tend to return higher results
for some themes, especially `Mental Health Issues' for the test set rather than
the dev set. A potential reason for this may be attributed to the fact that the
test set has been annotated in a similar manner to the classi cation models,
i.e., independent of the context of the entire documents. The BERT classi er
returned results above 0.50 for the themes `Contact with Agencies', `Re ections'
and `Indicative Behaviour' for the dev and test datasets with precision above
0.60 and recall above 0.70.
In the preceding section we evaluated the performance of the classi cation
approaches against the annotations generated by the creators of the thematic
framework, who we refer to as the expert annotators. By creating a classi er
that uses the annotations generated by expert annotators as a `ground truth', we
aim to produce uni ed and comparable results across generations that are not
susceptible to variations in annotations created by di erent human annotators
interpreting the coding framework. Going further, we judge the predictive power
of the models by comparing their performance against the annotations of expert
validators : independent social scientists who did not participate in the creation of
the thematic annotation framework. We aim to measure the ability of the learned
models to conserve the knowledge of the expert annotators versus if the task was
performed manually by independent social scientists who were not creators of the
framework (Section 5.1). In this way, we will be able to judge whether automated
approaches are reliable for labeling the reports.</p>
        <p>We perform three main types of analysis. First, we compare the performance of
the classi ers against the annotations of expert validators. Secondly, we compare
the performance of the classi ers for di erent length of sentences to observe the
classi ers suitability for various sequence lengths. We also measure the e ect of
the training dataset size on the performance of the models (Section 5.2). Thirdly
and nally, we look at the e ect of the number of training instances versus the
amount of context provided per instance on the performance of the classi ers
(Section 5.3).
5.1</p>
      </sec>
      <sec id="sec-4-3">
        <title>Expert Validators vs Classi ers</title>
        <p>The initial thematic framework was developed by annotating passages of the
documents rather than individual sentences. However, our classi ers are trained
with sentences. In order to fairly judge the predictive power of the models against
human annotators for annotating sentences and passages of the reports, we
performed a study comparing the performance of the classi cation models versus
two independent expert validators on sentence- and passage-level. For these
purposes we used two datasets | one consisting of sentences and one consisting
of passages. The sentence set consisted of a sample of 100 randomly chosen
sentences, while the passage set consisted of a 100 passages, each containing three
sentences. The sentence set was extracted from the dev set while the passage set
was extracted from the test set (see Table 1). We measured the inter-annotator
agreement for predicting themes using Cohen's kappa (see Table 4). We also
compare the average F1 measure per theme between the expert validators and
the best performing classi er (BERT).</p>
        <p>The Cohen's kappa scores showed moderate agreement between the validators
with an average score 0.40 on sentence and a passage level. The highest level of
agreement is for `Mental Health Issues' theme. However, the average expert F1
for this theme is surprisingly low. The reason for the discrepancy between the
Cohen's kappa score and the F1 measure is the occurrence of sentences which
mention mental health problems such as `depression'. Such sentences are labeled
by the expert validators as `Mental Health Issues', however their true label is
di erent because of the surrounding context. Surprisingly, a large portion of these
sentences were correctly classi ed by BERT. The average F1 score for the expert
validators signi cantly improves for passage-level classi cation with average F1
= 0.60 in comparison to sentence-level annotations where an average F1 = 0.45
(see Table 4). This suggests that humans need more context | i.e., to see the
sentences embedded in paragraphs | to classify sentences correctly, compared
to deep learning models that can generalize better in these cases with limited
context thanks to what they learned from the training set.
5.2</p>
      </sec>
      <sec id="sec-4-4">
        <title>E ect of sentence length and training size</title>
        <p>Experiments comparing the best-performing classi ers for di erent sentence
length and training set size showed that BERT performed better than the
baseline method for any length of sentences. Further, BERT gave higher results
than fastText and the baseline for shorter sentences. For long sentences, BERT
and fastText had very similar performance with a di erence less than 1% (see
Fig.2). The comparison between the classi cation models performance for di erent
sizes of training set revealed that deep learning models (i.e., BERT) are highly
in uenced by the size of the training set in comparison to linear models such as
the baseline and fastText (see Fig.2). BERT performed worse than the baseline
for the very small training set while fastText gave similar performance to the
baseline. However, BERT's performance almost doubled as more sentences were
added to the training set while GNB performance was not that heavily in uenced
by the size of the training data, especially for a training set with more than 1,000
sentences.
5.3</p>
      </sec>
      <sec id="sec-4-5">
        <title>Sentences vs Passages</title>
        <p>In this section, we extend the analysis from Section 5.1 by looking at the e ect of
context versus the number of training instances provided for the classi er models.
In this experiment, we gradually increase the length of the training instances in
order to judge the importance of the training size versus the context (in terms
of passage length). We evaluate the models using sentences and passages where
each test passage consisted of three sentences (see Fig. 3). The test sets for
these experiments were extracted from the dev set while the training sentences
and passages were extracted from the training set. Results showed that the
performance of deep learning models is more in uenced by the amount of the
training instances rather than the length of the training passages. Further, models
trained on sentence-level with a higher volume of training data give better results
when tested on small paragraphs than classi ers trained on passages but with
less training data available. This signi es the importance of higher volume of
labelled data for reaching good classi ers performance.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>Through this work, we explored the problem of predicting the main themes in
safeguarding reports using supervised machine learning approaches. We analysed
the performance of state-of-the-art classi ers, feature extraction and feature
integration techniques which allowed us to identify classi cation methods suitable
for domain-speci c documents. Results showed that state-of-the-art deep learning
model performance is highly dependent on the size of the training data in
comparison to linear models as BERT's performance is worse than a simple Naive
Bayes baseline and fastText for very small training datasets. Further, training
word embeddings onto the speci c domain, even when the size of the corpus is
very small, lead to much higher results in comparison to pre-trained embeddings.
This shows the importance of targeting pre-trained models to the speci c corpus
despite its size. The study comparing the expert validators' performance versus
the automated models showed that the thematic analysis can be challenging even
for subject-matter experts without prior knowledge of the thematic annotation
framework. Further, humans need more knowledge about the context surrounding
a sentence, compared to deep learning approaches. Experiments showed that
BERT and fastText performance is more a ected by the size of the training
data rather than the amount of context given. On this respect, sentence-level
classi cation provides more training data and ne-grained distinction between
themes, which in turn allows for an easier expansion of the models and faster
annotation.</p>
      <p>In the future, we want to improve theme detection for the safeguarding
documents by using generative language models for arti cially augmenting the
sparse data of the corpus. We will use the additional data as a training set in
order to improve classi er performance. Further, we plan to look into developing
and using knowledge graphs for improving classi cation. This will help re ne
the query functionality of the application and help improve the identi cation of
similar documents and common trends in the safeguarding collection.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ali</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Text classi cation based on fuzzy radial basis function</article-title>
          .
          <source>Iraqi Journal for Computers and Informatics</source>
          <volume>45</volume>
          (
          <issue>1</issue>
          ),
          <volume>11</volume>
          {
          <fpage>14</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>5</volume>
          ,
          <issue>135</issue>
          {
          <fpage>146</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . ArXiv abs/
          <year>1810</year>
          .04805,
          <issue>16</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camacho-Collados</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Ribaupierre</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Preece</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Go simple and pre-train on domain-speci c corpora: On the role of training data for text classication</article-title>
          .
          <source>In: Proceedings of the 28th International Conference on Computational Linguistics</source>
          . pp.
          <volume>5522</volume>
          {
          <issue>5529</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Preece</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Ribaupierre</surname>
          </string-name>
          , H.:
          <article-title>Knowledge extraction from a small corpus of unstructured safeguarding reports</article-title>
          .
          <source>In: European Semantic Web Conference</source>
          . pp.
          <volume>38</volume>
          {
          <fpage>42</fpage>
          . Springer, Portoroz, Slovenia (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>K.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsieh</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>Liblinear: A library for large linear classi cation</article-title>
          .
          <source>Journal of machine learning research 9(Aug)</source>
          ,
          <year>1871</year>
          {
          <year>1874</year>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gururangan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marasovic</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swayamdipta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beltagy</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Downey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          :
          <article-title>Don't stop pretraining: Adapt language models to domains and tasks</article-title>
          . arXiv preprint arXiv:
          <year>2004</year>
          .
          <volume>10964</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Bag of tricks for e cient text classi cation</article-title>
          .
          <source>In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>2</volume>
          ,
          <string-name>
            <given-names>Short</given-names>
            <surname>Papers</surname>
          </string-name>
          . pp.
          <volume>427</volume>
          {
          <fpage>431</fpage>
          . Association for Computational Linguistics, Valencia,
          <source>Spain (April</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          ,
          <fpage>3111</fpage>
          {
          <volume>3119</volume>
          (10
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , et al.:
          <article-title>Scikit-learn: Machine learning in python</article-title>
          .
          <source>Journal of machine learning research 12(Oct)</source>
          ,
          <volume>2825</volume>
          {
          <fpage>2830</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amodei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Language models are unsupervised multitask learners</article-title>
          .
          <source>OpenAI Blog</source>
          <volume>1</volume>
          (
          <issue>8</issue>
          ),
          <volume>9</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rees</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dehaghani</surname>
          </string-name>
          , R.:
          <article-title>Making connections: a multi-disciplinary analysis of domestic homicide, mental health homicide and adult practice reviews</article-title>
          .
          <source>The Journal of Adult Protection</source>
          <volume>21</volume>
          (
          <issue>1</issue>
          ),
          <volume>16</volume>
          {
          <fpage>26</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sainz</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rigau</surname>
          </string-name>
          , G.:
          <article-title>Ask2transformers: Zero-shot domain labelling with pre-trained language models</article-title>
          .
          <source>arXiv preprint arXiv:2101.02661</source>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Spasic</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenwood</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Preece</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elwyn</surname>
          </string-name>
          , G.:
          <article-title>Flexiterm: a exible term recognition method</article-title>
          .
          <source>Journal of biomedical semantics 4</source>
          (
          <issue>1</issue>
          ),
          <volume>27</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Turker, R., Zhang,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Koutraki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Sack</surname>
          </string-name>
          , H.:
          <article-title>Knowledge-based short text categorization using entity and category embedding</article-title>
          .
          <source>In: European Semantic Web Conference</source>
          . pp.
          <volume>346</volume>
          {
          <fpage>362</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michael</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          :
          <article-title>Glue: A multitask benchmark and analysis platform for natural language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>07461</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Wolf</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Debut</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanh</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaumond</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delangue</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cistac</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rault</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Louf</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Funtowicz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brew</surname>
          </string-name>
          , J.:
          <article-title>Huggingface's transformers: State-ofthe-art natural language processing</article-title>
          . ArXiv abs/
          <year>1910</year>
          .03771 (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>