<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Winterthur, Switzerland, April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Automated Requirements Demarcation using Large Language Models: An Empirical Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kaishuo Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Feier Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehrdad Sabetzadeh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Electrical Engineering and Computer Science (EECS), University of Ottawa</institution>
          ,
          <addr-line>800 King Edward Ave, Ottawa, ON, K1N 6N5</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>8</volume>
      <issue>2024</issue>
      <abstract>
        <p>Requirements demarcation is concerned with the identification of software and systems requirements in technical natural-language specifications. Manually distinguishing requirements from non-requirements is a labour-intensive task and is prone to omissions and errors. Therefore, building a reliable and accurate automated method for requirements demarcation is important. This research evaluates auto-encoder large language models [1], auto-regressive large language models [2], few-shot learning, and ensembling methods for accurate demarcation of requirements. Specific architectures considered are DeBERTa, Llama2, and few-shot learning with RoBERTa. Our work empirically compares these approaches and determines which one yields the best accuracy. Our experimental results show that DeBERTa yields the best performance. While Llama2 requires more computational resources and training time compared to DeBERTa, it has an accuracy deficit of approximately 1% across diferent metrics. This result suggests that auto-encoder models may be more suitable for requirements demarcation than auto-regressive models. Further, we observe that the few-shot learning approach has the worst performance among the alternatives considered. Finally, we find that ensembling leads to minor performance improvements compared to a single model. We make all the artifacts developed as part of this research available online: https:// github.com/ KaishuoWang/ Automated-Requirements-Classification .</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;requirements demarcation</kwd>
        <kwd>natural language processing</kwd>
        <kwd>ensemble learning</kwd>
        <kwd>transformers</kwd>
        <kwd>empirical evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Distinguishing requirements from non-requirements in technical documents is a dificult and
labour-intensive task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Manually classifying text across thousands of pages to identify
requirements is also prone to omissions and errors. Automating requirements demarcation,
defined as the task of distinguishing requirements from non-requirement sentences [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], is thus
important for streamlining downstream analysis, traceability, risk identification, and testing [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Table 1 shows examples to illustrate sentences classified as requirements and non-requirements.
      </p>
      <p>With the recent emergence of large language models, an interesting question arises as to
whether one can improve upon the accuracy of existing natural language processing techniques</p>
      <p>NPAC SMS shall provide post-collection audit analysis tools that can produce
detailed reports on data items relating to system intrusions.</p>
      <p>NANC Version 1.6, released on 11/12/97, contains changes from the NANC
FRS Version 1.5.</p>
      <p>User selects option to associate file types with editors (ALT 1).</p>
      <p>Label
Requirement</p>
      <p>
        Requirement
Non-requirement
Non-requirement
for requirements demarcation. This paper reports on an empirical examination comparing
several alternative technologies. We evaluate two state-of-the-art models, DeBERTa and Llama2,
as well as few-shot learning with RoBERTa. We also examine an ensemble approach combining
these methods to study whether demarcation accuracy over complex documents can be further
improved. Based on our experimental results, we observe that DeBERTa slightly outperforms
our replication of Bashir et al.’s approach [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The accuracy results reported by Bashir et al.
are nonetheless slightly higher than our replication, most likely due to slight diferences in
hyperparameter optimization. Consequently, our accuracy results from DeBERTa are slightly
below those reported by Bashir et al. In view of this finding, we hypothesize that DeBERTa
will likely outperform Bashir et al.’s approach if its hyperparameters are optimized according
to the process used by Bashir et al., details of which we did not have and could not exactly
replicate. As for Llama2, we observe that the model has lower performance than DeBERTa,
which could be attributed to the small size of the training data, the lack of context, and the
absence of a training data selection strategy. We further observe that few-shot learning has a
notable performance deficit compared to both DeBERTa and Llama2. In addition, we find that
simply integrating the predictions of each approach using their normalized F1 score does not
lead to significant improvement in performance. More complex ensemble approaches should
thus be considered in the future.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Work</title>
      <p>Automating the demarcation of requirements has seen considerable interest in the field of
software engineering. This section presents a brief overview of requirement demarcation
techniques alongside their enabling techniques, highlighting the evolution from traditional
machine learning approaches to the latest advancements in deep learning and transformer-based
models.</p>
      <p>
        Early requirements demarcation methods, e.g., work by Abualhaija et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], use traditional
machine learning such as Support Vector Machines (SVM), Logistic Regression (LR), and Naive
Bayes (NB). These methods require careful feature engineering and show limitations in handling
the linguistic nuances and complexities inherent in requirements text.
      </p>
      <p>
        Long Short-Term Memory (LSTM) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] networks revolutionized deep learning for sequential
data. LSTMs have gating mechanisms that allow them to retain or forget information, reducing
issues like vanishing gradients. This enables models to process longer sequences while capturing
context. New word embeddings like FastText (FT) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and Global Vectors (GloVe) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] have been
used as well, with FT aggregating subword n-grams for better morphology understanding
and GloVe using co-occurrence statistics to capture semantic relationships. Overall, LSTMs
allow sequential modelling of longer texts, while new word embeddings like FT and GloVe
enable more semantic understanding. In addition, Winkler et al.[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] proposed a method that
utilizes convolutional neural networks for automated requirements classification. Their method
achieved 0.73 in accuracy and 0.89 in recall on a real-world automotive requirements specification
      </p>
      <p>
        The emergence of Transformer-based models marks an important milestone in NLP.
Transformers like BERT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] have brought about major advances in NLP through self-attention and
bidirectional context modelling. Rather than processing text linearly, self-attention weighs the
significance of each word relative to all others, learning context more efectively. BERT
specifically introduces bidirectionality, understanding context based on surrounding text on both
sides of a word. This enables capturing nuances and relationships that were previously dificult
to capture. BERT uses WordPiece tokenization [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] to break words into subword units, allowing
it to handle unseen words by representing them as known subwords. Therefore, BERT and its
variants generally perform well at understanding context and complex dependencies, which are
key attributes for accurately demarcating the content of complex requirements documents.
      </p>
      <p>
        Bashir et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] propose few-shot learning with sentence transformers [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], utilizing the
SetFit framework [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] to fine-tune various pre-trained Sentence Transformer models for the
task of demarcating requirements. This process involves a dual-step training approach: initially,
ifne-tuning the Sentence Transformer model on a limited dataset using a contrastive training
method, followed by training a Logistic Regression model to act as a classification head on the
embeddings generated by the fine-tuned Sentence Transformer. The evaluation of this model
entails generating sentence embeddings from unseen examples and then predicting the class
label with the Logistic Regression model.
      </p>
      <p>
        In Bashir et al.’s work, the bert-base-uncased pipeline obtained the best performance, with a
Macro average F1 score of 0.83 on the Dronology dataset [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This pipeline will be used as our
baseline and compared to our approaches.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Alternative Approaches</title>
      <p>Figure 1 provides an overview of the alternative models we examine in this paper. Recognizing
the potential for synergy, we consider two ensemble systems among the alternatives. The first
ensemble system integrates the strengths of DeBERTa and Llama2, while the second system
leverages the strength of all three models. This approach not only allows us to compare three
methods separately but also to investigate potential enhancements brought about by ensembling.</p>
      <sec id="sec-3-1">
        <title>3.1. DeBERTa</title>
        <p>
          DeBERTa [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] improves upon the BERT architecture in several key ways. DeBERTa’s defining
feature is its disentangled attention mechanism. Unlike BERT, which processes content and
position information jointly within a single self-attention mechanism, DeBERTa separates the
two. This allows the model to learn the representation of word positions and the content more
efectively, giving it a more refined understanding of the meaning that word order and position
contribute to a sentence. Such a capability is especially critical in requirement demarcation
tasks, where the position of terms can dramatically change their meaning, and then change
the classification. In addition, DeBERTa introduces an Enhanced Mask Decoder (EMD). BERT
adds absolute position encodings directly in the input layer. In contrast, DeBERTa captures
only relative position information within the Transformer layers, and applies absolute position
information later, right before predicting masked tokens. Thus, EMD is better at predicting the
original tokens from masked ones by using contextual information more eficiently. Given these
architectural improvements and DeBERTa’s demonstrated superior performance over BERT in
various natural language processing benchmarks, including text classification, it is interesting
to examine the application of DeBERTa as a potentially more efective alternative to BERT.
3.2. Llama2
Llama2 [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is a transformer-based auto-regressive large language model. This type of model
is often applied to text generation tasks, such as machine translation, generative
questionanswering systems, and virtual assistants. In this project, we aim to examine the performance
of this model on a text classification task (i.e., requirements demarcation) and compare it with
BERT and other machine learning-based classification methods. Llama2 has a substantially
larger number of parameters than BERT variants. The version we used in this project contains 7
billion parameters, compared to 345 million and 110 million parameters in DeBERTa-large and
BERT-base, respectively. With more parameters, Llama2 is better at generalizing knowledge
learned from large amounts of training data to new tasks. At the same time, more parameters
may improve context understanding and robustness to ambiguity. These improvements will
give models a deeper understanding of context and remove ambiguities common in natural
language. In addition, Llama2 combines multi-task learning to jointly train models on diferent
natural language understanding tasks. In other words, the Llama2 model is not only trained on a
large text corpus but is also fine-tuned for multiple natural language understanding tasks. This
enables the model to learn from diverse and rich data sources and improve its generalization
capabilities. These improvements by Llama2 present an opportunity for a potentially more
accurate and streamlined approach to classification tasks.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Few-shot Learning</title>
        <p>As a machine learning technology that has emerged in recent years, few-shot learning enables
the model to learn from a small number of samples. Compared to supervised machine learning,
few-shot learning only requires a small number of data points to learn the information in
the data and generalize it to new tasks, making it useful when the amount of labelled data is
small. Since engineering specification documents are usually confidential information within
a company or organization, we hypothesize that applying few-shot learning to the task of
requirements demarcation can solve the problem of small data volumes.</p>
        <p>
          One recent development in the field of few-shot learning is SetFit [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] – an eficient and
prompt-free framework for the few-shot fine-tuning of sentence-transform models. Compared
to other few-shot learning methods, SetFit requires no prompts at all and can generate rich
embeddings directly from text examples. In addition, SetFit does not require large models like
GPT or Llama to achieve high accuracy. This means that it requires less computing resources
and training time.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Ensemble System</title>
        <p>Ensemble learning is a machine learning technique that combines the predictions from two or
more models. Compared to using single models, the ensemble technique exploits the predictive
power of single models to achieve more accurate predictions and better performance. Moreover,
the ensemble technique also improves the robustness of the model by reducing the spread or
dispersion of the predictions and model performance.</p>
        <p>To ensemble the predictions of DeBERTa, Llama2, and Few-shot Learning, we use the
normalized macro F1 score of each model as weight and multiply with their predictions to get the
ifnal prediction  :
 = 1 × 1 + 2 × 2 + 3 × 3
(1)
where 1, 2, 3 are the weights of each model and 1, 2, 3 are the predictions of each
model.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Implementation</title>
      <p>
        For this project, we utilize the PyTorch and HuggingFace transformers library. During the
ifne-tuning process of DeBERTa, we add a classification header on top of DeBERTa. This is a
common way of fine-tuning, where the weights of a neural network in the classification head
are updated via back propagation. However, this method requires a lot of computing resources
and time, so it is not suitable for fine-tuning the more complex Llama2 model. Therefore,
we use Low Rank Adaptation (LoRA) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], an improved fine-tuning method that has been
proven efective, in the process of fine-tuning Llama2. LoRA avoids catastrophic forgetting, a
phenomenon that occurs when knowledge of a pre-trained model is lost during fine-tuning,
by fine-tuning two smaller matrices that approximate the weight matrix of a large pre-trained
language model. This method greatly reduces the number of trainable parameters, enabling us
to keep the computational resources and time required for fine-tuning within an acceptable
range. For few-shot learning, we use the SetFit [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] framework and RoBERTa-large model. We
ifne-tuned the few-shot learning model with diferent numbers of labelled examples for each
label, including 8, 16, 24, and 32. The model fine-tuned with 24 labelled datapoints (per class)
yielded the best performance. Therefore, we use 24 randomly selected samples from each label
(requirement or non-requirement) for fine-tuning the few-shot learning model.
      </p>
      <p>We fine-tuned the DeBERTa and few-shot learning models using a single Nvidia V100 16GB
GPU, and the Llama2 model was fine-tuned using a single Nvidia A100 40GB GPU on Colab.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Empirical Evaluation</title>
      <sec id="sec-5-1">
        <title>5.1. Research Questions</title>
        <p>
          The goals of our evaluation are three-fold. First, we aim to investigate whether there is suficient
rationale for replacing auto-encoder models such as BERT and DeBERTa with more complex
auto-regressive models like Llama2 for the task of requirements demarcation. Settling this
question necessitates an examination of the trade-of between the potentially increased accuracy
from the more complex models and the less resource-intensive nature of earlier auto-encoder
models. In this comparison, the BERT model was used as baseline to compare with DeBERTa
and Llama2, since the BERT model yielded the best performance in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Second, we seek to
study the performance of few-shot learning for requirements demarcation to conduct a more
thorough comparison with DeBERTa and Llama2. And third, in order to enhance the reliability
of the model as a replacement for manual work, we aim to analyze whether the accuracy of
requirements demarcation can be further improved by utilizing ensemble learning techniques.
To achieve these objectives, we present the following research questions (RQs):
        </p>
        <p>RQ1: Which model architecture - DeBERTa, Llama or BERT - yields the most accurate
requirements demarcation results?</p>
        <p>RQ2: Can few-shot learning achieve on-par performance with state-of-the-art approaches?
RQ3: Can ensembling improve upon individual models?</p>
        <p>
          To answer the research questions, we designed two sets of experiments. In the first set of
experiments, we utilize stratified five-fold cross-validation on the Dronology dataset (discussed
in Section 5.2) with DeBERTa, Llama2, and a RoBERTa-based few-shot learning model to provide
a fair and direct comparison with the work of Bashir et al [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. In the second set of experiments,
we fine-tune DeBERTa, Llama2, and the same RoBERTa-based few-shot learning model with the
dataset combining the Dronology and PURE datasets (discussed in Section 5.2) and compared
their performance.
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Description of Dataset</title>
        <p>
          The dataset we use for our evaluation combines two existing labelled datasets: the Dronology
dataset from [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and a manually extracted and labelled dataset from the PURE dataset by
Ivanov et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>
          The Dronology dataset consists of 398 entries of various classes, including components, design
definitions , sub-task, and requirements. This dataset was processed and labelled by Bashir et
al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. They first labelled all non-requirement classes as non-requirement, then deleted 19 entries
without text. After processing, the dataset contains 99 requirements and 280 non-requirements.
To mitigate imbalance, the authors use stratified 5-fold cross-validation.
        </p>
        <p>
          The second dataset was labelled by Ivanov et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], from the PURE dataset developed
by Ferrari et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. They manually extracted 7,745 sentences from 30 of the 79 natural
language requirements documents where 4,145 sentences are requirements and 3,600 are
nonrequirements. To improve labelling accuracy, they employed a manual labelling process by a
subject-matter expert and independently verified by additional experts, and any disputed data
was removed. We consider this dataset as a measure against imbalanced and small amounts of
data.
        </p>
        <p>
          For the first set of experiments, discussed in Section 5.1, we use the Dronology dataset in
the replication package provided by Bashir et al [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] where the dataset was partitioned into
ifve subsets using 5-fold cross-validation. For the second set of experiments, we merged the
subsets to form a single dataset and then combined this merged set with the PURE dataset
to enlarge the dataset size and address label imbalance. By combining these two datasets, we
obtain 8,124 sentences, which contain 4,244 requirements and 3,880 non-requirements with
an average length of 134 words. Of these, we randomly selected 30% as the testing set (2,438
sentences), while the remaining 70% was used for fine-tuning (5,686 sentences).
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Results and Discussion</title>
        <p>
          Tables 2 and 3 show our experimental results for all the approaches. The Weighted Avg. and
Macro Avg. columns are the weighted average and macro average of each evaluation metric.
The Training Time column indicates the total training for each approach. We do not report
training times for the two ensemble systems because no training is necessary for these systems.
The results for BERT are obtained from our replication of bert-base-uncased on the Dronology
dataset using stratified five-fold cross-validation and the same configuration used by Bashir et
al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], to the extent that we could recreate the configuration based on the paper.
        </p>
        <p>RQ1: As shown in Table 2 where we utilized the stratified five-fold cross-validation to evaluate
all the models, DeBERTa achieved the best performance with a 82.87% macro F1 score. This
raises the prospect that DeBERTa will outperform BERT if one could achieve the same level of
performance as Bashir et al. have observed with BERT.</p>
        <p>In relation to Llama2 and few-shot learning, we note that these models show lower
performance than BERT. This deficit can be attributed to several factors. First, the small size of the
training set, with only 304 instances per fold, might be insuficient for these models to efectively
learn the task, particularly for Llama2, which typically performs better with larger datasets.
Second, the lack of context in the data can be detrimental for requirements classification tasks,
as context plays an important role in determining whether a sentence is a requirement or not.
Without suficient context, models may struggle to distinguish between neutral sentences and
actual requirements, leading to performance degradation. Third, the training data selection
strategy used in the few-shot learning approach could be suboptimal if the selected instances
are not representative of the entire dataset or do not cover the diversity of classes and patterns
present in the data. In future work, to mitigate these issues, one could consider solutions
that increase the size of the training set, incorporate context information, and employ more
sophisticated training data selection strategies for few-shot learning, such as clustering-based
techniques, manual curation, or active learning approaches.</p>
        <p>Regarding training time, compared to BERT, which took 11 minutes to train, DeBERTa and
Llama2 required significantly more training time, at 1 hour and 7 minutes and 1 hour and 8
minutes respectively. Therefore, while DeBRETa exhibits better performance, the substantially
longer training time suggests that careful consideration needs to be given to the accuracy versus
time trade-of.</p>
        <p>RQ2: The few-shot learning method exhibited lower performance in both sets of experiments,
suggesting that it is unable to match the capabilities of the other models we examined for the
requirements demarcation task. However, considering that it requires only 0.8% of the training
data (48 used by few-shot learning compared to 5686 used by DeBERTa and Llama2) and has a
relatively shorter training time compared to other models, we believe it is a viable option to
explore when one has to cope with very small training data.</p>
        <p>RQ3: We first attempted to construct an ensemble system by combining the predictions of
DeBERTa and Llama2, each weighted at 0.5. As shown in Table 3, ensembling DeBERTa and
Llama2 models improves upon individual models in terms of overall accuracy and weighted
scores, achieving higher accuracy (94.79%) and weighted scores compared to the individual
models. The ensemble system’s high weighted F1 of 95.53% indicates its ability to correctly
classify the majority class instances. However, it has lower macro scores (precision, recall, and
f1) compared to the individual DeBERTa and Llama2 models, indicating the weaknesses of
DeBERTa and Llama2 in handling minority classes may be amplified in the ensemble. Finally,
we note that adding the few-shot learning model to the ensemble did not provide significant
improvements, potentially due to its lower individual performance. The second ensemble system,
which includes DeBERTa, Llama2, and few-shot learning, shows a slightly lower accuracy of
92.86% compared to the first ensemble system, with lower weighted and macro scores.</p>
        <p>Overall Conclusion. In conclusion, our experiments demonstrated that DeBERTa outperformed
other models, albeit requiring a longer training duration compared to BERT. As the volume
of training data increased, Llama2 and few-shot learning exhibited comparable performance
with DeBERTa. Moreover, the results revealed that while ensemble systems generally enhanced
overall performance, they faced challenges in accurately classifying minority classes. Moving
forward, it is crucial to carefully evaluate class imbalance and ensemble architectures to address
this particular issue.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Threats to Validity</title>
        <p>Internal Validity. Due to resource limitations, our hyperparameter tuning was not exhaustive,
meaning there could still be undiscovered model configurations that improve performance.
Furthermore, despite the labelled data having been preprocessed and deemed to be of high
quality, it is possible that there still remain unusual, anomalous, or improperly labelled examples
within the combined requirements dataset. Such data outliers and noise could skew model
performance. In future work, robustly detecting and removing possible labelling errors through
outlier analysis and visual data profiling could increase data integrity.</p>
        <p>External Validity. While our evaluation yielded rather conclusive results on our dataset,
the question of whether these findings would generalize to diferent datasets and varied criteria
for requirements demarcation necessitates further case studies.</p>
        <p>Construct Validity. Currently, our models only classify text into binary requirement and
nonrequirement categories. However, real-world specifications contain a diverse array of semantic
types, such as constraints, assumptions, and metadata descriptors. In complex documents, many
such more nuanced labels can exist. By training and evaluating solely on a binary classification
task, our models may not capture the full richness within specifications. Expanding the label set
beyond binary categories could better measure model capabilities and efectiveness for realistic
applications.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This paper benchmarked two large pre-trained language models, DeBERTa and Llama2, as
well as a RoBERTa-based few-shot learning model, for automatically demarcating software and
systems requirements in technical specifications. In future work, we plan to create broader
training data covering more domains to address the limitations in domain transfer. Furthermore,
the integration method used in the ensemble system could be further improved to better take
advantage of the complementary traits of DeBERTa and Llama2.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Stone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Albert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almahairi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Babaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bashlykov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhargava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhosale</surname>
          </string-name>
          , et al.,
          <source>Llama</source>
          <volume>2</volume>
          :
          <article-title>Open foundation and fine-tuned chat models</article-title>
          ,
          <source>arXiv preprint arXiv:2307.09288</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bashir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abbas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saadatmand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Enoiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bohlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lindberg</surname>
          </string-name>
          ,
          <article-title>Requirement or not, that is the question: A case from the railway industry</article-title>
          , in: International Working Conference on Requirements Engineering: Foundation for Software Quality, Springer,
          <year>2023</year>
          , pp.
          <fpage>105</fpage>
          -
          <lpage>121</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Abualhaija</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sabetzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Briand</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Vaz,</surname>
          </string-name>
          <article-title>A machine learning-based approach for demarcating requirements in textual specifications</article-title>
          ,
          <source>in: 2019 IEEE 27th International Requirements Engineering Conference (RE)</source>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Staudemeyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. R.</given-names>
            <surname>Morris</surname>
          </string-name>
          ,
          <article-title>Understanding lstm-a tutorial into long short-term memory recurrent neural networks</article-title>
          , arXiv preprint arXiv:
          <year>1909</year>
          .
          <volume>09586</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Grave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , T. Mikolov,
          <article-title>Enriching word vectors with subword information, Transactions of the association for computational linguistics 5 (</article-title>
          <year>2017</year>
          )
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          , Glove:
          <article-title>Global vectors for word representation</article-title>
          ,
          <source>in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Winkler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vogelsang</surname>
          </string-name>
          ,
          <article-title>Automatic classification of requirements based on convolutional neural networks</article-title>
          ,
          <source>in: 2016 IEEE 24th International Requirements Engineering Conference Workshops (REW)</source>
          , IEEE,
          <year>2016</year>
          , pp.
          <fpage>39</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schuster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Norouzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Macherey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krikun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Macherey</surname>
          </string-name>
          , et al.,
          <article-title>Google's neural machine translation system: Bridging the gap between human and machine translation</article-title>
          ,
          <source>arXiv preprint arXiv:1609.08144</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          , arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>10084</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Tunstall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U. E. S.</given-names>
            <surname>Jo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Korat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wasserblat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pereg</surname>
          </string-name>
          ,
          <article-title>Eficient few-shot learning without prompts</article-title>
          ,
          <source>arXiv preprint arXiv:2209.11055</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cleland-Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vierhauser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bayley</surname>
          </string-name>
          ,
          <string-name>
            <surname>Dronology:</surname>
          </string-name>
          <article-title>An incubator for cyber-physical system research</article-title>
          , arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>02423</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          , W. Chen, Deberta:
          <article-title>Decoding-enhanced bert with disentangled attention</article-title>
          , arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>03654</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wallis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Allen-Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Lora:
          <article-title>Low-rank adaptation of large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2106.09685</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>V.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sadovykh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Naumchev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bagnato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yakovlev</surname>
          </string-name>
          ,
          <article-title>Extracting software requirements from unstructured documents</article-title>
          ,
          <source>in: International Conference on Analysis of Images, Social Networks and Texts</source>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>17</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. O.</given-names>
            <surname>Spagnolo</surname>
          </string-name>
          , S. Gnesi,
          <article-title>Pure: A dataset of public requirements documents</article-title>
          ,
          <source>in: 2017 IEEE 25th International Requirements Engineering Conference (RE)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>502</fpage>
          -
          <lpage>505</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>