<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic IMRAD Classification of Citation Contexts: Comparing Text Representations and Machine Learning Classifiers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francis Lareau</string-name>
          <email>lareau.francis@courrier.uqam.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>SCOLIA '26: Second International Workshop on Scholarly Information Access</institution>
          ,
          <addr-line>SCOLIA</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Université de Sherbrooke</institution>
          ,
          <addr-line>2500, boulevard de l'Université Sherbrooke (Québec), J1K 2R1</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Université du Québec à Montréal</institution>
          ,
          <addr-line>1430, rue Saint-Denis, Montréal (Québec), H2X 3J8</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <fpage>31</fpage>
      <lpage>38</lpage>
      <abstract>
        <p>Identifying the rhetorical function of citations within scientific articles is essential to advance citation analysis and develop intelligent writing assistance tools. Citation contexts serve distinct argumentative roles depending on their location within the standardized IMRAD structure (Introduction, Methods, Results, Discussion). To our knowledge, no comprehensive evaluation of machine learning approaches for automated IMRAD citation classification exists. This study systematically evaluated 172 combinations of representation-classifiers in a corpus of 303,947 annotated citation contexts from the scientific literature to predict the placement of sections. We compared nine text representation methods-ranging from classical bag-of-words to fine-tuned BERT variants and OpenAI embeddings-with five classifier categories (linear, neural, tree-based, instance-based, and unsupervised). The results show that domain-specific fine-tuning substantially enhances performance, with BERT fine-tuned on IMRAD classification achieving a weighted F-measure of 0.753 when paired with linear classification. Linear classifiers generally outperformed neural and other approaches, suggesting that IMRAD-specific linguistic patterns create linearly separable feature spaces. Embeddings derived from an earlier OpenAI model achieved competitive performance (F1=0.722), indicating that LLM-derived representations capture relevant structural information. Interestingly, embeddings from a newer OpenAI model yielded worse performance (F1=0.709), indicating that newer models do not uniformly improve structural representation quality. Unsupervised clustering methods performed poorly (F1 = 0.685), confirming that IMRAD section inference requires labelled training data. These ifndings establish the feasibility of automatic citation context classification and highlight the importance of task-specific model adaptation for scientific text analysis. The work has implications for bibliometric studies, intelligent authoring systems, and research evaluation frameworks seeking to quantify the rhetorical impact of citations across article sections.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;argument mining</kwd>
        <kwd>citation context</kwd>
        <kwd>IMRAD</kwd>
        <kwd>OpenAI</kwd>
        <kwd>embeddings</kwd>
        <kwd>text representation</kwd>
        <kwd>BERT</kwd>
        <kwd>machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The IMRAD structure—Introduction, Methods, Results, and Discussion—has become the standardized
format for scientific communication in most disciplines. This standardization facilitates knowledge
dissemination and allows readers to locate specific information within research articles rapidly. Within
this structure, citations serve a critical rhetorical function: they ground arguments in existing knowledge,
establish scholarly credibility, and signal the scientific contribution being advanced.</p>
      <p>Citation contexts—the passages in a text containing one or more references [1]—are not uniformly
distributed across IMRAD sections; rather, they serve distinct argumentative functions depending
on their location. Prior work has shown that citation contexts exhibit systematic lexical variation
across IMRAD sections, notably in the distribution of verbs surrounding citations. Complementary
large-scale analyses have also demonstrated that the position and age of references follow an almost
invariant pattern along the IMRAD structure, with introductions and discussions concentrating most
citations. Our study builds on these insights by treating IMRAD section identification as a supervised
classification problem at the sentence level [ 2, 3]. The ability to automatically predict the IMRAD section
from which a citation context originates could have important applications in research evaluation,
document analysis, and the design of intelligent tools for scientific writing.</p>
      <p>Despite the structural importance of IMRAD organization in scientific writing, little is known about
the computational feasibility of identifying citation context placement within this structure. Previous
studies have examined citation contexts broadly, including whether cited works are peripheral or
fundamental to the citing texts’ arguments [4, 5, 6]. To our knowledge, no one has systematically
evaluated multiple machine-learning approaches for IMRAD classification. This study addresses this
gap by comparing text representation techniques and classification methods to predict afiliation of
citation contexts with IMRAD sections.</p>
      <p>This research is guided by the following primary research question: Can machine learning models
reliably predict the IMRAD section from which a citation context originates? We test three
specific hypotheses. First, we posit that citation contexts contain suficient linguistic and structural
markers to enable accurate IMRAD classification. Second, we hypothesize that advanced text
representation methods inspired by large language models, such as GPT-based embeddings, outperform
classical approaches including bag-of-words and doc2vec. Third, we expect neural network classifiers
to demonstrate superior performance compared to linear, tree-based, instance-based, and unsupervised
approaches. To test these hypotheses, we compared nine distinct text representation methods with five
classification methodologies (45 combined models total) on a corpus of 303,947 annotated citation
contexts from scientific articles. We employed the In-text Reference Corpus (InTeReC), a well-established
dataset with explicit IMRAD section labels, enabling robust evaluation across representation-classifier
combinations.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methods</title>
      <p>As illustrated in Figure 1, our experimental pipeline consists of corpus preprocessing, vector
representation of the text, training classification families of models, and subsequent evaluation.</p>
      <sec id="sec-2-1">
        <title>2.1. Data and Corpus Description</title>
        <p>The In-text Reference Corpus (InTeReC) served as the foundation for this study [7]. InTeReC
comprises 314,023 citation contexts extracted from 90,071 open-access scientific articles published in
seven Public Library of Science (PLOS) journals between 2001 and 2013. Each article was automatically
partitioned into IMRAD sections using XML tags, and sentences containing explicit citations were
extracted and annotated with their source section (Introduction, Methods, Results, Discussion). The
corpus exhibits class imbalance, with Discussion sections being most represented (37.5%), followed by
Introduction (23.5%), Methods (17.7%), and Results (9.0%). We removed 10,076 sentences labelled with
overlapping categories from the original corpus (Methods &amp; Results, Results &amp; Discussion), resulting in
the final annotated dataset of 303,947 sentences. The corpus was split into training (80%, n=243,157)
and evaluation sets (20%, n=60,790). All experiments were conducted with a fixed random seed so
that corpus splits and model initializations remained identical across configurations. We ran each
representation–classifier combination once, rather than repeating training with multiple seeds. This
design ensures a fair, controlled comparison between models while keeping the overall computational
cost low.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Representation Methods</title>
        <p>We used classical and modern vectorial text representations. The first, BOW, is a classic “bag of words"
representation in which words are simply counted. The second, DBOW, is a doc2vec representation
of the dbow variant. This model is generated using a simple neural network that learns to predict
a masked word from its context. The third, OAI1, is an early model developed by OpenAI formally
known as “text-embedding-ada-002". The fourth, OAI2, is the newer model of the OpenAI family named
“text-embedding-3-large". Finally, the last one is a series of BERT-type representations generated using
Transformers pre-trained on an extensive corpus of text. Bert is a bidirectional model that considers both
the context and the word order in the text. We used the original representations of the basic BERT model
- BERT Base – and the large BERT model – BERT Large. Then, three (3) fine-tuned representation
models are used. The first, BERT NLISTS, is based on BERT Large and has been fine-tuned on natural
language inference (NLI) and semantic textual similarity (STS) tasks [8]. The second, BERT MPNet,
is based on BERT Base and is fine-tuned on a masked and permutated language modelling network
(MPNet) task [9]. The last, BERT IMRAD, is based on the BERT-base model and is fine-tuned for
IMRAD classification using BERT for sequence classification. 1</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Classification Methods</title>
        <p>Various machine learning classification techniques enable us to perform the task at hand.</p>
        <p>
          Clustering models group similar data points into clusters based on their similarity. The principle
underlying these classifiers is to identify emergent groupings in the data automatically. This classification
does not require labelled training data and can therefore be used for unsupervised learning tasks.2 Note
that clustering-based classification can be used in cases where labelled training data is available, such
as to evaluate whether the classes or labels assigned to the data can match automatically detectable
clusters without a priori knowledge of the classes to be discovered. When this type of approach solves
a classification problem, it indicates that the feature space can be readily decomposed into implicit
clusters corresponding to the classes of interest. We tested Gaussian mixture and K-means [
          <xref ref-type="bibr" rid="ref1">12</xref>
          ].
        </p>
        <p>
          Instance-based classification models are a class of algorithms whose classification principle is the
similarity of new cases to instances observed during training. These models are generally nonparametric,
in that they require no assumptions about the distribution of the underlying data. These use all cases
of the training dataset when modelling, and new data points are ranked based on their similarity to
the training data. We tested the methods of the K-nearest neighbors [
          <xref ref-type="bibr" rid="ref2">13</xref>
          ], radius neighbors [
          <xref ref-type="bibr" rid="ref3">14</xref>
          ], and
nearest centroid [
          <xref ref-type="bibr" rid="ref4">15</xref>
          ].
        </p>
        <p>
          Linear classification models use linear decision boundaries to separate classes in feature space.
Additionally, kernel-based linear classification enables the classification of non-linearly separable
data in the original feature space by mapping the input data to a higher-dimensional space in which
linear separation is possible. We tested stochastic gradient descent [
          <xref ref-type="bibr" rid="ref5">16</xref>
          ], support vector machines [
          <xref ref-type="bibr" rid="ref6">17</xref>
          ]
(with and without kernel-based approaches), logistic regression [
          <xref ref-type="bibr" rid="ref7">18</xref>
          ], Ridge regression [
          <xref ref-type="bibr" rid="ref6 ref8">17, 19</xref>
          ], and
passive-aggressive classification [
          <xref ref-type="bibr" rid="ref9">20</xref>
          ]. 3
        </p>
        <p>
          Tree-based classification models use branched decision structures to classify data. The principle of
tree classifiers is to recursively divide the space of features into subsets that are more homogeneous
with respect to the target variable. We tested the decision trees [
          <xref ref-type="bibr" rid="ref10">21</xref>
          ], random forests [
          <xref ref-type="bibr" rid="ref11">22</xref>
          ], extremely
randomized trees [
          <xref ref-type="bibr" rid="ref12">23</xref>
          ], and gradient boosting [
          <xref ref-type="bibr" rid="ref13">24</xref>
          ].
        </p>
        <p>
          Neural classification models are a type of machine learning inspired by the structure of the human
brain. The principle of neural classification is to approximate a function that maps input data to
output labels by adjusting the weights of connections between neurons. This is achieved by iteratively
1We use Gensim [10] for BOW and DBOW; OpenAI’s API for OAI, and the Hugging Face API for BERT-type corpus
representation models [11].
2Clustering-based classification does not identify which category corresponds to which label. However, we approximate the
performance on the labelled data by assigning each label to the cluster that globally maximizes the overall F1 score.
3Although some algorithms contain ‘regression’ in their names, we use their standard classification variants that output
discrete IMRAD labels.
propagating the input data through the network and adjusting the weights based on the errors between
the predicted and actual labels. We tested the perceptron [
          <xref ref-type="bibr" rid="ref14">25</xref>
          ], the multilayer perceptron [
          <xref ref-type="bibr" rid="ref15">26</xref>
          ], and BERT
for sequence classification. 4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Evaluation Methodology</title>
        <p>Model performance was assessed using the weighted F-measure (F1). This metric was chosen to
address class imbalance and to provide a single interpretable performance score. The weighted F1  ()
aggregates the class-specific F 1-scores by weighting each class according to its relative frequency in the
dataset. For each class , the weight is defined as / ∑︀
=1  , where  is the number of instances in
class  and  is the total number of classes. The class-specific F 1 score is computed as the harmonic
mean of precision and recall. The resulting weighted F-measure is therefore given by:
 (︃
 () = ∑︁
=1</p>
        <p>∑︀=1 
)︃ (︂ 2 Precision Recall )︂
·</p>
        <p>Precision + Recall</p>
        <p>We computed the weighted F1 for all possible combinations of representation and classification
methods, yielding 172 models.5 For each family of models (a given representation method and classifier
type), we summarize performance by the maximum weighted F1 within that family. We further report
the overall best-performing configuration as the maximum of these family-wise maxima.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>Analysis of the 45 joint representation-classifier type combinations revealed substantial variation in
performance (Table 1). The maximum weighted F1 scores ranged from 0.292 (clustering with DBOW
representation) to 0.753 (linear classification with BERT IMRAD representation).</p>
      <p>Representation Model Performance: The highest-performing representation methods, based
on the maximum F1 score across families of classifiers, were BERT IMRAD (F1=0.753), followed by
OAI1 text-embedding-ada-002 (F1=0.722), BERT Base (F1=0.711), BERT Large (F1=0.7094), and OAI2
text-embedding-3-large (F1=0.709). Conversely, the poorest-performing representation methods,
identiifed using the minimum of the maximum F1 scores across families of classifiers, were DBOW (F1=0.410),
BERT MPNet (F1=0.627), BOW (F1=0.635), and BERT NLISTS (F1=0.689). Notably, fine-tuning BERT
directly on the IMRAD classification task yielded substantially superior performance compared to
non-fine tuned BERT models, which themselves outperformed models fine-tuned on transfer learning
tasks (MPNet, NLISTS). The relatively strong performance of the non-fine tuned OpenAI embedding
model—without task-specific fine-tuning—suggests that LLM-derived representations capture valuable
domain-general linguistic patterns relevant to IMRAD classification. Interestingly, the most recent
OpenAI model exhibited inferior performance on this task compared to the earliest model (F1=0.709 vs
0.722).</p>
      <p>Classification Method Performance: Performance varied substantially across classification
approaches when assessed using the maximum of maxima across representation methods. Linear
classifiers achieved the highest performance (F1=0.753), closely followed by tree -based classifiers
(F1=0.752). Neural classifiers ranked third (F1=0.745), while instance -based classifiers exhibited lower
performance (F1=0.732). Clustering methods performed notably poorly, with a maximum F1 score of
0.685. The superior performance of supervised methods relative to unsupervised clustering suggests
that IMRAD section placement involves class-specific patterns that require labelled training data to
capture efectively.</p>
      <p>Overall, the optimal configuration paired BERT IMRAD with linear classification (F1=0.753). This
model outperformed other top configurations, such as linear classification with early OpenAI
embeddings (F1=0.722) or neural classification with BERT IMRAD (F1=0.745).
4We use Hugging Face for BERT for sequence classification, and Scikit-learn with default parameters for other methods
5Nine (9) representation techniques multiplied by nineteen classification methods, plus one end-to-end model.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
      <sec id="sec-4-1">
        <title>4.1. Key Findings and Interpretation</title>
        <p>Our results provide several important insights into IMRAD classification of citation contexts:</p>
        <p>Eficacy of Domain-Specific Fine-Tuning: The marked superiority of BERT IMRAD (F1=0.753)
over other BERT variants and general-purpose embeddings demonstrates that task and domain-specific
ifne-tuning substantially enhances performance for specialized text classification tasks. Fine-tuning
BERT for this IMRAD classification task produced a more separable representation space than either
generic BERT models or fine-tuning on unrelated tasks (NLISTS, MPNet). This finding suggests that the
benefits of transfer learning depend critically on the alignment between the fine-tuning objective and
the target task.</p>
        <p>Prevalence of Linearity in the Feature Space: The slightly superior performance of linear
classifiers relative to neural approaches (F1 = 0.753 vs. 0.745) suggests that the fine -tuned BERT
representation yields a feature space where relatively simple, approximately linear decision boundaries
already capture most of the signal needed for IMRAD classification. This does not imply that neural
networks are fundamentally less capable, but rather that, in our setting, linear models may be easier
to optimize reliably: linear classifiers with convex loss functions have a single global optimum and
are well handled by standard solvers, whereas neural networks involve non-convex objectives and can
converge to sub-optimal solutions under default hyperparameters and limited tuning.</p>
        <p>Limited Benefits of Unsupervised Approaches: The poor performance of clustering methods
(F1=0.685) indicates that the IMRAD structure cannot be reliably identified through unsupervised
learning. This suggests that while implicit partitions may exist in the feature space, they do not
align with IMRAD categories without explicit labelled training data. Although citation contexts across
distinct IMRAD sections may display similar surface-level linguistic patterns, accurately identifying their
diferences requires supervised learning methods capable of capturing subtle discriminative features.</p>
        <p>LLM-Derived Embeddings as a Practical Alternative: The strong out-of-the-box performance
of OpenAI’s text-embedding-ada-002 (F1=0.722) suggests that embeddings derived from large language
models capture suficient linguistic information for reasonable IMRAD classification without
taskspecific fine-tuning. While this approach underperforms BERT IMRAD, it ofers a practical advantage
for researchers lacking the computational resources or labelled data for fine-tuning. Furthermore, our
results indicate that increased model recency and architectural sophistication do not necessarily yield
superior performance for this task. Specifically, the more recent text -embedding-3-large model did not
outperform the oldest model in our experiments (F1 = 0.709 vs. 0.722).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Limitations and Considerations</title>
        <p>Several limitations warrant acknowledgment. First, this study evaluated models on a single corpus
(PLOS journals, 2001-2013); generalization to other disciplines, time periods, or publication venues
remains untested. Second, citation contexts in scientific articles may difer systematically across genres
(e.g., editorials, reviews, policy documents, etc.). Third, the class imbalance in the InTeReC corpus may
advantage majority classes (Discussion). Fourth, we rely on a single random seed and one run per
model, so our scores do not capture variability arising from stochastic training efects. Future work
should complement our findings with multi-seed experiments and confidence intervals.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Implications for Research and Practice</title>
        <p>These findings hold significance for multiple communities. For Scientific Text Mining, the demonstrated
feasibility of automatic IMRAD classification enables new analyses of how citations function across
article sections, potentially revealing disciplinary diferences in citation practices and argumentative
structures. For Argument Mining, automated IMRAD classification provides a structural foundation
for moving beyond generic citation analysis toward richer rhetorical network graphs. For Intelligent
Writing Assistance, these models could be integrated into writing tools to provide authors with
realtime feedback about citation context placement, supporting more rhetorically efective positioning of
citations. For Bibliometric Studies, automated IMRAD classification could enhance citation analysis
by contextualizing citations within their structural location, enabling more granular investigation of
research influence and knowledge transfer patterns.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Future Research Directions</title>
        <p>Several avenues merit further exploration, including cross-domain evaluation to assess whether models
trained on PLOS articles generalize to citations from other domains. Another promising direction is
the fine -tuning of large language model embeddings, particularly OpenAI embeddings, to determine
whether task-specific adaptation can reduce the performance gap with BERT IMRAD. Ensemble methods
also warrant investigation, as combining linear and neural classifiers may balance interpretability with
expressive power. In addition, systematic error analysis could help identify linguistic phenomena
responsible for misclassifications and guide refinements in feature engineering. Finally, evaluating
temporal generalization would clarify whether models trained on older articles remain efective on
contemporary publications, thereby testing the temporal stability of IMRAD-specific linguistic markers.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>This study demonstrates that machine learning can efectively predict the IMRAD section from which a
citation context originates, with optimal performance achieved through domain-specific fine-tuning of
BERT combined with linear classification (F1=0.753). The strong performance of fine-tuned models
underscores the value of task-specific adaptation, while the competitive performance of non-fine tuned
LLM-derived embeddings indicates that modern language models capture relevant linguistic structure
without explicit fine-tuning. Linear classifier superiority suggests that IMRAD classification leverages
relatively interpretable linguistic patterns, contrasting with the opacity of deep neural approaches. These
ifndings advance our understanding of how machine learning can be applied to structure recognition in
scientific writing and open pathways for practical applications in research evaluation, writing assistance,
and bibliometric analysis.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This study was funded by the Canadian Social Sciences and Humanities Research Council (grant number
756-2024-0557). Parts of this work were developed within the framework of the author’s PhD thesis.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author used Grammarly to check grammar, spelling and
improve prose. After using this tool, the author reviewed and edited the content as needed and takes
full responsibility for the publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lloyd</surname>
          </string-name>
          ,
          <article-title>Least squares quantization in PCM</article-title>
          ,
          <source>IEEE Transactions on Information Theory</source>
          <volume>28</volume>
          (
          <year>1982</year>
          )
          <fpage>129</fpage>
          -
          <lpage>137</lpage>
          . doi:
          <volume>10</volume>
          .1109/TIT.
          <year>1982</year>
          .
          <volume>1056489</volume>
          , conference Name:
          <source>IEEE Transactions on Information Theory.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Fix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J. L.</given-names>
            <surname>Hodges</surname>
          </string-name>
          ,
          <string-name>
            <surname>Discriminatory Analysis - Nonparametric Discrimination</surname>
          </string-name>
          : Consistency Properties,
          <source>Technical Report ADA800276</source>
          , California Univ. Berkeley,
          <year>1951</year>
          . URL: https://apps.dtic. mil/sti/citations/ADA800276, section: Technical Reports.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Cover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hart</surname>
          </string-name>
          ,
          <article-title>Nearest neighbor pattern classification</article-title>
          ,
          <source>IEEE Transactions on Information Theory</source>
          <volume>13</volume>
          (
          <year>1967</year>
          )
          <fpage>21</fpage>
          -
          <lpage>27</lpage>
          . URL: http://ieeexplore.ieee.org/document/1053964/. doi:
          <volume>10</volume>
          .1109/TIT.
          <year>1967</year>
          .
          <volume>1053964</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Rocchio</surname>
          </string-name>
          ,
          <article-title>Relevance feedback in information retrieval</article-title>
          , in: G. Salton (Ed.),
          <article-title>The Smart retrieval system - experiments in automatic document processing</article-title>
          , Prentice-Hall, Englewood Clifs, NJ,
          <year>1971</year>
          , pp.
          <fpage>313</fpage>
          -
          <lpage>323</lpage>
          . URL: https://www.bibsonomy.org/bibtex/c18d843e34fe4f8bd1d2438227857225.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Robbins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Monro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Stochastic</given-names>
            <surname>Approximation</surname>
          </string-name>
          <string-name>
            <surname>Method</surname>
          </string-name>
          ,
          <source>The Annals of Mathematical Statistics</source>
          <volume>22</volume>
          (
          <year>1951</year>
          )
          <fpage>400</fpage>
          -
          <lpage>407</lpage>
          . URL: https://projecteuclid.org/journals/annals-of-mathematical-statistics/ volume-22/issue-3/A-
          <string-name>
            <surname>Stochastic-</surname>
          </string-name>
          Approximation-Method/
          <year>10</year>
          .1214/aoms/1177729586.full. doi:
          <volume>10</volume>
          . 1214/aoms/1177729586, publisher: Institute of Mathematical Statistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Vapnik</surname>
          </string-name>
          ,
          <article-title>Pattern recognition using generalized portrait method, Automation</article-title>
          and
          <string-name>
            <given-names>Remote</given-names>
            <surname>Control</surname>
          </string-name>
          (
          <year>1963</year>
          )
          <fpage>774</fpage>
          -
          <lpage>780</lpage>
          . URL: https://www. semanticscholar.org/paper/Pattern-recognition
          <article-title>-using-generalized-portrait-</article-title>
          <string-name>
            <surname>Vapnik</surname>
          </string-name>
          /
          <year>7cabbdf6a7288d15e26fa6ea504009bab3d1edf4</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Berkson</surname>
          </string-name>
          ,
          <article-title>Application of the Logistic Function to Bio-Assay</article-title>
          ,
          <source>Journal of the American Statistical Association</source>
          <volume>39</volume>
          (
          <year>1944</year>
          )
          <fpage>357</fpage>
          -
          <lpage>365</lpage>
          . URL: https://www.jstor.org/stable/2280041. doi:
          <volume>10</volume>
          .2307/2280041, publisher: [American Statistical Association, Taylor &amp; Francis, Ltd.].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Hoerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Kennard</surname>
          </string-name>
          , Ridge Regression:
          <article-title>Biased Estimation for Nonorthogonal Problems</article-title>
          , Technometrics
          <volume>12</volume>
          (
          <year>1970</year>
          )
          <fpage>55</fpage>
          -
          <lpage>67</lpage>
          . URL: https://www.jstor.org/stable/1267351. doi:
          <volume>10</volume>
          .2307/1267351, publisher: [Taylor &amp; Francis, Ltd., American Statistical Association, American Society for Quality].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>K.</given-names>
            <surname>Crammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Dekel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Keshet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shalev-Shwartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer</surname>
          </string-name>
          , Online
          <string-name>
            <surname>Passive-Aggressive</surname>
            <given-names>Algorithms</given-names>
          </string-name>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>7</volume>
          (
          <year>2006</year>
          )
          <fpage>551</fpage>
          -
          <lpage>585</lpage>
          . URL: http://jmlr.org/papers/v7/ crammer06a.html.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J. N.</given-names>
            <surname>Morgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Sonquist</surname>
          </string-name>
          ,
          <article-title>Problems in the Analysis of Survey Data, and a Proposal</article-title>
          ,
          <source>Journal of the American Statistical Association</source>
          <volume>58</volume>
          (
          <year>1963</year>
          )
          <fpage>415</fpage>
          -
          <lpage>434</lpage>
          . URL: http://www.tandfonline.com/doi/ abs/10.1080/01621459.
          <year>1963</year>
          .
          <volume>10500855</volume>
          . doi:
          <volume>10</volume>
          .1080/01621459.
          <year>1963</year>
          .
          <volume>10500855</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>Breiman</surname>
          </string-name>
          , Random Forests,
          <source>Machine Learning</source>
          <volume>45</volume>
          (
          <year>2001</year>
          )
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          . URL: https://doi.org/10.1023/A: 1010933404324. doi:
          <volume>10</volume>
          .1023/A:
          <fpage>1010933404324</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>P.</given-names>
            <surname>Geurts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ernst</surname>
          </string-name>
          , L. Wehenkel,
          <article-title>Extremely randomized trees</article-title>
          ,
          <source>Machine Learning</source>
          <volume>63</volume>
          (
          <year>2006</year>
          )
          <fpage>3</fpage>
          -
          <lpage>42</lpage>
          . URL: https://doi.org/10.1007/s10994-006-6226-1. doi:
          <volume>10</volume>
          .1007/s10994-006-6226-1.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <article-title>Greedy function approximation: A gradient boosting machine</article-title>
          .,
          <source>The Annals of Statistics</source>
          <volume>29</volume>
          (
          <year>2001</year>
          ). URL: https://projecteuclid.org/journals/annals-of-statistics/volume-29/issue-5/ Greedy-function
          <article-title>-approximation-A-gradient-boosting-</article-title>
          <source>machine/10</source>
          .1214/aos/1013203451.full. doi:
          <volume>10</volume>
          .1214/aos/1013203451.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rosenblatt</surname>
          </string-name>
          ,
          <article-title>The perceptron: A probabilistic model for information storage and organization in the brain</article-title>
          ,
          <source>Psychological Review</source>
          <volume>65</volume>
          (
          <year>1958</year>
          )
          <fpage>386</fpage>
          -
          <lpage>408</lpage>
          . doi:
          <volume>10</volume>
          .1037/h0042519, place: US Publisher: American Psychological Association.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Rumelhart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <article-title>Learning Internal Representations by Error Propagation</article-title>
          ,
          <source>Technical Report ADA164453</source>
          , University of San Diego, California,
          <year>1985</year>
          . URL: https://apps.dtic. mil/sti/citations/ADA164453, section: Technical Reports.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>