<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>S. Buk);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Multiparametric profiling of a linguistic construction: linguoquantitative and machine-learning aspects</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Solomija Buk</string-name>
          <email>solomija@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viktoriia Zhukovska</string-name>
          <email>victoriazhukovska@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksandr Mosiiuk</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ivan Franko National University of Lviv</institution>
          ,
          <addr-line>Universytetska str. 1, Lviv, 79000</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Zhytomyr Ivan Franko State University</institution>
          ,
          <addr-line>Velyka Berdychivska str. 40, Zhytomyr, 10008</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper discusses the results of the linguoquantitative multiparametric profiling of linguistic constructions, focusing on English 'detached nonfinite/ nonverbal with explicit subject'constructions and adopting cognitive-quantitative construction grammar as a theoretical and methodological foundation. Despite extensive research into the linguistic diversity of the syntactic patterns under analysis, a comprehensive parametrization of their linguistic profiles to identify the determining properties that influence their linguistic behavior in present-day English has yet to be carried out. Thus, the statistical platform R was used to achieve two goals: 1) to perform a linguoquantitative parametrization of the formal properties of English 'detached nonfinite/ nonverbal with explicit subject'-constructions based on the corpus data; 2) to verify the results of the linguoquantitative parametrization in a machine learning model and establish the properties with the most significant capacities to differentiate between the DNF/NVESconstructions beyond the corpus. The results of the study prove that the operationalized parameters (factors/factor values) demonstrate different determinative capacity; thus, the linguistic profiles of the DNF/NVES-constructions are distinguished by high, medium, and low linguistic homogeneity, which allows their classification according to the degree of proximity/remoteness in the constructional network. The operationalized linguistic parameters are quite reliable in differentiating types of the DNF/NVES-constructions beyond the corpus.</p>
      </abstract>
      <kwd-group>
        <kwd>cognitive-quantitative construction grammar</kwd>
        <kwd>clause-level construction</kwd>
        <kwd>multiparametric profiling</kwd>
        <kwd>statistical analysis</kwd>
        <kwd>machine-learning experiment 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The development of modern linguistics, particularly construction grammar, is accompanied
by a discussion about increasing the objectivity of research data and finding ways to
improve research precision [1, p. 149; 2]. As a result, the methodology for analyzing
linguistic phenomena is being refined and statistically reliable tools are being actively
employed to verify scientific theories and hypotheses. Traditional methods of language
analysis are being complemented by innovative quantitative corpus-linguistic
methods [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. The use of objective linguistic quantitative analysis tools specifically
designed for processing large amounts of language data increases the degree of evidential
support for the obtained results, revealing new data that would be difficult to identify
through conventional empirical and interpretive approaches. Cognitive-quantitative
construction grammar studies apply advanced quantitative-corpus methods to
parameterize a construction’s linguistic profile [
        <xref ref-type="bibr" rid="ref5 ref6 ref7 ref8">5-8</xref>
        ].
      </p>
      <p>This study aims to discuss the results of the linguoquantitative multiparametric profiling
of linguistic constructions, focusing on English ‘detached nonfinite/ nonverbal with explicit
subject’-constructions (DNF/NVES-constructions) and adopting cognitive-quantitative
construction grammar as a theoretical and methodological foundation. With this in mind,
the following objectives are attained: 1) to perform a linguoquantitative parametrization of
the formal properties of English ‘detached nonfinite/ nonverbal with explicit
subject’constructions based on the corpus data; 2) to verify the results of the linguoquantitative
parametrization in a machine learning model and establish the properties with the greatest
capacity to differentiate between the DNF/NVES-constructions beyond the corpus. By
integrating quantitative corpus linguistics and machine learning approaches, this study
advances understanding of the determining properties that define the linguistic behavior of
the analyzed constructions in present-day English.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>The parameterization method, widely used in engineering and exact sciences, is also applied
in various fields of linguistics, such as linguistic modeling (Yatsenko (2011)), lexicology and
lexicography (Boychuk (2011), Ivakhnenko (2016), Kupriianov (2019)), stylistics and
genre studies (Romanchenko, Stryi (2022)), analysis of the individual writer’s style (Buk
(2021), Davydenko (2014), Tkachenko (2018)), cognitive linguistics (Harmash (2015)),
corpus studies (Luchyk, Ostapova (2017)), and forensic linguistics (Azhniuk (2017)).
Quantitative corpus-based studies employ parameterization to identify, define, and
quantify the essential properties (“diagnostic features” [9, p. 35]) of a linguistic unit.</p>
      <p>English ‘detached nonfinite/nonverbal with explicit subject’-constructions ([[AUGwith]
[NPthe bats] [XPtaking turns to be the starved victim]]; [[NPheart] [XPthumping]];
[[AUGdespite] NP[oil] [XPbeing the lifeblood of industrial (modern) society]]; [[AUGwhat
with] [NPmy three sons] [XPbeing away in the Army]]) as complex clause-level
constructions possess several idiosyncratic properties that set them apart from other
complex syntactic units. Different aspects of these syntactic patterns in both diachrony and
synchrony have been studied from the standpoint of various linguistic approaches and
frameworks such as traditional grammar (Stump (1985), Quirk, Greenbaum, Leech,
Svartvik (1985), Kortmann (1991)), generative grammar (Riemsdijk (1981), Beukema,
Hoekstra (1984), Felser, Britain (2007), Nakagawa (2011)), corpus linguistics (van de Pol
(2012, 2014), van de Pol &amp; Petré (2015)), systemic functional grammar (He, Wu (2015), He,
Yang (2015)), and construction grammar (Riehemann, Bender (1999), Bouzada-Jabois,
Guerra (2016)). In addition, the analyzed syntactic units have been considered in the
dimensions of linguotypology (Haff (2012), Hasselgård (2012)), translation studies
(Davydiuk (2010)) and discourse structure analysis (Asher, Lascarides (2003)). Although
several studies have been undertaken, the linguistic versatility of nonfinite/nonverbal
syntactic patterns with an explicit subject in English raises a number of questions that have
not yet been finally resolved. Primarily, most research has concentrated on the qualitative
rather than quantitative aspects of these units, resulting in a gap in understanding their
functional and contextual characteristics. Moreover, a comprehensive parametrization of
their linguistic profiles to identify the determining properties that influence their linguistic
behavior in present-day English and may reflect speakers’ preferences in categorizing their
linguistic experience has yet to be performed.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Theoretical and methodological background</title>
      <p>
        Cognitive-quantitative construction grammar (CQCxG) is a novel research framework in
cognitive-quantitative grammar studies. This framework triangulates the theoretical and
methodological underpinnings of cognitive-semiotic grammar approaches with analytical
and research tools of quantitative corpus linguistics to investigate general and idiosyncratic
properties of linguistic constructions. Cognitive-quantitative construction grammar
revitalizes the traditional concept of a construction, promoting it to the status of the basic
unit for language representation and analysis. Constructions are conceptualized as holistic
semiotic models, emergent cognitively entrenched symbolic units conventionally used in a
language community, and exhibit pairings of generalized form and meaning/function (plane
of expression and plane of content) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Constructions embrace all language levels, from
morphemes and abstract clausal patterns to text types and genres, ultimately forming an
organized inventory of constructional networks (constructicon), constantly updated and
adjusted by language usage [
        <xref ref-type="bibr" rid="ref11 ref12 ref13">11-13</xref>
        ]. An in-depth examination of the linguistic properties of
a particular linguistic construction can be performed by analyzing its essential
form/meaning properties (prosodic, morphological, syntactic, semantic, distributional,
functional, pragmatic, etc.).
      </p>
      <p>
        The most effective analytical and research tool for examining the essential properties of
a construction is a comprehensive methodology for multiparametric constructional
profiling. In quantitative corpus studies, ‘profiling’ refers to the process of establishing
specific linguistic properties at a particular language level based on quantitative indicators
of this property (parameter) realization in the corpus, whereas the ‘profile’ of a linguistic
unit is a set of established quantitative indicators [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. The set of these properties
determines the linguistic behavior of a construction.
      </p>
      <p>Multiparametric profiling is based on the procedure of a linguoquantitative
parameterization. The procedure entails identifying and statistically verifying a set of
essential linguistic properties (parameters/ factors/ factor values) of the plane of
expression (form) and the plane of content (meaning/function) of a linguistic construction
that constitute its linguistic profile. Thus, a construction’s linguistic profile is an inventory
of its formal and semantic properties (parameters) (morphosyntactic, positional, relational,
referential, distributional, syntactic-functional, collocational-collexeme and
cognitivesemantic), along with corresponding quantitative indicators obtained through their
linguoquantitative verification in corpus data. Linguistic parameters of a construction are
realized in linguistic features at a particular language level – factors, which are then
manifested in specific language categories – factor values.</p>
      <p>
        Nonfinite/nonverbal syntactic patterns with an explicit subject are considered
syntagmatically and semantically complex clause-level constructions in CQCxG and are
referred to as “D(etached) N(on)F(inite)/N(on)V(erbal) (with) E(xplicit)
S(ubject)”constructions (DNF/NVES-constructions). The argument-predicate structure of the
DNF/NVES-constructions minimally consists of a predicate expressed by a nonfinite
(NF)/nonverbal (NV) phrase (XP) and a subject (the external argument of the
nonfinite/nonverbal predicate) expressed by a (pro)nominal phrase (NP). These
clauselevel constructions are partially schematic, represented by obligatory lexically unspecified
slots [SubjNP] and [PredNF/NV], with an open slot for an augmentor [Aug/ØAug] that in
present-day English is expressed by a limited number of units {AUG: with, without, despite,
what with}. The constructions represent a syntactically independent configuration,
detached from a matrix clause by intonation or a punctuation mark. The morphosyntactic
arrangement of the components is displayed as [[Aug/ØAug][SubjNP][PredNF/NV]] (e.g.,
[Augwith][Subjher eyesNP][PredNVopenAdj] (BNC, GOS); [Augdespite] [Subjdesparate
attemptsNP][PredNFto reviveto-Inf her] (BNC, JYB); [Augwhat_with][Subjdelays][PredNFgetting
startedPII] (BNC, HPP); [ØAug][SubjheartNP][PredNFthumpingPI widely] (BNC, EWH). The
DNF/NVES-constructions constitute a taxonomic constructional network in which
individual constructions are projected onto the network as nodes with different degrees of
schematicity, lexical specification, and productivity [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>The quantitative multiparametric constructional profiling methodology is employed to
establish the essential properties of the DNF/NVES-constructions that determine their
linguistic behavior in contemporary English. Multiparametric profiling of the
DNF/NVESconstructions entails linguoquantitative parameterization of their linguistic (formal and
semantic/functional) properties that comprise their constructional multiparametric
linguistic profiles, followed by verification of the obtained data through a machine learning
experiment.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment: corpus sample, statistical software R, and research algorithm</title>
      <p>
        The procedure for linguoquantitative parameterization of a constructional profile is applied
to a research sample of the DNF/NVES-constructions obtained from the British National
Corpus [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The sample includes 11,000 corpus contexts that instantiate five
DNF/NVESconstructions (dt-øaug-SubjPredNF/NV–cxn, dt-with-SubjPredNF/NV–cxn,
dt-despiteSubjPredNF/NV–cxn, dt-without-SubjPredNF/NV–cxn, dt-what_with-SubjPredNF/NV–
cxn), manifested in 35 predicate specifications {NF: VPPI, VPPII, VPto-Inf; NV: NP, AdjP,
AdvP, PP}. The sample size is adequate to be regarded as reliable for linguistic quantitative
profiling since the derived indicators are characterized by a 1,9% relative error. A 5% error
is considered acceptable in linguistic and statistical research, although an error of 20-30%
is also permitted [18, p. 28].
      </p>
      <p>In the current study, the primary focus is on the plane of expression (form) of the
investigated constructions, while the properties of the content plane need a different
methodology. The inventory of formal parameters (factors/factor values) is determined by
the linguistic and constructional nature of the constructions under study. The
DNF/NVESconstructions, structurally complex clause-level constructions, are distinguished by seven
parameters of the plane of expression, which define their morphosyntactic, relational,
referential, syntactic-functional, positional, and distributional properties. The inventory of
the specified parameters is not exhaustive, but it is sufficient for an objective examination
of the linguistic behavior of the DNF/NVES-constructions in contemporary English.</p>
      <p>A considerable number of parameters (factors/factor values) specified to describe
linguistic profiles of the DNF/NVES-constructions in combination with a large amount of
quantitative data cannot be objectively analyzed without using complex statistical
procedures and appropriate computer programs for statistical processing of linguistic data.
As a result, each linguistic parameter (factor / factor value) is submitted to computerized
quantitative verification and subsequent qualitative interpretation.</p>
      <p>
        One of the most widely used analytical tools for quantitative processing of empirical data
in Western corpus-oriented linguistics and usage-based construction grammar is the
statistical data analysis system R (R Development Core Team) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. It is a robust and freely
distributed statistical software environment for data analysis, providing researchers with a
comprehensive toolset for qualitative linguistic and statistical analysis and result
visualization [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The software environment enables users to manipulate extensive
amounts of multidimensional data, employing various processing techniques such as
visualization, primary data analysis, matrix graph construction, scatter plots, etc.
Additionally, it offers classification methods for organizing data, performing statistical
verification, and mathematical modeling.
      </p>
      <p>Parametrization of linguistic profiles of the DNF/NVES-constructions is carried out
according to the following algorithm.</p>
      <p>Step 1. Operationalization of the parameter by identifying the factors of its linguistic
manifestation and defining the values that a particular factor acquires at the appropriate
level of the linguistic structure.</p>
      <p>Step 2. Quantitative analysis of the realization (frequency) of a particular parameter
(factor/factor value) in the corpus.</p>
      <p>Step 3. Statistical analysis of the data obtained using multivariate analysis of variance
(MANOVA), one-factor analysis of variance (ANOVA), and Tukey's multiple comparison
method, quantified with the computer statistical data analysis system R.</p>
      <p>Step 4. Interpretation of quantitative indicators and identification of essential
parameters (factors/factor values) that determine the degree of proximity/remoteness
between constructions in a constructional network.</p>
      <p>Step 5. Verification of the linguoquantitative data in a machine experiment to establish
factors/factor values with the highest capacity to distinguish between the constructions
beyond the corpus.</p>
      <p>
        The application of the algorithm to the factors "Part of speech representation of the
subject" (SubjPOS) of the morphosyntactic parameter and the factor "Register Distribution"
(RegDSTN) of the distributional parameter of the DNF/NVES-constructions has already been
extensively discussed in our prior works [
        <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
        ]. The findings of our previous research
provide the foundation for using statistical methods and machine-learning approach in the
current study.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Results/ Discussion</title>
      <p>5.1. Linguistic profiles of the DNF/NVES-constructions: a computerized
linguoquantitative parametrization
The linguistic profiles of the DNF/NVES-constructions are parametrized through the
quantitative verification of 7 parameters (morphosyntactic, relational, referential,
syntactic-functional, positional, distributional, and punctuational), manifested in 12 factors
and 34 factor values as shown in Table 1. The quantitative data of the operationalized
parameters and their respective factors/ factor values retrieved from the BNC are
standardized by logarithmization. Subsequently, the null (H0) and alternative (H1)
statistical hypotheses are formulated for each of the identified factors:</p>
      <p>Н0: The quantitative differences between the analyzed DNF/NVES-constructions
(dt-øaug</p>
      <sec id="sec-5-1">
        <title>SubjPredNF/NV–cxn, dt-with-SubjPredNF/NV–cxn, dt-despite-SubjPredNF/NV–cxn, dt-without</title>
      </sec>
      <sec id="sec-5-2">
        <title>SubjPredNF/NV–cxn, dt-what_with-SubjPredNF/NV–cxn) within the "FACTOR" are insignificant,</title>
        <p>and any detected quantitative differences are random.</p>
        <p>Н1: The quantitative differences between the analyzed DNF/NVES-constructions
(dt-øaug</p>
      </sec>
      <sec id="sec-5-3">
        <title>SubjPredNF/NV–cxn, dt-with-SubjPredNF/NV–cxn, dt-despite-SubjPredNF/NV–cxn, dt-without</title>
      </sec>
      <sec id="sec-5-4">
        <title>SubjPredNF/NV–cxn, dt-what_with-SubjPredNF/NV–cxn) within the "FACTOR" are significant,</title>
        <p>and the differences found are essential.</p>
        <p>The formulated hypotheses are tested using multivariate analysis of variance
(MANOVA), and the quantified findings are displayed in Table 1.</p>
        <sec id="sec-5-4-1">
          <title>Positional Referential</title>
        </sec>
        <sec id="sec-5-4-2">
          <title>Distributional</title>
        </sec>
        <sec id="sec-5-4-3">
          <title>Syntacticfunctional</title>
        </sec>
        <sec id="sec-5-4-4">
          <title>Relational</title>
          <p>Punctuation marking
(PUNC)
Position to Matrix
Clause (SentPSN)
Coreference with
Matrix Clause
(CoREFR)
Discourse Mode
Distribution
(DiscMdDSTN)
Text Type Distribution
(TxtTpDSTN)
Syntactic Function to
Matrix Clause (FSYN)
Syntactic Relation
with Matrix Clause
(SynREL)
accepted
accepted</p>
          <p>The statistical analysis of the linguistic profiles of the DNF/NVES-constructions did not
reveal statistically significant differences in the realization of 3 factor values (Nominative
case (Nom), Accusative case (Acc), Full coreference (CorrefFull)) out of 32. The quantitative
correlations between these linguistic properties do not differentiate the linguistic profiles
of the DNF/NVES-constructions and suggest general regularities of the subject’s linguistic
embodiment and the reference relations between the DNF/NVES-constructions and the
corresponding matrix clauses.</p>
          <p>The one-way ANOVA indicates the existence of differences but does not explain where
these differences are most prominent. To solve this issue, the post hoc Tukey test is
employed to prevent erroneous rejection of the null hypothesis.</p>
          <p>The Tukey’s multiple comparison method detects the DNF/NVES-constructions that
exhibit statistically significant differences in the realization of factor values. The Tukey’s
test is quantified by comparing the indicators for a specific factor value in pairs of
constructions. It enables the establishment, with a 95% confidence level, which linguistic
features are determining for specific constructions. The multiple comparison method is used
to analyze ten pairs of the DNF/NVES-constructions: 1) dt-what_with-SubjPredNF/NV–cxn and
dt-despite-SubjPredNF/NV–cxn; 2) dt-with-SubjPredNF/NV–cxn and dt-despite-SubjPredNF/NV–
cxn; 3) dt-øaug-SubjPredNF/NV–cxn and dt-despite-SubjPredNF/NV–cxn; 4)
dt-withoutSubjPredNF/NV–cxn and dt-despite-SubjPredNF/NV–cxn; 5) dt-with-SubjPredNF/NV–cxn and
dtwhat_with-SubjPredNF/NV–cxn; 6) dt-øaug-SubjPredNF/NV–cxn and
dt-what_withSubjPredNF/NV–cxn; 7) dt-without-SubjPredNF/NV–cxn and dt-what_with-SubjPredNF/NV–cxn; 8)
dt-øaug-SubjPredNF/NV–cxn and dt-with-SubjPredNF/NV–cxn; 9) dt-without-SubjPredNF/NV–cxn
and dt-with-SubjPredNF/NV–cxn; 10) dt-without-SubjPredNF/NV–cxn and
dt-øaugSubjPredNF/NV–cxn.</p>
          <p>According to the degree of proximity/remoteness by the number of statistically
significant differences in the realization of the specified factor/factor values, we distinguish
constructional linguistic profiles with high, medium, and low linguistic homogeneity.</p>
          <p>The constructions dt-despite-SubjPredNF/NV–cxn, dt-without-SubjPredNF/NV–cxn,
dtwhat_with-SubjPredNF/NV–cxn form a group with high linguistic homogeneity, showing no
statistically significant differences in the analyzed factors/ factor values realization. The
medium linguistic homogeneity group consists of constructions such as
dt-øaugSubjPredNF/NV–cxn and dt-with-SubjPredNF/NV–cxn, which revealed statistically significant
differences in 6 factor values. The group with low linguistic homogeneity includes two
subgroups of the constructions: 1) subgroup dt-with-SubjPredNF/NV–cxn and
dt-despiteSubjPredNF/NV–cxn, dt-without-SubjPredNF/NV–cxn, dt-what_with-SubjPredNF/NV–cxn, where
significant differences were recorded between the with- and what_with-augmented
constructions by 31 factor values, between the with- and despite- augmented constructions
by 30, and between the with- and without-augmented constructions by 27; 2)
dt-øaugSubjPredNF/NV–cxn та dt-despite-SubjPredNF/NV–cxn, dt-without-SubjPredNF/NV–cxn,
dtwhat_with-SubjPredNF/NV–cxn, between which differences in 13, 17, and 20 factor values
were registered, respectively.</p>
          <p>
            The statistical analysis of linguistic profiles of the DNF/NVES-constructions
demonstrates that some factors/factor values significantly influence their linguistic
behavior. Determining parameters (factors/factor values) of the plane of expression of the
analyzed constructions define the degree of proximity and remoteness of the constructions.
To validate the obtained results, a machine learning experiment is conducted to establish
the operationalized factors with the most significant capacity to differentiate between the
DNF/NVES-constructions beyond the corpus.
5.2. Linguistic profiles of the DNF/NVES-constructions: a machine-learning model
The use of machine learning technologies alongside traditional statistical approaches in
language study is becoming more prevalent [
            <xref ref-type="bibr" rid="ref23 ref24">23, 24</xref>
            ]. Researchers have employed
algorithms for machine learning to investigate strategies to increase classifier accuracy in
evaluating readability [
            <xref ref-type="bibr" rid="ref25">25</xref>
            ], as well as assess the prediction powers of ML systems in
crosslinguistic vowel categorization [
            <xref ref-type="bibr" rid="ref26">26</xref>
            ]. These approaches gave been also used to address
problems in the field of Natural Language Processing (NLP) [27]. Drawing on our previous
research [
            <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
            ], we utilize linear discriminant analysis (LDA) to determine the factors
with the greatest potential for separation between the DNF/NVES-constructions outside the
corpus (the British National Corpus) in present-day English. The specialized package MASS
[28, 29] is used to build the model for linear discriminant analysis in R and all graphs are
produced with ggplot2 library for R programming language [30].
          </p>
          <p>Figure 1 displays the results of the distribution of the DNF/NVES-constructions for each
factor, where significant differences were found with a one-factor ANOVA. All data are best
divided along the first two axes LD1 and LD2, simplifying the graphical representation of
information and enabling result comparison.</p>
        </sec>
      </sec>
      <sec id="sec-5-5">
        <title>Subject Determiner (SubjDET)</title>
        <p>Model accuracy = 0,71</p>
      </sec>
      <sec id="sec-5-6">
        <title>Punctuation marking (PUNC) Model accuracy = 0,4</title>
      </sec>
      <sec id="sec-5-7">
        <title>Position to Matrix Clause (SentPSN) Model accuracy = 0,71</title>
      </sec>
      <sec id="sec-5-8">
        <title>Coreference with Matrix Clause (CoREFR) Model accuracy = 0,69</title>
      </sec>
      <sec id="sec-5-9">
        <title>Discourse Mode Distribution (DiscMdDSTN) Model accuracy = 0,4</title>
      </sec>
      <sec id="sec-5-10">
        <title>Text Distribution (TxtTpDSTN) Model accuracy = 0,66</title>
      </sec>
      <sec id="sec-5-11">
        <title>Syntactic Function to Matrix Clause (FSYN) Model accuracy = 0,63</title>
        <p>• – dt-despite-SubjPredNF/NV–cxn,
• – dt-what_with-SubjPredNF/NV–cxn
• – dt-with-SubjPredNF/NV–cxn,
• – dt-øaug-SubjPredNF/NV–cxn,
• – dt-without-SubjPredNF/NV–cxn</p>
      </sec>
      <sec id="sec-5-12">
        <title>Syntactic Relation with Matrix Clause (SynREL) Model accuracy = 0,91</title>
        <p>Figure 1 shows a distinct separation of dt-with-SubjPredNF/NV–cxn and
dt-øaugSubjPredNF/NV–cxn from other constructions submitted to the linear analysis. However, in the
factor "Syntactic Function to Matrix Clause" (FSYN), the dt-despite-SubjPredNF/NV–cxn is also
distinguished.</p>
        <p>The confusion matrices are generated for all specified factors to validate the results
reached from the analysis of graphic materials. Due to space limitations, only the confusion
matrices with the highest and lowest model accuracy are presented in Table 3 and Table 4.
The analysis of the confusion matrices reveals that for with- and øaug-augmented
constructions the Recall and Precision values are pretty high, particularly for models with
an overall accuracy of 0,7 and higher. It suggests a high probability of their correct
extraction by the created machine learning models. The specified factors may be ranked
based on the model accuracy: the factor ‘Syntactic Relation with Matrix Clause’ with the
model’s accuracy of 0,91 is characterized by the highest capacity to differentiate the
analyzed constructions. The factors ‘Subject Determiner’ (0,71), ‘Position to Matrix Clause’
(0,71), ‘Coreference with Matrix Clause’ (0,69), ‘Text Distribution’ (0,66), ‘Syntactic Function
to Matrix Clause’ (0,63) are less reliable in differentiating between the
DNF/NVESconstructions. The factors ‘Punctuation marking’ (0,4) and ‘Discourse Mode Distribution’
(0,4) show the lowest differentiating capacity.</p>
        <p>The results of the machine learning experiment are very similar to those obtained by the
linguoquantitative parameterization of the constructional profiles. However, the overall
efficiency of the constructed machine learning model to solve the problem of distinguishing
the types of the DNF/NVES-constructions beyond the corpus is not sufficient. The model
effectively distinguishes the dt-øaug-SubjPredNF/NV–cxn and dt-with-SubjPredNF/NV–cxn
constructions, despite insufficient overall accuracy. However, dt-despite-SubjPredNF/NV–cxn,
dt-what_with-SubjPredNF/NV–cxn, and dt-without-SubjPredNF/NV–cxn are more challenging to
classify. Provided an effective model is constructed, the specified factors of the
operationalized linguistic parameters can be used to distinguish types of the
DNF/NVESconstructions in the present-day English usage.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>The formal parameters (factors/factor values) are determined by the linguistic and
constructional nature of the DNF/NVES-constructions as complex clausal constructions. The
analysis includes seven parameters of the expression plane of the DNF/NVES-constructions,
which define their morphosyntactic, relational, referential, syntactic-functional, positional,
punctuational, and distributional features. The operationalized parameters (factors/factor
values) reveal different determinative capacities; thus, the linguistic profiles of the
DNF/NVES-constructions are characterized by a certain degree of proximity/remoteness in
the constructional network, which allows their categorization according to the degree of
linguistic homogeneity: 1) high (dt- despite-SubjPredNF/NV–cxn, dt- without-SubjPredNF/NV–
cxn, dt- what_with-SubjPredNF/NV–cxn); 2) medium (dt- øaug-SubjPredNF/NV–cxn and
dt- with-SubjPredNF/NV–cxn); 3) low (subgroup dt- with-SubjPredNF/NV–cxn and dt-
despiteSubjPredNF/NV–cxn, dt- without-SubjPredNF/NV–cxn, dt- what_with-SubjPredNF/NV–cxn;
subgroup dt- øaug-SubjPredNF/NV–cxn and dt- despite-SubjPredNF/NV–cxn, dt-
withoutSubjPredNF/NV–cxn, dt- what_with-SubjPredNF/NV–cxn). The specified factors of the
operationalized linguistic parameters can be utilized to differentiate types of the
DNF/NVES-constructions beyond the corpus in current English usage. Differences in the
quantitative realization of individual factors/factor values within one parameter of a
specific DNF/NVES-construction are determined by intra-constructional variability. In
contrast, quantitative differences in the realization of factors/factor values within one
parameter between different DNF/NVES-constructions are determined by
interconstructional variability.</p>
      <p>The results of this study indicate a need for further research on the discussed issues. In
future studies, it will be interesting to generate other types of machine learning models and
assess their effectiveness in distinguishing between the types of the
DNF/NVESconstructions based on the data sets for the operationalized parameters/factors/factor
values extracted from the corpus.
[27] V. Addanki, S. Durgapu, K. Dorasanaiah, S. Abhishek, Safeguarding SMS: A Dynamic Duo
Approach to Tackle Spam Using LDA and QDA. 2023 Innovations in Power and
Advanced Computing Technologies (i-PACT). IEEE, 2023 1–6. doi:
10.1109/iPACT58649.2023.10434334.
[28] Cran.r-project.org. Package MASS, 2024. URL:
https://cran.rproject.org/web/packages/MASS/MASS.pdf
[29] M. Kuhn, J. Wing, S. Weston, A. Williams, C. Keefer, A. Engelhardt, T. Cooper, et al.</p>
      <p>Package “caret”: Classification and Regression Training. 2023. URL:
https://cran.rproject.org/web/packages/caret/caret.pdf
[30] H. Wickham, ggplot2. Elegant Graphics for Data Analysis, Springer, 2016. doi:
10.1007/978-3-319-24277-4.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Janda</surname>
          </string-name>
          ,
          <source>Cognitive Linguistics in the Year</source>
          <year>2015</year>
          ,
          <article-title>Cognitive Semantics 1 (</article-title>
          <year>2015</year>
          )
          <fpage>131</fpage>
          -
          <lpage>154</lpage>
          . doi:
          <volume>10</volume>
          .1163/
          <fpage>23526416</fpage>
          -
          <lpage>00101005</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Janda</surname>
          </string-name>
          , Quantitative Perspectives in Cognitive Linguistics,
          <source>Review of Cognitive Linguistics</source>
          <volume>17</volume>
          (
          <issue>1</issue>
          ) (
          <year>2019</year>
          )
          <fpage>7</fpage>
          -
          <lpage>28</lpage>
          . doi:
          <volume>10</volume>
          .1075/rcl.00024.jan.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kortmann</surname>
          </string-name>
          ,
          <source>Reflecting on the Quantitative Turn in Linguistics, Linguistics</source>
          <volume>59</volume>
          (
          <issue>5</issue>
          ) (
          <year>2021</year>
          )
          <fpage>1207</fpage>
          -
          <lpage>1226</lpage>
          . doi:
          <volume>10</volume>
          .1515/ling-2019-0046.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Winter</surname>
          </string-name>
          , Statistics for
          <string-name>
            <given-names>Linguists. An</given-names>
            <surname>Introduction Using</surname>
          </string-name>
          <string-name>
            <given-names>R</given-names>
            ,
            <surname>Routledge</surname>
          </string-name>
          , New York, London,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Krawczak</surname>
          </string-name>
          ,
          <article-title>The role of verb polysemy in constructional profiling: A cross-linguistic study of give in the dative alternation</article-title>
          , in: M.
          <string-name>
            <surname>Bouveret</surname>
          </string-name>
          (Ed.), Constructional Approaches to Language, John Benjamins Publishing Company, Amsterdam, Philadelphia,
          <year>2021</year>
          . pp.
          <fpage>75</fpage>
          -
          <lpage>96</lpage>
          . doi:
          <volume>10</volume>
          .1075/cal.29.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Casal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shirai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <article-title>English verb-argument construction profiles in a specialized academic corpus: Variation by genre and discipline</article-title>
          ,
          <source>English for Specific Purposes</source>
          <volume>66</volume>
          (
          <year>2022</year>
          )
          <fpage>94</fpage>
          -
          <lpage>107</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.esp.
          <year>2022</year>
          .
          <volume>01</volume>
          .004.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Zhukovska</surname>
          </string-name>
          ,
          <source>Quantitative Corpus-Based Methods for Construction Grammar Research</source>
          , Zhytomyr Ivan Franko State University Journal. Philological Sciences [
          <article-title>Visnyk Zhytomyrskoho derzhavnoho universytetu imeni Ivana Franka</article-title>
          . Filolohichni nauky]
          <volume>1</volume>
          (
          <issue>99</issue>
          ) (
          <year>2023</year>
          )
          <fpage>93</fpage>
          -
          <lpage>104</lpage>
          . doi:
          <volume>10</volume>
          .35433/philology.1(
          <issue>99</issue>
          ).
          <year>2023</year>
          .
          <volume>93</volume>
          -
          <fpage>104</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Wyroślak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Glynn</surname>
          </string-name>
          ,
          <article-title>Disentangling constructional networks: integrating taxonomic effects into the description of grammatical alternations</article-title>
          , Linguistics
          <string-name>
            <surname>Vanguard</surname>
          </string-name>
          (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .1515/lingvan-2023-0035.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Speelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Geeraerts</surname>
          </string-name>
          ,
          <article-title>Causes for causatives: The case of Dutch doen and laten</article-title>
          , in: T.
          <string-name>
            <surname>Sanders</surname>
          </string-name>
          , E. Sweetser (Eds.),
          <source>Causal Categories in Discourse and Cognition</source>
          . Mouton de Gruyter, Berlin, New York,
          <year>2009</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>204</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hilpert</surname>
          </string-name>
          , Constructional Approaches, in: B.
          <string-name>
            <surname>Aarts</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Bowi</surname>
          </string-name>
          , G. Popova (Eds.),
          <source>The Oxford Handbook of English Grammar</source>
          , Oxford University Press, Oxford,
          <year>2020</year>
          , pp.
          <fpage>106</fpage>
          -
          <lpage>123</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hoffmann</surname>
          </string-name>
          ,
          <source>Construction Grammar: The Structure of English</source>
          , Cambridge University Press, Cambridge.
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Diessel</surname>
          </string-name>
          , The Constructicon:
          <article-title>Taxonomies and Networks (Elements in Construction Grammar)</article-title>
          , Cambridge University Press, Cambridge,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T</given-names>
            <surname>Ungerer</surname>
          </string-name>
          , S. Hartmann, Constructionist Approaches: Past, Present,
          <source>Future (Elements in Construction Grammar)</source>
          , Cambridge University Press, Cambridge,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Speelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Grondelaers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Geeraerts</surname>
          </string-name>
          ,
          <article-title>Profile-Based Linguistic Uniformity as a Generic Model for Comparing Language Varieties</article-title>
          ,
          <source>Computers and the Humanities</source>
          <volume>37</volume>
          (
          <issue>3</issue>
          ) (
          <year>2003</year>
          )
          <fpage>317</fpage>
          -
          <lpage>337</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Divjak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Gries</surname>
          </string-name>
          , Ways of trying in Russian: Clustering Behavioral Profiles,
          <source>Corpus Linguistics and Linguistic Theory</source>
          <volume>2</volume>
          (
          <issue>1</issue>
          ) (
          <year>2006</year>
          )
          <fpage>23</fpage>
          -
          <lpage>60</lpage>
          . doi:
          <volume>10</volume>
          .1515/CLLT.
          <year>2006</year>
          .
          <volume>002</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Zhukovska</surname>
          </string-name>
          ,
          <article-title>Constructional Modeling in the Formalism of Cognitive-Quantitative Construction Grammar</article-title>
          , Messenger of Kyiv National Linguistic University. Series “Philology” [Visnyk Kyivskoho Natsionalnoho Universytetu. Seriia “Filolohiia”]
          <volume>26</volume>
          (
          <issue>2</issue>
          ) (
          <year>2023</year>
          )
          <fpage>51</fpage>
          -
          <lpage>62</lpage>
          . doi:
          <volume>10</volume>
          .32589/
          <fpage>2311</fpage>
          -
          <lpage>0821</lpage>
          .2.
          <year>2023</year>
          .
          <volume>297670</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <source>British National Corpus (BNC)</source>
          ,
          <year>2004</year>
          . URL: https://www.englishcorpora.org/bnc/.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>V. S.</given-names>
            <surname>Perebyinis</surname>
          </string-name>
          ,
          <article-title>Statistical methods for linguists [Statystychni metody dlia linhvistiv]</article-title>
          .
          <source>Nova knyha. Vinnytsia</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R Core</given-names>
            <surname>Team</surname>
          </string-name>
          ,
          <string-name>
            <surname>R:</surname>
          </string-name>
          <article-title>A language and environment for statistical computing</article-title>
          ,
          <year>2024</year>
          . https://www.r-project.
          <source>org/.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E. G. M.</given-names>
            <surname>Hui</surname>
          </string-name>
          , Learn R for Applied Statistics.
          <article-title>With Data Visualizations, Regressions</article-title>
          , and Statistics, Apress, Berkeley,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhukovska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mosiiuk</surname>
          </string-name>
          ,
          <source>Statistical Software R in Corpus-Driven Research and Machine Learning</source>
          ,
          <source>Information Technologies and Learning Tools</source>
          <volume>86</volume>
          (
          <issue>6</issue>
          ) (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhukovska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mosiiuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Buk</surname>
          </string-name>
          ,
          <article-title>Register Distribution of English Detached Nonfinite/Nonverbal with Explicit Subject Constructions: a Corpus-Based and Machine-Learning Approach</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          ,
          <volume>3396</volume>
          ,
          <year>2023</year>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>R.</given-names>
            <surname>Baayen</surname>
          </string-name>
          ,
          <article-title>Analyzing linguistic data</article-title>
          . Cambridge University Press, Cambridge,
          <year>2008</year>
          . doi:
          <volume>10</volume>
          .1017/CBO9780511801686.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Th</surname>
          </string-name>
          . Gries, Statistics for linguistics with R. Mouton de Gruyter, Berlin, New York,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>N.</given-names>
            <surname>Sukhija</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Priya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Arya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kohli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Arya</surname>
          </string-name>
          ,
          <source>Hybrid Ensemble Stacking Model for Gauging English Transcript Readability International Journal of Performability Engineering</source>
          <volume>19</volume>
          (
          <issue>11</issue>
          ) 2023
          <fpage>719</fpage>
          -
          <lpage>727</lpage>
          . doi:
          <volume>10</volume>
          .23940/ijpe.23.11.
          <year>p2</year>
          .
          <fpage>719727</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>G. P.</given-names>
            <surname>Georgiou</surname>
          </string-name>
          ,
          <article-title>Comparison of the Prediction Accuracy of Machine Learning Algorithms in Crosslinguistic Vowel Classification</article-title>
          .
          <source>Scientific Reports</source>
          <volume>13</volume>
          (
          <issue>15594</issue>
          )
          <year>2023</year>
          1-
          <fpage>10</fpage>
          doi: 10.1038/s41598-023-42818-3.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>