<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>ORCID:</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Disciplinary Variation in Syntactic A Corpus Analysis of Professional Academic Writing Complexity:</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Javier Pérez-Guerra</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elizaveta A. Smirnova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HSE University</institution>
          ,
          <addr-line>38 Studencheskaya Street, Perm, 614070</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Vigo</institution>
          ,
          <addr-line>Campus Universitario, Vigo, E-36310</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This study deals with the analysis of syntactic complexity in professional academic writing and is based on a corpus of so-called 'hard' and 'soft' papers published in leading international journals. We aim at describing the main complexity features of academic discourse and testing the hypothesis that there is considerable disciplinary variation in linguistic complexity. We conclude that, first, clausal complexity strategies are more prevalent in the 'hard' sciences, while phrasal-complexity features dominate in the 'soft' ones. Second, the data reveal a continuum across subdisciplines within the broad categories of 'soft' and 'hard' genres with respect to the adoption of complexity strategies.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Corpus analysis</kwd>
        <kwd>disciplinary variation</kwd>
        <kwd>academic discourse</kwd>
        <kwd>academic writing</kwd>
        <kwd>syntactic complexity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The phenomenon of complexity has been extensively approached in corpus linguistics over the
recent years. Specifically, the complexity of writing has been studied in terms of the comparison of
L2 and L1 writing [e.g. 1], correlations between text complexity, language proficiency and task types
[e.g. 2], and the development of text complexity after intensive instruction [e.g. 3]. However,
complexity in professional academic writing has been relatively under-researched to date despite the
potential pedagogical implications of such studies. In this respect, we contend that following the
linguistic conventions of a particular discipline plays a crucial role in identifying the writers as
experts in their own discourse communities [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. From this perspective, a research article can serve as
a benchmark for optimal academic writing, providing learners with “a rich and authentic introduction
to the complexities and nuances of the genre” [5: 3]. This study reports the empirical analysis of
linguistic complexity features which aims, first, to describe the complexity features of research
articles written by professional authors and, second, to test the hypothesis that linguistic complexity
varies across disciplines.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Data and methodology</title>
      <p>The analysis of linguistic complexity in professional academic writing has been conducted on a
775,000-word corpus of research papers in four ‘soft’ arts and social sciences (business studies,
linguistics, history and political science), and four ‘hard’ life and physical sciences (mathematics,
engineering, chemistry and physics) which were published in leading peer-review journals indexed in
Scopus Quartile 1, in 2016 and 2017. Once collected, the texts were manually cleared from tables,
formulas, graphs, charts, metadata and reference lists for further analysis. The size and details of the
corpus are given in Table 1.</p>
      <sec id="sec-2-1">
        <title>No. texts</title>
      </sec>
      <sec id="sec-2-2">
        <title>Word totals</title>
        <p>Journals
16
18
13
17
64
10
10
10
11
41
97,947
95,852
98,430
99,003
391,232
95,350
95,603
99,303
93,366
383,622</p>
        <sec id="sec-2-2-1">
          <title>Cell Chemical Biology (CCB)</title>
          <p>Chem
Physics Letters B (PL)</p>
          <p>Reviews in Physics (RP)</p>
          <p>Compositio Matematica (CM)
The Journal of Differential Geometry (JDG)</p>
          <p>Automatica (Auto)
Materials Characterisation (MC)</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>The Journal of Management (JM) The Journal of Management Studies (JMS) Applied Linguistics (AL) Lingua (Ling)</title>
          <p>Contemporary European History (CEH)
The Journal of Modern History (JMH)</p>
          <p>Political Analysis (PA)
World Politics (WP)</p>
          <p>In this study we undertake both the quantitative analysis of measures automatically generated by
the complexity analyser and the qualitative scrutiny of a number of syntactic patterns associated with
syntactic complexity. Firstly, to accomplish the quantitative analysis, the corpus texts were processed
using Lu’s L2 Syntactic Complexity Analyser (hereafter L2SCA). L2SCA provided the 14 indices
given in Table 2 along with their descriptions, as in Lu [6: 43]. Such indices were categorised into: (i)
metrics of structural complexity: indices reporting the length of units (sentences, T-units, clauses2),
measured by counting the number of words; (ii) metrics of syntactic complexity: indices reflecting
syntactic depth and dependency, that is, those based on coordination and subordination ratios as well
as on clausal/T-unit embedding within other superordinate units; and (iii) metrics of categorial
complexity: indices expressing the pervasiveness of nominal and verbal categories in the text.</p>
          <p>
            At the second stage of the analysis, we carried out the qualitative analysis of the clausal and the
phrasal complexity features, based on the taxonomy in Staples et al. [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]. The features are:
sentencefinal adverbial clauses of different types, wh complement clauses, verb + that-clauses, nouns,
attributive adjectives, premodifying nouns and of-genitives. The analysis of such features required
extensive manual disambiguation of the data examples.
2 The notion of a T-unit is extensively used in complexity studies and is defined as “the shortest terminable units into which a connected
discourse can be segmented without leaving any residue” [7: 34]. Bardovi-Harllg [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ] notes that a T-unit normally comprises an independent
along with its dependent clauses. For example, the expression This would certainly continue to be the case with the CNT, but the UGT fared
differently thanks to the support of the PSOE, its European partners and even the Spanish government, who had a strong interest in
weakening the Communists (CEH-2016-4) consists of one sentence, two T-units (This would certainly continue to be the case with the CNT
and the UGT fared differently thanks to the support of the PSOE, its European partners and even the Spanish government, who had a strong
interest in weakening the Communists) and three clauses (This would certainly continue…, …but the UGT fared differently… and …who
had a strong interest…).
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>The automated complexity indices are given in Table 3.</p>
      <p>In an attempt to determine the relative weights of the complexity indices, a binomial linear
regression analysis was applied to the data, implemented via the function ‘glm’ (‘stats’ package, R
Core Team 2020). We operationalised a (backward-steps) reduction of the number of indices that led
to the model in (1), with only the indices VPT (Verb phrases per T-unit), DCS (Dependent clause
ratio), TS (T-unit/sentence ratio) and CPT (Coordinate phrases per T-unit). Both the C(oncordance)
0.918 and Nagelkerke R2 0.653 discrimination indices indicate that the model is very good at
explaining the variation.
(1)</p>
      <p>Definitive glm model (‘***’: 0,001, ‘*’: 0,05)</p>
      <p>Estimate Std, Error z value Pr(&gt;|z|)
(Intercept) -25,9115 4,0116 -6,459 1,05e-10 ***
vpt 3,4756 1,0531 3,300 0,000966 ***
dcs 10,6276 5,0567 2,102 0,035580 *
ts 10,1416 2,6373 3,845 0,000120 ***
cpt 3,8392 0,7312 5,250 1,52e-07 ***</p>
      <p>The interpretation of the findings revealed by the statistical analysis of the complexity indices per
broad discipline, that is, hard and soft sciences, is as follows. The reduction of the indices led to a
model with only 4 indices evincing different dimensions of linguistic complexity:</p>
      <p>(i) syntactic complexity mirrored by pervasive coordination, as reflected by the index CPT, which
calculates the ratio of coordinated phrases per T-unit</p>
      <p>(ii) syntactic complexity determined by subordination within clausal units, as evinced by the index
DCC, which expresses the amount of subordinate dependent clauses in matrix clauses, and in
sentences, which has been corroborated by the statistical significance of the index TS, a telling
indicator of the ratio of T-units per sentence</p>
      <p>(iii) categorial complexity associated with the frequency of, specifically, verbal constituents in
Tunits, here captured by the index VPT.</p>
      <p>
        Random Forests have demonstrated, on the one hand, that, out of the indices that proved to be very
strong in the model, those measures evincing complexity triggered by coordination (CPT) and by the
profusion of verbal categories (VPT), contribute to the variation of hard versus soft science to a
greater extent than DCC and TS. On the other hand, the probability of higher values in the four
complexity indices increases in academic writings categorised as soft science. In other words, greater
ratios of coordination, subordination and the ‘verby’ status of texts can be taken as proxies for the
categorisation of a research paper within the domain of social sciences and humanities. These results
are in line with Biber et al, [10: 29] when they claim that “complexity is not a single unified construct,
and it is therefore not reasonable to suppose that any single measure will adequately represent this
construct”. However, some remarks are in order here as regards the interpretation of our findings in
light of the conclusions drawn by Biber and colleagues. In their multidimensional analysis of
academic writing versus other more informal genres, Biber et al, [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] found that high(er) phrasal
complexity and low(er) clausal complexity are characteristic features of academic English (as well as
of newspaper and magazine writings). By contrast, the type of complexity evinced in personal,
professional (even academic) spoken genres, as well as in popular written (novels, personal essays)
discourse, is fundamentally clausal. Specifically, they contend that T-unit- and subordination-based
(i,e, clausal) measures are not typical of academic writing but of conversational discourse, whereas
nominal/prepositional (i,e, phrasal) measures are good indicators of academic writing. The statistical
modeling of the complexity indices reported in this section has shown that subordination,
coordination and the ‘verby’ status of sentences (or, better, T-units) are defining features of soft
academic writing. As we see it, this conclusion does not invalidate a dominantly phrasal
characterisation of academic writing when compared to more informal speech-based/related
discourse, but gives support to the multifaceted nature of academic writing.
      </p>
      <p>Subsequently, a more qualitative analysis of the frequencies of the features associated with clausal
and phrasal complexity was carried out. The results of the such an analysis are shown in Figure 2,
which provides the normalised frequencies (per 100,000 words) of the features.</p>
      <p>All the differences in the use of the complexity features in hard and in soft sciences were found to
be statistically significant at the level of 1%, except that of verb+that-clauses, which was significant
at the 5% level. As can be seen in Figure 2, adverbial clauses were found to be more common in the
corpus of the hard-science papers. A closer look at the types of adverbial clauses extensively
employed in life and physical sciences revealed that the most frequently used one is the conditional
clause, which accounts for almost a third of all adverbial clauses. This type of adverbial clauses is
typically used in the comments for various calculations, formulas and theorems (see example 1). As
regards the two features evincing complementation strategies, wh-clauses prevail in the soft research
papers, whereas that-clauses are more frequent in the hard disciplines. Finally, the data demonstrates
that, overall, phrasal complexity features, particularly, adjectival and prepositional phrases prevail in
the soft-science texts, while nominal categories are more frequent in the hard sciences, particularly in
chemistry, where they are used in long names of chemical entities and processes (see example 2).
(1) The next lemma expresses the important fact that if qC &gt; 0 and if the excess measured
relative to C is much smaller than the excess measured relative to pairs of planes with
higher-dimensional axes… (JDG-2017-3).
(2) In addition, methyliminodiacetic acid (MIDA)-protected boronate esters were well tolerated
(Chem-2016-4)</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>This study has tackled the analysis of linguistic complexity in professional academic writing in
English. The analysis of automated indices of complexity in a corpus of research articles published in
leading journals in hard (mathematics, chemistry, physics, engineering) and soft (linguistics, history,
business, political science) science papers led to the following conclusions. Soft sciences demonstrate
a significantly larger number of features associated with syntactic complexity, subordination and
coordination ratios than the hard-science genre. The data have also revealed that the
clausalcomplexity indices, in particular, the occurrence of sentence-final adverbial clauses, are significantly
more frequent in the corpus of the hard-science papers. Phrasal complexity, measured here by the
amount of adjectival and prepositional phrases, proved to prevail in the soft-science category, whereas
the hard-science texts exhibited greater ratios of nominal categories.</p>
      <p>
        An in-depth description of linguistic complexity in professional academic texts, along the lines of
analyses of objectively depicted indices, can benefit the teaching of EAP/ESP writing in terms of
guiding the production of discipline-specific language-learning materials that will address the needs
of learners of different sciences in a more effective way. From the perspective of Data Driven
Learning (DDL) approaches [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], EAP/ESP practitioners could employ teaching materials with
examples from research papers in a particular discipline or group of disciplines (hard vs soft) with the
purpose of helping students learn how to meet the necessary language and stylistic conventions
established in a specific discipline. In this vein, concordance lines with the most common finite
adverbial clauses could for example be employed to demonstrate the way in which clausal complexity
is achieved and realised in hard sciences, while occurrences of adjectival and prepositional phrases
from papers in soft disciplines would serve as an illustration of the type of phrasal complexity in this
domain.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lambert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nakamura</surname>
          </string-name>
          ,
          <article-title>Proficiency‐related variation in syntactic complexity: A study of English L1 and L2 oral descriptive discourse</article-title>
          .
          <source>International Journal of Applied Linguistics</source>
          <volume>29</volume>
          (
          <issue>2</issue>
          ) (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          . doi:
          <volume>10</volume>
          .1111/ijal.12224
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Casal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Syntactic complexity and writing quality in assessed first year L2 writing</article-title>
          .
          <source>Journal of Second Language Writing</source>
          <volume>44</volume>
          (
          <year>2019</year>
          )
          <fpage>51</fpage>
          -
          <lpage>62</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.jslw.
          <year>2019</year>
          .
          <volume>03</volume>
          .005
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Mazgutova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kormos</surname>
          </string-name>
          ,
          <article-title>Syntactic and lexical development in an intensive English for Academic Purposes programme</article-title>
          .
          <source>Journal of Second Language Writing</source>
          <volume>29</volume>
          (
          <year>2015</year>
          )
          <fpage>3</fpage>
          -
          <lpage>15</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.jslw.
          <year>2015</year>
          .
          <volume>06</volume>
          .004
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hyland</surname>
          </string-name>
          ,
          <article-title>As can be seen: Lexical bundles and disciplinary variation</article-title>
          .
          <source>English for Specific Purposes</source>
          <volume>27</volume>
          (
          <issue>1</issue>
          ) (
          <year>2008</year>
          )
          <fpage>4</fpage>
          -
          <lpage>21</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.esp.
          <year>2007</year>
          .
          <volume>06</volume>
          .001
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. F.</given-names>
            <surname>Kelly-Laubscher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Muna</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>van der Merwe, Using the research article as a model for teaching laboratory report writing provides opportunities for development of genre awareness and adoption of new literacy practices</article-title>
          .
          <source>English for Specific Purposes</source>
          <volume>48</volume>
          (
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.esp.
          <year>2017</year>
          .
          <volume>05</volume>
          .002
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <article-title>A corpus-based evaluation of syntactic complexity measures as indices of college-level ESL writers' language development</article-title>
          .
          <source>TESOL Quarterly</source>
          <volume>45</volume>
          (
          <issue>1</issue>
          ) (
          <year>2011</year>
          )
          <fpage>36</fpage>
          -
          <lpage>62</lpage>
          . doi:
          <volume>10</volume>
          .5054/tq.
          <year>2011</year>
          .240859
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K. W.</given-names>
            <surname>Hunt</surname>
          </string-name>
          ,
          <article-title>Differences in grammatical structures written at three grade levels: The structures to be analysed by transformational methods</article-title>
          .
          <source>Report no. CRP-1998</source>
          . Tallahasser: Florida State University,
          <year>1964</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Bardovi-Harlig</surname>
          </string-name>
          ,
          <article-title>A second look at T-unit analysis: Reconsidering the sentence</article-title>
          .
          <source>TESOL quarterly 26(2)</source>
          (
          <year>1992</year>
          )
          <fpage>390</fpage>
          -
          <lpage>395</lpage>
          . doi:
          <volume>10</volume>
          .2307/3587016
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Staples</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Egbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Biber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <article-title>Academic writing development at the university level: Phrasal and clausal complexity across level of study, discipline, and genre</article-title>
          .
          <source>Written Communication</source>
          <volume>33</volume>
          (
          <issue>2</issue>
          ) (
          <year>2016</year>
          )
          <fpage>149</fpage>
          -
          <lpage>183</lpage>
          . doi:
          <volume>10</volume>
          .1177%2F0741088316631527
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Biber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Poonpon</surname>
          </string-name>
          ,
          <article-title>Should we use characteristics of conversation to measure grammatical complexity in L2 writing development? TESOL Quarterly 45(1) (</article-title>
          <year>2011</year>
          )
          <fpage>5</fpage>
          -
          <lpage>35</lpage>
          . doi:
          <volume>10</volume>
          .5054/tq.
          <year>2011</year>
          .244483
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D.</given-names>
            <surname>Biber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Poonpon</surname>
          </string-name>
          ,
          <article-title>Pay attention to the phrasal structures: Going beyond T-units - A response to WeiWei Yang</article-title>
          .
          <source>TESOL Quarterly</source>
          <volume>47</volume>
          (
          <issue>1</issue>
          ) (
          <year>2013</year>
          )
          <fpage>192</fpage>
          -
          <lpage>201</lpage>
          . doi:
          <volume>10</volume>
          .1002/tesq.84
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Johns</surname>
          </string-name>
          ,
          <article-title>Should you be persuaded: two samples of data-driven learning materials</article-title>
          .
          <source>English Language Research Journal</source>
          <volume>4</volume>
          (
          <year>1991</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>