<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Common Derivation for Parsing and Generation with Expectation-Based Minimalist Grammars (e-MGs)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Computational Linguistics</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Theoretical Syntax</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavia cristiano.chesi@iusspavia.it</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Merge only Merge</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Expectation-based Minimalist Grammars (e-MGs) are simplified versions of the (Conflated) Minimalist Grammars, (C)MGs, formalized by Stabler (Stabler 1997; Stabler 2011; Stabler 2013) and Phase-based Minimalist Grammars, PMGs (Chesi 2007; Chesi 2005; Stabler 2011). The crucial simplification consists of driving structure building only using lexically encoded categorial top-down expectations. The commitment on a topdown procedure (in e-MGs and PMGs, as opposed to (C)MGs, Chomsky, 1995; Stabler, 2011) allows us to define a core derivation that is the same in both parsing and generation (Momma &amp; Phillips 2018).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Minimalism
        <xref ref-type="bibr" rid="ref11 ref12">(Chomsky 1995; Chomsky 2001)</xref>
        is
an elegant transformational grammatical
framework that defines structural dependencies in
phrasal (i.e. hierarchical) terms simply relying on
one core structure building operation, Merge, that
combines lexical items and the result of other
Merge operations. (1).a is the representative result
of two ordered Merge operations (i.e. Merge(γ,
Merge(α, β)) both taking the items α, β and γ
directly from the lexicon, while (1).b relies on the
so called Internal Merge (Move): the re-Merge of
an item that was already merged in the structure.
As result, Move connects the item at the edge of
the structure (β) with a trace (_β), a phonetically
empty copy of the item that in a previous Merge
operation combined with a hierarchically lower
item (α in (1).b). In both (Conflated) Minimalist
and Phase-based Minimalist Grammars ([C]MGs
and PMGs respectively) Merge and Move are
feature-driven operations, that is, a successful
operation must be triggered by the relevant (categorial)
features matching, and, once these features are
used, they get deleted. Consequently, a feature
pair is always responsible for each operation
        <xref ref-type="bibr" rid="ref28">(unless specific features are left unerased after a
successful operation, as in raising predicates and
successive cyclic movement, Stabler 2011)</xref>
        . One
crucial difference between PMGs and MGs is that
while MGs operate from-bottom-to-top, as
indicated in (2), PMGs structure building operations
apply top-down as schematized in (3)1:
(2)
(3)
      </p>
      <p>Merge(α=X, Xβ) = [α [α=X Xβ]] MGs
Move(+Yα, [… β-Y …]) =</p>
      <p>
        [α [β-Y [+Yα [… β-Y …]]]]
Merge(α=X, Xβ) = [α=X [Xβ]]
Move([α=S +Y[Y Z β]]) =
[α=S +Y[Y Z β] S[… (=Z [Z β]) …]]
PMGs
Another relevant difference between the two
approaches is related to the implementation of
Move: MGs use the “+/-” feature distinction and
the same deletion procedure after matching, while
PMGs do not use “-” features and simply assume
that both “+” and “=” select categorial features,
which are deleted after Merge. In PMGs, “+”
features force memory storage and hence the
movement (downward) of the licensed item, until the
relevant prominent category identifying the
moved item (Z in (3)) is selected. If no proper
selection is found, the sentence is ungrammatical.
CMG as well dispenses the grammar with the
+/feature distinction and only relies on select
features (=X), but it must assume that feature
deletion can be procrastinated (again, for instance, in
* Copyright ©️ 2021 for this paper by its author. Use
permitted under Creative Commons License
Attribution 4.0 International (CC BY 4.0).
1 α and β are lexical items, =X indicates the selection
of X, where X is a categorial feature. Lexical items are
tuples consisting of selections/expectations (=X) and
categories (X, i.e. selected/expected features); for
convenience, select features are expressed by
rightward subscripts, and categories as leftward subscripts.
Similarly, Move is driven by licensing (-Y, leftward
subscripts) and licensors (+Y, rightward subscripts)
features
        <xref ref-type="bibr" rid="ref28">(Stabler 2011)</xref>
        .
raising predicates). Despite the fact that, from a
generative point of view, all these formalisms are
equivalent and they all fall under the so called
mildly-context sensitive domain
        <xref ref-type="bibr" rid="ref28">(Stabler 2011)</xref>
        , it
is worth to appreciate the dynamics of structure
building “on-line”, namely how the derivation
unrolls, word by word. Taking the MGs lexicon (4),
the expected constituents in (1) are built adding
items to the left-edge of the structure at each
Merge/Move application, as described in (5).
(4) LexMG = {[Yα=X], [X -Zβ], [γ=Y +Z]}
(5) i. Merge(Yα=X, X -Zβ) = [Yα [α=X X -Zβ]]
ii. Merge(γ=Y +Z, [Yα [α -Zβ]]) =
      </p>
      <p>[γ=Y +Z [Yα [α -Zβ]]]
iii. Move ([γ+Z [α [α -Zβ]]]) =</p>
      <p>[[-Zβ] γ+Z [γ [α [α _β]]]
An equivalent structure is obtained in PMGs2 as
shown in (7). Notice a minimal difference in the
lexicon (6): the absence of the “-” features.
(6) LexPMG = { [Yα=X], [Z X β], [γ+Z =Y] }
(7) i.</p>
      <p>Merge(Z Xβ, γ+Z =Y) = [[Z Xβ] γ+Z =Y]
ii. Merge([[β] γ =Y], Yα =X) =</p>
      <p>
        [[β] γ [(γ) =Y [Yα =X]]
iii. Move([[β] γ [(γ) [α =X]], Xβ) =
[[β] γ [(γ) [α=X [(α) X_β]]]]
Xβ → M
M = {Xβ}
Xβ ← M
The result of the two derivations is (strongly)
equivalent in hierarchical (and dependency)
terms. The simplicity, in pre-theoretical terms, of
the two descriptions is comparable: while PMGs
must postulate the M storage to implement Move
(as result of the missing selection of a categorial
feature), MGs must postulate independent
workspace to build nontrivial left-branching structures,
for instance before merging a multi-word subject
like “the boy” with its predicate (e.g., “runs”).
Furthermore, both formalisms must restrict the
behavior either of the M buffer operativity or the
accessibility to the -f features to limits the Move
operation
        <xref ref-type="bibr" rid="ref19">(e.g., island constraints, Huang, 1982)</xref>
        .
1.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Top-Down is Better</title>
      <p>
        There are at least three reasons to commit
ourselves to the top-down orientation instead of
remaining agnostic or relying on the mainstream
Minimalist brick-over-brick (from-bottom-to-top)
approach
        <xref ref-type="bibr" rid="ref5">(Chesi 2007)</xref>
        : First, the order in which
2 Move is implemented using a Last-In-First-Out
addressable memory buffer M, where the item (β) with
unselected categorie(s) (X) is stored (“Xβ → M”) and
retrieved (“Xβ ← M”) when selected (i.e. “=X”).
the structure is built is grossly transparent with
respect to the order in which the words are
processed in real-life tasks, both in generation and in
parsing in PMGs, but not in MGs.
      </p>
      <p>
        Second, in PMGs, the simple processing order
of multiple expectations is sufficient to
distinguish between sequential (the last expectation of
a given lexical item) and nested expectations (any
other expectation): The first qualifies as the
transparent branch of the tree (i.e. it is able to license
pending items from the superordinate selecting
item), while constituents licensed by nested
expectations qualify as configurational islands
        <xref ref-type="bibr" rid="ref2 ref6">(Bianchi &amp; Chesi 2006; Chesi 2015)</xref>
        . Moreover,
successive cyclic movement is easily described in
PMGs without relying on feature checking at any
step or non-deterministic assumptions on features
deletion
        <xref ref-type="bibr" rid="ref6">(Chesi 2015)</xref>
        contrary to (C)MGs.
      </p>
      <p>A third logical reason to prefer the top-down
orientation over the bottom-up alternative is
related to the unicity of the root node in tree graphs.
As anticipated, the creation of complex (binary)
branching structures poses a puzzle for (C)MGs:
Independent workspaces must be postulated,
namely [the boy] and [sings … ] phrases must be
created before one can merge with the other:
(8) [VP [DP the boy] [V sings [DP a song]]]</p>
      <p>
        This is the case of “complex” subject or adjunct
(i.e., non-projecting constituents which are simply
composed by more than one lexical item) that
must be the result of (at least) one independent
Merge operation, before this can merge with the
relevant predicate (e.g. [V sings …]3 in (8)).
Processing these constituents represents a major
difference between (bottom-up) MGs and
(topdown) PMGs derivations. While MGs must
decide where to start from (and both solutions are
possible and forcefully logically independent
from parsing or generation, which undeniably
proceed “left-right”), PMGs take advantage of the
“single root condition”
        <xref ref-type="bibr" rid="ref24">(Partee, Meulen &amp; Wall
1993: 439)</xref>
        and avoid this problem:
(9) In every well-formed constituent structure
tree, there is exactly one node that
dominates every node.
      </p>
      <p>
        As indicated in (3), the binary operation Merge
simply produces a hierarchical dependency in
which the dominating (asymmetrically
C3 Considering the inflection “-s” as part of the lexical
element or by (head) moving the root “sing-“ to T is
uninfluential here. This sort of head movement is
implemented lexically in e-MGs (e.g. [T (=V V) eats …].
commanding, in the sense of
        <xref ref-type="bibr" rid="ref20">Kayne 1994</xref>
        ) item is
above the dominated (C-commanded) one. This is
compatible with Stabler notation (10).a-b and
plainly solves the ambiguity of the nature of the
“label” of the constituent
        <xref ref-type="bibr" rid="ref26">(Rizzi 2016)</xref>
        . In this
sense, PMGs (and the e-MGs discussed later) can
adopt directly a more concise description, that is
(10).c, more transparent with respect to the
(Universal) Dependency approach
        <xref ref-type="bibr" rid="ref23">(Nivre et al. 2017)</xref>
        :
Elements are “dependent” when they Merge.
(10) a. MGs
b. (C)MGs
      </p>
      <p>c. (P/e-)MGs
α
&lt;
α=X
=
α=X βX α=X βX βX
=
The higher node (possibly the root) is always the
selecting item (a probe, in minimalist terms), and
it is the first item to be processed. This does not
necessarily imply that this item is linearized
before the selected category (the goal, in minimalist
terms): if the selecting node has multiple selection
needs, it must remain to the right-edge of the
structure to license, locally, the other(s) selection
expectation(s). E.g., if [α=X =Y], [Xβ] and [Yγ], then:
(11) [α=X =Y [Xβ] [(α=Y) [Y γ]]]
In this case, &lt;α, β, γ&gt; would be the default
linearization, but it is easy to derive &lt;β, α, γ&gt; instead,
assuming a simple parameterization on spell-out
in case of multiple select features.</p>
      <p>Here, I will argue that we can push further this
intuition and only rely on (categorial)
expectations, encoded in the lexical items, to guide the
derivation. This leads to the so-called
expectationbased Minimalist Grammars (e-MGs).</p>
      <p>In the following sections, I will sketch a simple
formalization for e-MGs (§2), and the core
derivation algorithm (§3) that would be used both in
Generation and Parsing tasks (§3.2).
2</p>
      <sec id="sec-2-1">
        <title>The Grammar</title>
        <p>
          As (C/P)MGs, e-MGs include a specification of a
lexicon (Lex) and a set of functions (F), the
structure building operations. The lexicon, in turn, is a
finite set composed by words each consisting of
phonetic/orthographic information (Phon) and a
combination of categorical features (Cat),
4 As in MGs, lexical items could be specified both for
phonetic (Phon) and semantic features (Sem). In
eMGs, expectations (=/+X) and expectees (X)
correspond to MGs selectors/licensors and
selectees/licensees respectively. Agreement features indicate
categorial values to be unified
          <xref ref-type="bibr" rid="ref8">(Chesi 2021)</xref>
          .
expressing expect(ations), expected and
agreement categories4. In the end, an optional set of
Parameters (P)
          <xref ref-type="bibr" rid="ref8">(see Chesi 2021)</xref>
          , inducing minimal
modifications to the structure building operations
F and, possibly, to the Cat set, under the fair
assumption that F and Cat are universal. More
precisely, any e-MG is a 5-tuple such that:
(12) G = (Phon, Cat, Lex, F, P), where
        </p>
        <p>Phon, a finite set of phonetic/orthographic
features (i.e., orthographic forms
representing words, e.g., “the”, “smiles”)
Cat, a finite set (morphosyntactic categories,
that can be expect, expected or agreement
features e.g., “D”, “V”… “gen(der)”,
“num(ber)”, “pl(ural)” etc.)
Lex, a set of expressions built from Phon and</p>
        <p>Cat (the lexicon)
F, a set of partial functions from tuples of
expressions to expressions (the structure
building operations)
P, a finite set of minimal transformations of
F and/or Cat (the parameters), producing
F' and Cat', respectively.
2.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Lexical Items and Categories</title>
      <p>Each lexical item l in Lex, namely each word, is a
4-tuple defined as follows5:
(13) l = (Ph, Exp(ect), Exp(ect)ed, Agr(ee)),
Phon, from Phon in G (e.g., “the”)
Exp, a finite list of ordered features from Cat
in G (the category/ies that the item
expects will follow, e.g., =N)
Exped is a finite list of ordered features from
Cat in G (the category/ies that should be
licensed/expected, e.g., N)
Agr(ee) is a structured list of features from</p>
      <p>Cat in G (e.g., gen.fem, num.pl)
All Exp(ect), Exp(ect)ed and Agr(ee) features are
then subsets of Cat in G. In Agr, for instance, a
feminine gender specification (gen.fem) expresses
a subset relation (i.e., “feminine”  “gender”).</p>
      <p>
        For sake of simplicity, each l will be
represented as [Expected(; Agree) Phon =/+Expect] as in (14):
(14) [D the =N], [N; num.pl dogs], [T barks =D]
5 This is the simplest possible implementation.
Attribute-Value Matrices, as in HPSH
        <xref ref-type="bibr" rid="ref25">(Pollard &amp; Sag 1994)</xref>
        or TRIE/compact trees exploiting the sequence of
expectations
        <xref ref-type="bibr" rid="ref29 ref7 ref9">(Chesi 2018; Stabler 2013)</xref>
        are possible
implementations.
      </p>
      <p>We refer to the most prominent (i.e., the first)
Expected feature as the Label (L) of the item. E.g.,
the label L of “the” will be D, while the label of
“barks” will be T. Similarly, let us call S (for
select) the first Expect feature and R the remaining
Expect(actions) (if any).
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Structure Building Operations</title>
      <p>
        Given lx an arbitrary item such that lx = (Px,
Lx/Expedx, Sx/Rx/Expx, Agrx) we can define MERGE
as follows:
(15) MERGE(l1(S1), l2(L2)) =
{1, [ 1( 1)[ 2( 2)]]   1 =  2}
0  ℎ
MERGE is implemented as the usual binary
function that is successful (it returns “1”) and creates
the dependency (asymmetric C-command or
inclusion, in set theoretic terms) (10).c, namely [l1
[l2]], if and only if the label of the subsequent item
(l2) is exactly the one expected by the preceding
item (l1), namely S1 = L2. This is probably both too
strict in one sense (adjuncts are not properly
selected) and too permissive in another (certain
elements must agree to be merged). In the first case,
I assume that [l1 [l2]] can be formed even if S1 is
not =X but +X: while =X corresponds to
functional selection
        <xref ref-type="bibr" rid="ref18">(in compositional semantics terms
Heim &amp; Kratzer 1998)</xref>
        , +X corresponds to an
intersective compositional interpretation (e.g.
adjuncts and restrictive relative clauses). As for the
agreement constraint, I postulate an extra
(possibly parametrized) condition on MERGE, namely
the sharing (inclusion) of the relevant Agr features
associated to some specific categories.
      </p>
      <p>
        The auxiliary functions necessary to implement
Agreement are AGREE and UNIFY and can be
minimally defined as follows:
(16) AGREE(l1(L1), l2(L2)) =
{1    1 ∧  2 ∈  { } → 
0  ℎ
(17) UNIFY( 1(agr1),  2(agr2)) =
1,  , ∀ :  1∀ :  2  ∩ b   ⊆ 
{1,  , ∀ :  1 ∀ :  2  ∩ b   ⊆ 
0  ℎ
(  1,  2)}
}
Unification is simply expressed as an inclusion
relation returning true and the most specific feature
for any possible featural intersection between l1
and l2 Agr features6. Notice that Agreement is a
6 UNIFY(num, num.pl) = num.pl; UNIFY(, num.pl) =
num.pl; UNIFY(gen.f, num.pl) = gen.f, num.pl, since
gen and num are distinct agree subsets. On the other
hand, UNIFY([gen.f, num.sg], num.pl) would fail.
conditional, parametrized option, that is, it only
involves specific categories (possibly specified in
the parameter set P): if the L category belongs to
the Agreement set (Agr) in P for the grammar G,
unification will be attempted, otherwise
agreement will be trivially successful. The fact that
AGREE should apply in conjunction with MERGE
is straightforward in the D-N domain: in most
Romance languages, in which gender and number are
shared between the determiner and the noun, we
assume that D selects N
        <xref ref-type="bibr" rid="ref14">(this happens also for
intermediate functional specifications, according to
the cartographic intuition, Cinque 2002)</xref>
        . This is
less evident in the Subject – Predicate case, in SV
language, where the predicate should select (then
precede) D. Since the subject is clearly processed
(i.e. merged) before T, in canonical SV sentences,
and it does not select T, a re-merge operation
should be considered (e.g. case checking). This
remerge
        <xref ref-type="bibr" rid="ref12">(inducing the locality of Agree, pace
Chomsky 2001)</xref>
        is logically and empirically sound
        <xref ref-type="bibr" rid="ref1">(movement and agreement can be related and
parametrized, Alexiadou &amp; Anagnostopoulou 1998)</xref>
        .
In this case, re-merge must be preceded by MOVE,
an operation that stores in memory an item which
is “not fully” expected (i.e. there are exped2
features remaining) by the previous MERGE:
(18) MOVE(l1(M1), l2(L2)) =
{1,  ℎ( 1,  2( ℎ 2=∅))   2 ≠ ∅)}
0  ℎ
The definition of MOVE tells us that an item (l2)
must be moved (pushed7) into the memory buffer
(M1) of the superordinate item (l1) if it still has
expected features to be selected (L2 ≠ ). Notice that
item moved in M1 is not an exact copy of l2: the
used features (including Phon) will not be stored
in memory. This definition produces the expected
derivation if it applies right after MERGE, that is,
once the item l2 is properly (at least partially)
selected; in this case, if l2 still has exp(ect)ed
features to be licensed, it must hold in the memory
buffer of the selecting item, waiting for a proper
selection of what has become the new l2 label (i.e.
L2). (Re-)Merge is then when agreement will be
attempted (i.e. if MERGE(l1, l2) in §3, should then
be interpreted as if MERGE(l1, l2)  AGREE(l1, l2)
then… for specific parameterized categories). In
the end, the top-down derivation in SV languages
would unroll as follows: the subject (a DP) is first
7 PUSH and POP are trivial functions operating on
arrays: insert (PUSH) / remove (POP) an item to/from the
first available slot of a stack or a priority queue.
selected by a superordinate item (presuppositional
subject position, situation topic, focus etc.)8 then
it gets (partially) stored in the M buffer of the
selecting item in virtue of the unselected D features,
then re-merged as soon as a proper predicate,
expressing the relevant T category requiring
agreement (T should be included in the parameterized
Agreement), is merged and properly selects a D
argument (or it selects a V that later selects D).
The content of the memory buffer is transmitted
(inherited) through the last selected expectation,
namely when the expecting and the expectee
items successfully merge and the expecting item
has no more expectations (R1 ≠ ).
      </p>
      <p>
        If the expecting item has expectations, then the
expected item constitutes a nested expansion, and
the inheritance mechanism is blocked:
(19) INHERIT(l1(M1), l2(M2)) =
{1,  2 ⟸  1  MERGE( 1,  2) ∧  1 ≠ ∅)}
0  ℎ
The M buffer of the last selected item that does
not have other expectations (namely a right
phrasal edge, i.e., S=) must be empty (i.e.,
M=). If not, the derivation fails (i.e., it stops)
since a pending item remains unlicensed:
(20) SUCCESS(lx(Sx, Mx)) =
{1,    = ∅ →   = ∅)}
  ℎ
Notice that the sequential item must be properly
selected (=SX). If this is not the case, the
inheritance would transmit the content of the memory
buffer of the superordinate phase into the memory
buffer of an adjunct or a restrictive relative clause,
which clearly qualify as (right-branching) islands.
Therefore, the “restrictive” (since feature driven)
MERGE definition in (15) seems correct and
empirically more accurate than “free Merge”
        <xref ref-type="bibr" rid="ref13">(Chomsky, Gallego &amp; Ott 2019: 238)</xref>
        .
3
      </p>
      <sec id="sec-4-1">
        <title>The Derivation Algorithm</title>
        <p>We can now define the full-fledged top-down
derivation algorithm which is common both to
generation and to parsing tasks (§3.2). Consider cn to
be the current node, exp the list of pending
expectations and mem the ordered list of items in
memory. We initialize our procedure by picking
up an arbitrary node from G.Lex as cn. Being cn
the root node of our derivation(al tree) and w the
array of words we want to produce/recognize, we
can define the function DERIVE(cn, w) as follows:
while cn.exp &amp; w
while cn.mem
foreach cn.mem[i] in cn.mem
if MERGE(cn.exp[0], cn.mem[i])</p>
        <p>POP(cn.exp)</p>
        <p>POP(cn.mem)
else break
if MERGE(cn.exp[0], w[0])</p>
        <p>POP(cn.exp)
if w[0].exped</p>
        <p>MOVE(cn, w[0])
if w[0].exp
cn = w[0]
INHERIT(exp[0], w[0])</p>
        <p>SUCCESS(w[0])
POP(w)
if not cn.exp
while !cn.exp &amp; (cn != root)</p>
        <p>cn = cn.father
else fail</p>
        <p>
          Informally speaking, as long as we have lexical
items to consume (w), we loop into the set of
expectations of cn (cn.exp), first attempting to
Merge items from (cn.)mem (if any), as in the
active filler strategy
          <xref ref-type="bibr" rid="ref16">(Frazier &amp; Clifton 1989)</xref>
          , then
consuming words in the input (being w[0] the first
available word). Remember that each word has
exp(ect)ed features (the first being the label L),
exp(ectations) and agr(eement) features. Cns have
their own mem that can be inherited only by the
last expected item, and, apart from the root node,
a father. The derivation is then a depth-first,
leftright (i.e., real-time) strategy to derive a structure
given a grammar, a root node, and a sequence of
lexical items to be integrated.
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>The Complexity of Lexical Ambiguity</title>
      <p>
        Ignoring Parameters, the derivation procedure in
§3 should face lexical ambiguity: the same Phon
in w[n] might be associated to multiple items l in
Lex with different features; the default option is to
initialize a new derivational tree for any
ambiguous item in Lex. Given an ambiguity rate m in Lex,
the derivation procedure would have an
exponential order of complexity O(mn). We can mitigate
this, either by selecting the element(s) bringing
only coherent (i.e. expected) categories
        <xref ref-type="bibr" rid="ref3 ref30">(a
categorial priming strategy, Ziegler et al., 2019)</xref>
        or to use
a statistical oracle, following
        <xref ref-type="bibr" rid="ref29">Stabler (2013)</xref>
        , to
8 We have various options to implement this selection:
a specific feature (+focus, +topic, +presupposed etc.)
can be added to the relevant item (but this would lead
to a proliferation of lexical ambiguity, e.g. [D the …]
vs [FOC D the …]) or we assume that certain
superordinate items can select specific categories, without
deleting them (e.g. [+D ε FOC]). In this implementation, I
will pursue this second, more economic, alternative.
limit (or rank) the number of possible alternatives.
It is however important to stress that lexical
ambiguity is the major source of complexity in this
derivation: syntactic ambiguity is greatly
subsumed by the lexicon, being the source of
structural differences related to the set of categorial
expectations processed and to the order in which
lexical items are introduce in the derivation. With the
strict version of MERGE defined in (15), no
attachment ambiguity is allowed, since a matching
selection must be readily satisfied as soon as the
relevant configuration is created
        <xref ref-type="bibr" rid="ref7 ref9">(but see Chesi &amp;
Brattico 2018)</xref>
        . This is not the case if we would
admit “free merge” instead of
select/licensorsdriven merge: in the first case, admitting that
MERGE(l1(S1), l2(L2)) is possible also if S1 ≠ L2,
would produce a syntactic ambiguity which is
(exponentially) proportional to the number of items
merged in the structure. This is a crucial argument
to prefer feature-driven Merge. Notice, moreover,
that admitting that re-merge is also possible
without proper licensors/selectors, would quickly lead
to unbounded unstoppable recursion. This must be
prevented if we want to avoid the halting problem.
Therefore the licensors/selectors option seem to
be a more logical, self-contained, solution.
3.2
      </p>
    </sec>
    <sec id="sec-6">
      <title>Generation and Parsing</title>
      <p>As far as Generation is concerned, the procedure
described in §3 is integrally adopted and it is
sufficient to produce the expected sentence with the
associated, dependency-based, structural
description. As long as the sequence of words w is
concerned, once a root node is selected, it is easy to
imagine a dynamic function, instead of the static
ordered sequence w, that incrementally proposes
items to be integrated, given the history of the
derivation or, at least, the last expectation (a sort of
structural priming, possibly enriched with
semantic features if we add to the lexicon Sem(antic)
specifications in addition to Cat and Phon ones).</p>
      <p>Notice that the lexicon can include phonetically
empty categories; this is not a problem for the
generation procedure, that consumes input tokens
one by one, and then considers a phonetically
empty category on a par with phonetically
realized ones, namely each item should be postulated
as incoming token to be processed.</p>
      <p>
        From this perspective, the Parsing procedure is
minimally different since it must postulate a
phonetically empty item, for instance in pro-drop
languages, by deducting that the w sequence received
in input is incomplete/incompatible with specific
structural hypotheses. One proposal
        <xref ref-type="bibr" rid="ref3">(Brattico &amp;
Chesi 2020)</xref>
        relies on inflectional morphology as
an overt realization of unambiguous person and
number features cliticized on the predicate, hence
doubling the (null) subject. Otherwise, only after
a relevant category is selected (with its agreement
features) and unmatched by the current input, the
empty item could be postulated. This
non-determinism is exacerbated by the attachment/selection
ambiguity: given [l1 =/+X [l2 =/+X]], for instance, an
incoming item with X exp(ect)ed feature that
should be merged with l2 first, according to the
derivation algorithm provided in §3, could, in fact,
be merged also with l1, assuming that l2 =X
expectation can be satisfied with an empty item bearing
X as exp(ect)ed. Similarly, an adjunct marked with
Y exp(ect)ed category could be merged with both
l1 and l2 in [l1 [l2]] in case of lexical ambiguity ([l1],
[l1 +Y], [l2], [l2 +Y]). In this sense, the derivation
procedure in §3 is insufficient as a full-fledged
parsing strategy and must be integrated with
disambiguation routines dealing with the
possibilities just mentioned. It is however important to
stress that these disambiguation strategies do not
alter the general derivation procedure introduced
here, which remains the lowest common
denominator of Generation and Parsing in e-MGs.
4
      </p>
      <sec id="sec-6-1">
        <title>Conclusions</title>
        <p>
          The e-MGs formalization proposed here is a
simple (parametrized) framework for comparing
syntactic predictions directly with human parsing and
generation performance evidence. This is possible
since the core derivation algorithm is assumed to
be the same in both tasks
          <xref ref-type="bibr" rid="ref21">(token transparency,
Miller &amp; Chomsky 1963)</xref>
          . While there is little to
add to implement a full-fledged Generation
procedure (see §3.2), as long as the Parsing
perspective is concerned, the information asymmetry of
this task with respect to Generation requires extra
routines to be implemented, in addition to the
basic derivation algorithm: lexical ambiguity
must be resolved “on-line” and phonetically
empty items must be postulated when needed.
This creates an extra level of complexity which is
however manageable under the same derivational
perspective here presented: the core derivation is
sufficiently specified to operate independently
from parsing-specific disambiguation
assumptions which operate monotonically with respect to
MERGE, MOVE and AGREE. This is an ideal
foothold for metrics that aim at comparing the
predicted difficulty not only globally
          <xref ref-type="bibr" rid="ref15 ref17">(De Santo,
2020; Graf et al., 2017)</xref>
          but also “on-line” that is,
on a word by word basis
          <xref ref-type="bibr" rid="ref10 ref3 ref8">(Chesi &amp; Canal 2019;
Chesi 2021)</xref>
          .
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Implementation:</title>
      <p>https://github.com/cristianochesi/e-MGs</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Alexiadou</surname>
            ,
            <given-names>Artemis &amp; Elena</given-names>
          </string-name>
          <string-name>
            <surname>Anagnostopoulou</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <string-name>
            <surname>Parametrizing</surname>
            <given-names>AGR</given-names>
          </string-name>
          :
          <article-title>Word order, V-movement and EPP-checking</article-title>
          .
          <source>Natural Language &amp; Linguistic Theory. Springer</source>
          <volume>16</volume>
          (
          <issue>3</issue>
          ).
          <fpage>491</fpage>
          -
          <lpage>539</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Bianchi</surname>
            ,
            <given-names>Valentina &amp; Cristiano</given-names>
          </string-name>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Phases, left-branch islands, and computational nesting</article-title>
          .
          <source>Proceedings of the 29th Annual Penn Linguistics Colloquium</source>
          (University of Pennsylvania Working Papers in Linguistics)
          <volume>12</volume>
          .
          <fpage>1</fpage>
          .
          <fpage>15</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Brattico</surname>
            ,
            <given-names>Pauli &amp; Cristiano</given-names>
          </string-name>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>A top-down, parser-friendly approach to pied-piping and operator movement</article-title>
          .
          <source>Lingua. Elsevier</source>
          <volume>233</volume>
          (
          <issue>102760</issue>
          ).
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          . https://doi.org/10.1016/j.lingua.
          <year>2019</year>
          .
          <volume>102760</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          , Cristiano.
          <year>2005</year>
          .
          <article-title>Phases and Complexity in Phrase Structure Building</article-title>
          .
          <source>In Computational Linguistics in the Netherlands</source>
          <year>2004</year>
          :
          <article-title>Selected Papers of the 15th Meeting of Computational Linguistics in the Netherlands</article-title>
          ,
          <volume>59</volume>
          -
          <fpage>75</fpage>
          . UTRECHT: LOT. http://lotos.library.uu.nl/publish/issues/4/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          , Cristiano.
          <year>2007</year>
          .
          <article-title>An introduction to Phase-based Minimalist Grammars: why move is Top-Down from Left-to-Right</article-title>
          . In STIL - Studies in Linguistics - Vol.
          <volume>1</volume>
          , vol.
          <volume>1</volume>
          ,
          <fpage>38</fpage>
          -
          <lpage>75</lpage>
          . Siena: CISCL Press.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          , Cristiano.
          <year>2015</year>
          .
          <article-title>On directionality of phrase structure building</article-title>
          .
          <source>Journal of Psycholinguistic Research</source>
          <volume>65</volume>
          -89. https://doi.org/10.1007/s10936-014- 9330-6.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          , Cristiano.
          <year>2018</year>
          .
          <article-title>An efficient Trie for binding (and movement)</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), vol.
          <volume>2253</volume>
          . https://www.scopus.com/inward/record.uri?eid=
          <fpage>2</fpage>
          -
          <lpage>s2</lpage>
          .
          <fpage>0</fpage>
          -
          <lpage>85057729135</lpage>
          &amp;partnerID=
          <volume>40</volume>
          &amp;md5=
          <fpage>3c941a7524597857a24b64d671e</fpage>
          7239a.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          , Cristiano.
          <year>2021</year>
          .
          <article-title>Expectation-based Minimalist Grammars</article-title>
          . arXiv:
          <volume>2109</volume>
          .13871 [cs]. http://arxiv.org/abs/2109.13871 (
          <issue>2</issue>
          November,
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Chesi</surname>
          </string-name>
          ,
          <source>Cristiano &amp; PAULI JUHANI Brattico</source>
          .
          <year>2018</year>
          .
          <article-title>Larger than expected: constraints on pied-piping across languages</article-title>
          .
          <source>RGG. RIVISTA DI GRAMMATICA GENERATIVA</source>
          <year>2008</year>
          .
          <volume>4</volume>
          .
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Chesi</surname>
            ,
            <given-names>Cristiano &amp; Paolo</given-names>
          </string-name>
          <string-name>
            <surname>Canal</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Person Features and Lexical Restrictions in Italian Clefts. FRONTIERS IN PSYCHOLOGY</article-title>
          . https://doi.org/10.3389/fpsyg.
          <year>2019</year>
          .
          <volume>02105</volume>
          . https://www.frontiersin.org/articles/10.3389/fpsyg.
          <year>2019</year>
          .02105/full.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Chomsky</surname>
          </string-name>
          , Noam.
          <year>1995</year>
          .
          <article-title>The minimalist program</article-title>
          . Cambridge, MA: MIT press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Chomsky</surname>
          </string-name>
          , Noam.
          <year>2001</year>
          .
          <article-title>Derivation by phase</article-title>
          . In Michael Kenstowicz (ed.), Ken Hale:
          <article-title>A life in language</article-title>
          , 1-
          <fpage>52</fpage>
          . Cambridge (MA): MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Chomsky</surname>
            , Noam,
            <given-names>Ángel J Gallego &amp; Dennis</given-names>
          </string-name>
          <string-name>
            <surname>Ott</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Generative grammar and the faculty of language: Insights, questions, and challenges</article-title>
          .
          <source>Catalan Journal of Linguistics 229-261.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Cinque</surname>
          </string-name>
          , Guglielmo.
          <year>2002</year>
          .
          <article-title>Functional Structure in DP and IP: The Cartography of Syntactic Structures</article-title>
          , Volume
          <volume>1</volume>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>De Santo</surname>
          </string-name>
          , Aniello.
          <year>2020</year>
          .
          <article-title>MG Parsing as a Model of Gradient Acceptability in Syntactic Islands</article-title>
          .
          <source>In Proceedings of the Society for Computation in Linguistics</source>
          <year>2020</year>
          ,
          <fpage>59</fpage>
          -
          <lpage>69</lpage>
          . New York, New York: Association for Computational Linguistics. https://www.aclweb.org/anthology/2020.scil-
          <volume>1</volume>
          .7.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Frazier</surname>
            ,
            <given-names>Lyn &amp; Charles</given-names>
          </string-name>
          <string-name>
            <surname>Clifton</surname>
          </string-name>
          .
          <year>1989</year>
          .
          <article-title>Successive cyclicity in the grammar and the parser</article-title>
          .
          <source>Language and Cognitive Processes</source>
          <volume>4</volume>
          (
          <issue>2</issue>
          ).
          <fpage>93</fpage>
          -
          <lpage>126</lpage>
          . https://doi.org/10.1080/01690968908406359.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Graf</surname>
            , Thomas,
            <given-names>James Monette &amp; Chong</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Relative clauses as a benchmark for Minimalist parsing</article-title>
          .
          <source>Journal of Language Modelling</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ). https://doi.org/10.15398/jlm.v5i1.157. https://jlm.ipipan.waw.pl/index.php/JLM/article/view/157 (21 June,
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Heim</surname>
            ,
            <given-names>Irene &amp; Angelika</given-names>
          </string-name>
          <string-name>
            <surname>Kratzer</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Semantics in generative grammar (Blackwell Textbooks in Linguistics 13)</article-title>
          . Malden, MA: Blackwell.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>C.-T.</given-names>
          </string-name>
          <string-name>
            <surname>James</surname>
          </string-name>
          .
          <year>1982</year>
          .
          <article-title>Logical relations in Chinese and the theory of grammar</article-title>
          . Cambridge (MA): MIT.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Kayne</surname>
          </string-name>
          , Richard S.
          <year>1994</year>
          .
          <article-title>The antisymmetry of syntax (Linguistic Inquiry Monographs 25)</article-title>
          . Cambridge, Mass: MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>George A. &amp; Noam</given-names>
          </string-name>
          <string-name>
            <surname>Chomsky</surname>
          </string-name>
          .
          <year>1963</year>
          .
          <article-title>Finitary Models of Language Users</article-title>
          . In D. Luce (ed.),
          <source>Handbook of Mathematical Psychology</source>
          ,
          <volume>2</volume>
          -
          <fpage>419</fpage>
          . John Wiley &amp; Sons.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Momma</surname>
            ,
            <given-names>Shota &amp; Colin</given-names>
          </string-name>
          <string-name>
            <surname>Phillips</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The Relationship Between Parsing and Generation</article-title>
          .
          <source>Annual Review of Linguistics</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ).
          <fpage>233</fpage>
          -
          <lpage>254</lpage>
          . https://doi.org/10.1146/annurev-linguistics011817-
          <volume>045719</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Nivre</surname>
            , Joakim, Željko Agić, Lars Ahrenberg, Lene Antonsen, Maria Jesus Aranzabe, Masayuki Asahara,
            <given-names>Luma</given-names>
          </string-name>
          <string-name>
            <surname>Ateyah</surname>
          </string-name>
          , et al.
          <source>2017. Universal Dependencies 2.1.</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Partee</surname>
            ,
            <given-names>Barbara H.</given-names>
          </string-name>
          , Alice ter Meulen &amp;
          <string-name>
            <surname>Robert</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Wall</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Mathematical methods in linguistics (Studies in Linguistics and</article-title>
          Philosophy volume
          <volume>30</volume>
          ).
          <article-title>Corrected second printing of the first edition</article-title>
          . Dordrecht Boston London: Kluwer Academic Publishers.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Pollard</surname>
            ,
            <given-names>Carl</given-names>
          </string-name>
          <string-name>
            <surname>Jesse</surname>
          </string-name>
          &amp;
          <string-name>
            <surname>Ivan</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sag</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Head-driven phrase structure grammar (Studies in Contemporary Linguistics)</article-title>
          . Stanford : Chicago:
          <article-title>Center for the Study of Language</article-title>
          and Information ; University of Chicago Press.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Rizzi</surname>
          </string-name>
          , Luigi.
          <year>2016</year>
          .
          <article-title>Labeling, maximality and the headphrase distinction</article-title>
          .
          <source>The Linguistic Review. De Gruyter Mouton</source>
          <volume>33</volume>
          (
          <issue>1</issue>
          ).
          <fpage>103</fpage>
          -
          <lpage>127</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Stabler</surname>
          </string-name>
          , Edward.
          <year>1997</year>
          .
          <article-title>Derivational minimalism</article-title>
          . In Christian Retoré (ed.),
          <source>Logical Aspects of Computational Linguistics</source>
          ,
          <fpage>68</fpage>
          -
          <lpage>95</lpage>
          . Berlin, Heidelberg: Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Stabler</surname>
          </string-name>
          , Edward.
          <year>2011</year>
          .
          <article-title>Computational Perspectives on Minimalism</article-title>
          . In Cedric Boeckx (ed.),
          <source>The Oxford Handbook of Linguistic Minimalism</source>
          . Oxford University Press. https://doi.org/10.1093/oxfordhb/9780199549368.013.0027. http://oxfordhandbooks.com/view/10.1093/oxfordhb/9780199549368.001.0001/oxfordhb9780199549368-e-
          <volume>027</volume>
          (
          <issue>26</issue>
          <year>April</year>
          ,
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Stabler</surname>
          </string-name>
          , Edward.
          <year>2013</year>
          .
          <article-title>Two Models of Minimalist, Incremental Syntactic Analysis</article-title>
          .
          <source>Topics in Cognitive Science</source>
          <volume>5</volume>
          (
          <issue>3</issue>
          ).
          <fpage>611</fpage>
          -
          <lpage>633</lpage>
          . https://doi.org/10.1111/tops.12031.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Ziegler</surname>
          </string-name>
          , Jayden, Giulia Bencini, Adele Goldberg &amp; Jesse
          <string-name>
            <surname>Snedeker</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>How abstract is syntax? Evidence from structural priming</article-title>
          .
          <source>Cognition</source>
          <volume>193</volume>
          . 104045. https://doi.org/10.1016/j.cognition.
          <year>2019</year>
          .
          <volume>104045</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>