<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Multi-Strategy Approach to Crossword Clue Answer Retrieval and Ranking</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Zugarini</string-name>
          <email>azugarini@expert.ai</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Ernandes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>. Expert.ai</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. DIISM University of Siena</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Crossword clues represent an extremely challenging form of Question Answering, due to their intentional ambiguity. Databases of previously answered clues are a vital source for the retrieval of candidate answers lists in Automatic Crossword Puzzles (CPs) resolution systems. In this paper, we exploit language neural representations for the retrieval and ranking of crossword clues and answers. We assess the performances of several embedding models, both static and contextual, on Italian and English CPs. Results indicate that embeddings usually outperform the baseline. Moreover, the use of embeddings for retrieval allows different ranking strategies, which turned out to be complementary, and lead to better results when used in combination.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Le domande dei cruciverba
rappresentano una forma di Question
Answering particolarmente complessa a
causa della loro intenzionale ambiguita`. I
risolutori automatici di cruciverba
sfruttano ampiamente basi di dati di domande
precedentemente risposte. In questo
articolo proponiamo l’uso di embeddings per
la ricerca semantica di domande-risposte
da tali databases. Le performances sono
valutate in cruciverba di lingua sia
italiana che inglese, confrontando diversi tipi
di embeddings, sia contestuali che statici.
I risultati suggeriscono che la ricerca
semantica e` migliore della baseline. Inoltre,
l’utilizzo di embeddings permette di
applicare differenti strategie di retrieval, che,
Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
migliorano la qualita` dei risultati quando
usate congiuntamente.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        Crossword Puzzles (CPs) resolution is a popular
game. As almost any other human game, it is
possible to tackle the problem automatically. CPs
solvers frame it into a constraint satisfaction task,
where the goal is to maximize the probability of
iflling the grid with answers consistent with their
clues and coherent to the puzzle scheme. These
systems
        <xref ref-type="bibr" rid="ref12 ref7 ref9">(Littman et al., 2002; Ernandes et al.,
2005; Ginsberg, 2011)</xref>
        heavily rely on lists of
candidate answers for each clue. Candidates’ quality
is crucial to CPs resolution. If the correct answer
is not present in the candidates’ list, the Crossword
Puzzle cannot be solved correctly. Moreover, even
a poorly ranked correct answer can lead to a failure
in the crossword puzzle filling. Answers lists can
come from multiple solvers, where each solver is
typically specialized in solving different kinds of
clues, and/or exploits different source of
information. Such lists are mainly retrieved with two
techniques: (1) by querying the web with search
engines using clue representations; (2)
interrogating clue-answer databases that contain previously
answered clues. In this work, we focus on the
latter.
      </p>
      <p>
        In the problem of candidate answers retrieval
from clue-answer knowledge sources, answers are
ranked according to the similarity between a query
clue and the clues in the DB. The similarity is
provided by the search engine that assigns a score to
each retrieved answer. Several approaches have
been carried out to re-rank the candidates’ list by
means of learning to rank strategies
        <xref ref-type="bibr" rid="ref1 ref1 ref15 ref2 ref2 ref20 ref21 ref23 ref23">(Barlacchi et
al., 2014a; Barlacchi et al., 2014b; Nicosia et al.,
2015; Nicosia and Moschitti, 2016; Severyn et al.,
2015)</xref>
        . These approaches require a training phase
to learn how to rank and mostly differ for the
reranking model or strategy adopted. In particular,
pre-trained distributed representations and neural
networks are used for re-ranking clues in
        <xref ref-type="bibr" rid="ref23">(Severyn
et al., 2015)</xref>
        .
      </p>
      <p>The re-ranking of answer candidates attempts to
improve the quality of candidates’ lists, assuming
that the correct answer belongs to the list.
Differently from previous work, we aim at directly
retrieving richer lists of answer candidates from a
clue-answer database. In order to do so, we
exploit both static and contextual distributed
representations to perform a semantic search on the DB.
An embedding-based search extends the retrieval
to semantically related clues that may be phrased
differently. Moreover, it also allows us to map
in the same space questions and answers, which
opens the way for ranking answers directly based
on their similarity with respect to the query clue.
Our approach requires no training on CPs data and
it can be applied with any pre-trained embedding
model.</p>
      <p>In summary, the contributions of this work are:
(1) a semantic search approach to candidate
answer retrieval in automatic crossword resolution;
(2) two complementary retrieval methodologies
(namely QC and QA) detecting candidate answers
that when combined together (even naively)
produce a better set of candidates; (3) a comparison
between different pre-trained language
representations (either static or contextual).</p>
      <p>The paper is organized as follows. First, we
describe in Section 2 distributed representations of
language. In Section 3, we present the two answer
retrieval approaches proposed in this work. Then,
in Section 4 we outline the experiments in detail,
and discuss the obtained results. Finally, we draw
our conclusions in Section 5.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Language Representations</title>
      <p>
        Assigning meaningful representations to language
is a long standing problem. Since the inception
of the first text mining solutions, the bag-of-words
technique has been widely adopted as one of the
standard approaches to text representation.
Inverted indices and statistical weighting schemes
(as TF-IDF or BM25) are still to this day
commonly paired with bag-of-words, providing a
scalable and effective approach to document retrieval.
On the other hand, in the last decade, we have
assisted to tremendous progress in the field of
Natural Language Processing. Huge credit goes
to the diffusion of distributed representations of
words
        <xref ref-type="bibr" rid="ref17 ref17 ref18 ref18 ref19 ref3 ref5 ref6">(Bengio et al., 2003; Mikolov et al., 2013a;
Mikolov et al., 2013b; Collobert et al., 2011;
Mikolov et al., 2018; Devlin et al., 2018)</xref>
        learned
through Language Modeling related tasks on large
corpora.
      </p>
      <p>In general, the goal is to assign a fixed length
representation of size d, aka embedding, to a
textual passage s such that similar text passages
syntactically and/or semantically - are represented
closely in such space. An embedding model fe is
a function mapping s to a d-dimensional vector,
i.e: fe : s → Rd. Since language is a composition
of symbols (typically words), embedding models
ifrst tokenize text and then process such tokens in
order to compute the representation of such textual
passage.</p>
      <p>
        Nowadays, there are lots of embedding
models, and for some of them pre-trained
embeddings are available in a plethora of languages
        <xref ref-type="bibr" rid="ref10 ref19 ref25 ref26">(Yamada et al., 2020; Grave et al., 2018; Yang et
al., 2019)</xref>
        . Early methods like
        <xref ref-type="bibr" rid="ref17 ref18">(Mikolov et al.,
2013a)</xref>
        produce dense representations for single
tokens - mainly words - therefore further
processing is needed to obtain the actual representation of
s, when s is composed of multiple words. These
kinds of embeddings are also referred to as static
embeddings, since the representation of a token is
always the same regardless of the context in which
it appears. In
        <xref ref-type="bibr" rid="ref19">(Mikolov et al., 2018)</xref>
        , authors
extend
        <xref ref-type="bibr" rid="ref17 ref18">(Mikolov et al., 2013a)</xref>
        introducing n-gram
and sub-word information and in
        <xref ref-type="bibr" rid="ref1 ref11 ref2">(Le and Mikolov,
2014)</xref>
        , distributed representations are learned
directly for sentences and documents.
      </p>
      <p>
        Most of the proposed methods for contextual
embeddings were based on recurrent neural
language models
        <xref ref-type="bibr" rid="ref14 ref15 ref16 ref22 ref26 ref4">(Melamud et al., 2016; Yang et al.,
2019; Chidambaram et al., 2018; Mikolov et al.,
2010; Marra et al., 2018; Peters et al., 2018)</xref>
        ,
until the introduction of transformer architectures
        <xref ref-type="bibr" rid="ref13 ref24 ref6">(Vaswani et al., 2017; Devlin et al., 2018; Liu et
al., 2019)</xref>
        which are currently the state-of-the-art
models. In the next Section we will discuss how
such representations can be used to perform
semantic search. In the experiments, we will exploit
some of these embedding models - both static and
contextual.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Semantic Search</title>
      <p>Traditional CPs solvers rely on Similar Clue
Retrieval mechanisms. The idea is to find possible
[stan, totò, step,.... nasali]
ranking
[stan, ines, mike,.... nasali]
rankng
QC
Il nome di Laurel</p>
      <p>....</p>
      <p>Il comico Laurel
Le consonanti come la n
...</p>
      <p>Clue-Answer</p>
      <p>DB</p>
      <p>Query:
L'indimenticabile</p>
      <p>Laurel
&lt;Il nome di Laurel, stan&gt;</p>
      <p>....</p>
      <p>&lt;Il comico Laurel, stan&gt;
&lt;Le consonanti come la n, nasali&gt;
...</p>
      <p>stan
....
nasali
...</p>
      <p>QA
answers from clues in the database that are
similar to the given query. This is particularly
effective for crosswords, since the same clues tend to
be repeated over time, or may have little lexical
variations. Retrieval of similar clues is based on
search engines based on classical IR algorithms
such as TF-IDF or BM25, representing clues in the
database as documents to retrieve, given the target
clue as query.</p>
      <p>Here instead, we retrieve and rank documents
with semantic search. We propose two strategies,
namely QC and QA. QC is analogous to classical
similar clues retrieval systems, with the difference
that text is represented with a dense
representation. The approach retrieves and ranks from the
DB clues similar to the query and returns in
output the answers associated to those clues. QA,
instead, ranks the answers directly by computing the
cosine distance between the query and the answers
themselves. Intuitively, the latter approach ranks
well answers semantically correlated to the
question itself, particularly useful for clues about
synonyms. As we will show in Section 4, due to their
different nature, the list of candidates retrieved by
the two approaches are strongly complementary.
A sketch of the two approaches is outlined in Fig.
1. Let us describe them separately.
3.1</p>
      <sec id="sec-4-1">
        <title>Similar Clues Retrieval</title>
        <p>We are given a query clue which is a sequence of n
words q := (w1, . . . , wn), and a clue-answer DB
(C, A) constituted by M clue-answer pairs, where
C and A indicate the list of all the clues and
answers, respectively, while we denote a clue-answer
pair as: (c, a).</p>
        <p>We assign a fixed-length representation qe ∈ Rd
to the query clue q, computed with an embedding
model:
qe = fe(q).
(1)
For contextual embeddings fe is the model itself,
since they work directly on the sequence, whereas
for static embeddings we have to collapse n word
representations together into a single vector. For
simplicity, we simply average such embeddings.</p>
        <p>Analogously, each clue c ∈ C is encoded as in
Equation 1. Then, we measure the cosine
similarity between the query and each clue:
score(q, (c, a)) = cos(qe, fe(c)),
(2)
score(q, a) =
where cos(· , · ) denotes the cosine similarity. Thus,
we obtain a similarity score for each clue-answer
pair. In order to finally rank answers we average
all clue-answer pairs having the same answer:
1</p>
        <p>X score(q, (c, ak)), (3)
·
|A| ak∈A
where A indicates the set of clue-answer pairs
where the answer ak is equal to a. All the answers
in A are then ranked. Since we know a priori the
length of a query answer, candidates with incorrect
lengths are filtered out. We refer to this approach
as QC (Query-Clue).
3.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Similar Answers Retrieval</title>
        <p>Since we can map text into a fixed-length space,
we can also rank by measuring the similarity
between the query and the answer itself. The query
is encoded exactly as in Equation 1. In this case
however we only need the clue-answer DB to
retrieve the set of unique answers, denoted as A.
Similarly to Equation 2, we compute the cosine
similarity between query and answer embeddings:
score(q, a) = cos(qe, fe(a)),
(4)
for each a ∈ A, then we rank as in QC. We
call it QA (Query-Answer). It is important to
remark that QA is only feasible using latent
representations, traditional methods like TF-IDF are not
suited because of their sparsity of representations.
Moreover, QA is somewhat an orthogonal strategy
with respect to QC. We will see in Section 4, how
even a trivial ensemble of QA and QC is beneficial
to the performances.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>In the experiments we aim to prove the
effectiveness of semantic search to retrieve accurate lists
of candidate answers, and to show that the QA
approach carries out complementary information
that can increase the coverage of the retrieval.
4.1</p>
      <sec id="sec-5-1">
        <title>Experimental Setup</title>
        <p>We considered for our experiments three
well known embedding models, two static
(Word2Vec12, FastText3) and one contextual
(Universal Sentence Encoder4), briefly denoted
as W2V, FT and USE, respectively. We exploited
pre-trained models for all of them. In absence
of an Italian USE model, we used for the Italian
crosswords database, the multilingual version of
USE, that was trained on 16 languages (Italian
included). Embedding models are compared against
TF-IDF, which is a typical text representation in
document retrieval problems.</p>
        <p>To measure performances, we used well known
metrics of Retrieval systems. In particular we
considered Mean Hit at k (MH@k) and Mean
Reciprocal Rank (MRR). Hit at k is 1 if the
correct answer is within the first k elements of the
list, 0 otherwise. The hits at k are evaluated for
k = {1, 5, 20, 100}. MRR is defined as follows:
1 Pn 1
n q=1 rank(q) .
4.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Datasets</title>
        <p>
          We consider two different clue-answer databases
for our experimentation. In particular,
experiments were carried out on two languages, Italian
and English, respectively on CWDB dataset
          <xref ref-type="bibr" rid="ref1 ref2">(Barlacchi et al., 2014a)</xref>
          and New York Times
Crosswords. We apply the same pre-processing pipeline
in both corpora. (1) We discarded clue-answer
pairs having answers with more than three
characters, because they are typically about linguistic
puzzles and they are addressed differently in CPs
solvers. (2) Answer and clues containing special
characters are erased. (3) Text has been
lowercased and punctuation removed. (4) We kept only
answers appearing in at least two clues.
        </p>
        <p>1English: https://code.google.com/archiv
e/p/word2vec/</p>
        <p>2Italian: https://wikipedia2vec.github.io/
wikipedia2vec/
3https://fasttext.cc/
4https://tfhub.dev/google/collections
/universal-sentence-encoder/1
1.0
0.8
0.6
0.4
0.2
0.0</p>
        <p>USE</p>
        <p>
          TF-IDF
0
500
1000
1500
2000
2500
3000
English Crosswords. The data consist of a
collection of clue-answer pairs for crossword
puzzles published in the New York Times5 in 1997
and 2005, previously collected in
          <xref ref-type="bibr" rid="ref8">(Ernandes et
al., 2008)</xref>
          . Overall, there are about 61, 000
clueanswer pair samples. Clues, answers and
clueanswer pairs may occur multiple times. A clue
is generally a short sentence, while answers are
usually made up of a single word, but there are
cases of multi-word answers. In such a case the
answer is a string made of multiple words without
any word separator. After pre-processing we
obtain a corpus with 31, 808 pairs in which 27, 527
questions and 8, 324 answers are unique.
Italian Crosswords. The clue-answer database
for Italian was constructed from CWDB v0.1 it
corpus6
          <xref ref-type="bibr" rid="ref1 ref2">(Barlacchi et al., 2014a)</xref>
          . We combined
pairs from both train and test splits, since we did
not perform any training in our experiments and
we opportunely omitted the clue-answer pair itself
during its evaluation. From the original 62, 011
pairs, it remains 25, 545 pairs after pre-processing,
constituted of 5, 813 unique answers and 16, 970
unique questions.
4.3
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>Results</title>
        <p>All the results for Italian and English
crosswords are outlined in Tables 1 and 2,
respectively. From them, we can catch several
interesting insights. First of all, contextual
representations from Universal Sentence Encoders are
gen5https://www.nytimes.com/
6https://ikernels-portal.disi.unitn.i
t/projects/webcrow
MH@20
27.78
27.29
30.01
44.09
42.66
32.67
49.34
54.34
MH@100
42.62
43.42
45.17
49.54
57.38
46.64
63.35
69.00
MRR
12.66
12.75
14.25
31.46
25.65
20.20
32.12
EnsembleUSE− W 2V
erally the most effective ones, especially on
similar clues retrieval (QC), where both the query
and the elements to rank are textual sequences.
Nonetheless, Word2vec embeddings work
surprisingly well, outperforming FastText almost all the
times. Furthermore, they are the best ones on QA
search in Italian database. We believe the reason
why Word2Vec outperforms USE on Italian QA
is twofold. First, the advantage of contextual
embeddings is less evident in QA setup, indeed USE
brings less benefits on English QA as well.
Second, USE is a multilingual model, therefore its
embeddings are less specialized than Word2Vec
which was instead trained for Italian only.</p>
        <p>When comparing semantic search models
against the baseline (TF-IDF) - which is only
possible in QC - we can notice that, static
embeddings struggle to outperform it. Indeed, the sparse
nature of TF-IDF induces crisp similarity scores,
very high for clues sharing the same keywords,
extremely low for all the rest. On the contrary,
similarity scores are more blurred with dense
embeddings. As a consequence, TF-IDF achieves high
MH@1 and MH@5 scores (and MRR too).
However, TF-IDF leads to a poorer coverage when the
candidates list grows (MH@20 and MH@100).
This behavior is also evident in Fig. 2, where
we compare the cumulative distributions of
ranking with USE and TF-IDF. After the initial bump,
TF-IDF hits growth is almost linear (i.e. random),
whereas the Universal Sentence Encoder keeps
growing significantly.</p>
        <p>Ensembling QC and QA. Analyzing the
results, we observed that ranks from QA and QC
had low levels of overlaps. We reported in the last
line of Tables 1 and 2, performances of a naive
ensemble approach to combine QC and QA
strategies. Due to the limited levels of overlaps, we
decided to merge the two ranks taking the first K/2
ranks from each strategy to compute MH@K,
K = {5, 20, 100}7. We chose the best
embedding model on each strategy. Despite its
simplicity and the large room for improvements, the
ensemble significantly improved the performances in
both languages. This suggests possible directions
for further improving the retrieval of CPs solvers.</p>
        <p>7Since K=5 is not even, we took the first 3 ranks from QC
and the first two ranks from QA.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this paper, we proposed two different
semantic search strategies (QC and QA) for ranking and
retrieving answer candidates to CPs clues. We
exploited pre-trained state-of-the-art embeddings,
both static and contextual, to rank clue-answer
pairs from databases. Embedding-based retrieval
overcomes some of the limitations of inverted
indices models, leading to higher coverage ranks,
and allowing similar answers retrieval (QA).
Finally, we observed that, even a simple ensembling
that combines QC and QA, is effective and
improves overall retrieval performances.</p>
      <p>This opens further research directions, where
learning to rank methods could be exploited in
order to better combine candidate answer lists from
complementary approaches like QC and QA.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We thank Nicola Landolfi and Marco Maggini for
the great support and fruitful discussions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Gianni</given-names>
            <surname>Barlacchi</surname>
          </string-name>
          , Massimo Nicosia, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Moschitti</surname>
          </string-name>
          .
          <year>2014a</year>
          .
          <article-title>Learning to rank answer candidates for automatic resolution of crossword puzzles</article-title>
          .
          <source>In Proceedings of the Eighteenth Conference on Computational Natural Language Learning</source>
          , pages
          <fpage>39</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Gianni</given-names>
            <surname>Barlacchi</surname>
          </string-name>
          , Massimo Nicosia, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Moschitti</surname>
          </string-name>
          .
          <year>2014b</year>
          .
          <article-title>A retrieval model for automatic resolution of crossword puzzles in italian language</article-title>
          .
          <source>In The First Italian Conference on Computational Linguistics CLiC-it</source>
          <year>2014</year>
          , page 33.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          , Re´jean Ducharme, Pascal Vincent, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Jauvin</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>A neural probabilistic language model</article-title>
          .
          <source>Journal of machine learning research</source>
          ,
          <volume>3</volume>
          (Feb):
          <fpage>1137</fpage>
          -
          <lpage>1155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Muthuraman</given-names>
            <surname>Chidambaram</surname>
          </string-name>
          , Yinfei Yang, Daniel Cer, Steve Yuan,
          <string-name>
            <surname>Yun-Hsuan</surname>
            <given-names>Sung</given-names>
          </string-name>
          , Brian Strope, and
          <string-name>
            <given-names>Ray</given-names>
            <surname>Kurzweil</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Learning cross-lingual sentence representations via a multi-task dual-encoder model</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .12836.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Collobert</surname>
          </string-name>
          , Jason Weston, Le´on Bottou, Michael Karlen, Koray Kavukcuoglu, and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Kuksa</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Natural language processing (almost) from scratch</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>12</volume>
          (Aug):
          <fpage>2493</fpage>
          -
          <lpage>2537</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Ernandes</surname>
          </string-name>
          , Giovanni Angelini, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Gori</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Webcrow: A web-based system for crossword solving</article-title>
          .
          <source>In AAAI</source>
          , pages
          <fpage>1412</fpage>
          -
          <lpage>1417</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Ernandes</surname>
          </string-name>
          , Giovanni Angelini, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Gori</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>A web-based agent challenges human experts on crosswords</article-title>
          .
          <source>AI Magazine</source>
          ,
          <volume>29</volume>
          (
          <issue>1</issue>
          ):
          <fpage>77</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Matthew L Ginsberg</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Dr. fill: Crosswords and an implemented solver for singly weighted csps</article-title>
          .
          <source>Journal of Artificial Intelligence Research</source>
          ,
          <volume>42</volume>
          :
          <fpage>851</fpage>
          -
          <lpage>886</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Edouard</given-names>
            <surname>Grave</surname>
          </string-name>
          , Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Learning word vectors for 157 languages</article-title>
          .
          <source>In Proceedings of the International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Quoc</given-names>
            <surname>Le</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Distributed representations of sentences and documents</article-title>
          .
          <source>In International conference on machine learning</source>
          , pages
          <fpage>1188</fpage>
          -
          <lpage>1196</lpage>
          . PMLR.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Michael L Littman</surname>
          </string-name>
          ,
          <article-title>Greg A Keim,</article-title>
          and
          <string-name>
            <given-names>Noam</given-names>
            <surname>Shazeer</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>A probabilistic approach to solving crossword puzzles</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>134</volume>
          (
          <issue>1-2</issue>
          ):
          <fpage>23</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Yinhan</given-names>
            <surname>Liu</surname>
          </string-name>
          , Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Marra</surname>
          </string-name>
          , Andrea Zugarini, Stefano Melacci, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Maggini</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>An unsupervised character-aware neural approach to word and context representation learning</article-title>
          .
          <source>In International Conference on Artificial Neural Networks</source>
          , pages
          <fpage>126</fpage>
          -
          <lpage>136</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Oren</given-names>
            <surname>Melamud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Goldberger</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ido</given-names>
            <surname>Dagan</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>context2vec: Learning generic context embedding with bidirectional lstm</article-title>
          .
          <source>In Proceedings of the 20th SIGNLL conference on computational natural language learning</source>
          , pages
          <fpage>51</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <article-title>Toma´sˇ Mikolov, Martin Karafi a´t, Luka´sˇ Burget, Jan Cˇ ernocky`</article-title>
          , and
          <string-name>
            <given-names>Sanjeev</given-names>
            <surname>Khudanpur</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Recurrent neural network based language model</article-title>
          .
          <source>In Eleventh annual conference of the international speech communication association.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013a</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>arXiv preprint arXiv:1301</source>
          .
          <fpage>3781</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013b</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Edouard Grave, Piotr Bojanowski,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Puhrsch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Armand</given-names>
            <surname>Joulin</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Advances in pre-training distributed word representations</article-title>
          .
          <source>In Proceedings of the International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Nicosia</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Moschitti</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Crossword puzzle resolution in italian using distributional models for clue similarity</article-title>
          .
          <source>In IIR.</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Nicosia</surname>
          </string-name>
          , Gianni Barlacchi, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Moschitti</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learning to rank aggregated answers for crossword puzzles</article-title>
          .
          <source>In European Conference on Information Retrieval</source>
          , pages
          <fpage>556</fpage>
          -
          <lpage>561</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Matthew E.</given-names>
            <surname>Peters</surname>
          </string-name>
          , Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In Proc. of NAACL.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Aliaksei</given-names>
            <surname>Severyn</surname>
          </string-name>
          , Massimo Nicosia, Gianni Barlacchi, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Moschitti</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Distributional neural networks for automatic resolution of crossword puzzles</article-title>
          .
          <source>In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)</source>
          , pages
          <fpage>199</fpage>
          -
          <lpage>204</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
          <string-name>
            <surname>Łukasz Kaiser</surname>
            , and
            <given-names>Illia</given-names>
          </string-name>
          <string-name>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Ikuya</given-names>
            <surname>Yamada</surname>
          </string-name>
          , Akari Asai, Jin Sakuma, Hiroyuki Shindo, Hideaki Takeda, Yoshiyasu Takefuji, and
          <string-name>
            <given-names>Yuji</given-names>
            <surname>Matsumoto</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Wikipedia2Vec: An efifcient toolkit for learning and visualizing the embeddings of words and entities from Wikipedia</article-title>
          .
          <source>In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          , pages
          <fpage>23</fpage>
          -
          <lpage>30</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Yinfei</surname>
            <given-names>Yang</given-names>
          </string-name>
          , Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernandez Abrego, Steve Yuan, Chris Tar,
          <string-name>
            <surname>Yun-Hsuan Sung</surname>
          </string-name>
          , et al.
          <year>2019</year>
          .
          <article-title>Multilingual universal sentence encoder for semantic retrieval</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .04307.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>