<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Phonological Layers of Meaning: A Computational Exploration of Sound Iconicity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Gregor de Varda</string-name>
          <email>andreagregor.devarda@ studenti.unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlo Strapparava</string-name>
          <email>strappa@fbk.eu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Mind/Brain Sciences, University of Trento</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>FBK</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The present paper aims to investigate the nature and the extent of cross-linguistic phonosemantic correspondences within a computational framework. An LSTMbased Recurrent Neural Network is trained to associate the phonetic representation of a word, encoded as a sequence of feature vectors, to its corresponding semantic representation in a multilingual vector space. The processing network is tested, without further training, in a language that does not appear in the training set. The performance of the multilingual model is compared with a monolingual upper bound and a randomized baseline. After the quantitative evaluation of its performance, a qualitative analysis is carried out on the network's most effective predictions, showing an inhomogeneous distribution of phonosemantic information in the lexicon, influenced by semantic, syntactic, and pragmatic factors.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        The idea of a consistent relationship between
sound and meaning has held a particular
fascination over philosophers and linguists
        <xref ref-type="bibr" rid="ref21">(Plato, 1998)</xref>
        .
However, in recent times, this charming
hypothesis has progressively lost the interest of
scholars, especially in the post-Saussurean linguistic
tradition, which emphasized the arbitrariness in
such relation. The idea that sounds have
inherent meanings has recaptured its original
attractiveness in the field of cognitive sciences, where the
attention has initially focused on the link between
sound and shape. A prominent example of these
      </p>
      <p>
        Copyright c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
naturally biased mappings came from Ko¨hler’s
(1929) finding that, when asked to match two
novel shapes with the non-words ‘maluma’ and
‘takete’, English-speaking adults tended to label
as ‘maluma’ the curled shape, and as ‘takete’ the
sharp one. This germinal study paved the way to
several replications and expansions of its findings,
that reproduced Ko¨hler’s results in different
geocultural contexts
        <xref ref-type="bibr" rid="ref7">(Bremner et al., 2013)</xref>
        and at
different developmental stages
        <xref ref-type="bibr" rid="ref17">(Maurer et al., 2006)</xref>
        .
Since then, different studies have tackled the topic
of iconicity in language from a broader
perspective, showing that adults can associate visually
presented characters
        <xref ref-type="bibr" rid="ref16">(Koriat and Levy, 1977)</xref>
        and
auditorily presented words
        <xref ref-type="bibr" rid="ref2">(Berlin, 1995)</xref>
        of a
foreign language to their meaning, with an accuracy
above chance.
      </p>
      <p>
        Recently, linguistic iconicity has gone from
being a marginal – although appealing – matter to
being integrated into broader theories of language
evolution and acquisition. Indeed, rejecting the
assumption of an arbitrary mapping between sound
and meaning sensibly reduces the problem space
of language emergence, establishing constraints
on the consensus of word choice. Furthermore,
an iconic relation between a sound and its
referent might help with memory consolidation in
the process of language acquisition
        <xref ref-type="bibr" rid="ref24">(Sathian and
Ramachandran, 2019)</xref>
        . Ramachandran and
Hubbard (2001) speculate that phenomena as the one
reported by Ko¨hler might arise from neural
connections among adjacent cortical areas, where the
visual features of the referent, the appearance of
the speaker’s lips and the kinaesthetic features
of the articulation are combined. According to
their view, such neural connections would have
influenced both the phylogenetic evolution and
the ontogenetic development of language.
Although the previous findings are consistent with
this hypothesis, an alternative explanation must be
taken into account: the roots of these
correspondences could be grounded in the knowledge of
language, that allows children and adults to
generalize the regularities in sound-to-meaning
mappings from their native language to nonsense and
foreign words. Under this rationale,
phonosemantic relations would be implicitly learned from
general recurrences in already known languages. A
crucial aspect of this account lies in the fact that
it does not posit any preexisting disposition wired
in the human brain, moving the locus of
linguistic iconicity from the mind to language itself. A
natural question that arises from this perspective
is whether linguistic information alone is
sufficient to give rise to the phonaesthetic biases
presented in the literature. A computational
exploration of the phenomenon under scrutiny is a
feasible way to approach the subject. The idea that
phones have inherent meanings is relatively
understudied within the computational framework, and
most of the studies addressing the topic have
either focused on a single language
        <xref ref-type="bibr" rid="ref1 ref13 ref19 ref23 ref25">(Gutie´rrez et al.,
2016; Sagi and Otis, 2008; Abramova et al., 2013;
Monaghan et al., 2014; Tamariz, 2008)</xref>
        or on a
small set of concepts on a massively multilingual
scale
        <xref ref-type="bibr" rid="ref27 ref4">(Blasi et al., 2016; Wichmann et al., 2010)</xref>
        .
Surprisingly, no study to our knowledge has
tackled the topic through a deep learning methodology,
and no cross-linguistic investigation has been
performed on a lexicon-wide level. The purpose of
the present study is two-fold: first, we wish to
explore the idea of a cross-linguistic correspondence
between the phonetic and the semantic
representation of a word on the whole lexicon, without any
theory-driven restriction guiding our choice of the
lexical items. Then, we aim to examine whether
the meaning that is rooted in the sound that words
are made of is homogeneously distributed in the
lexicon. Ultimately, these two goals converge
toward the research question hinted above, namely,
whether linguistic information alone could
suffice for the extrapolation of the phonosemantic
biases reported in the present section. A possible
way to answer this question is to assess the
ability of a tabula rasa neural network to extend the
regularities captured in a set of given languages
to a previously unseen one. Although equipped
with clear structural priors, neural networks do
not conceal biases that resemble those assumed
to model the aforementioned phonosemantic
correspondences. If a processing network showed
the ability to induce cross-linguistic regularities in
sound-to-meaning mappings, this would suggest
that linguistic data contain a sufficient amount of
information to encode for phonosymbolic biases.
      </p>
      <p>The present study aims to explore the possibility
of a certain degree of cross-linguistic
correspondence between sound and meaning that is already
encoded in language. A Long Short-Term
Memory network (LSTM) is trained on four languages
to associate the sequence of sounds that compose
a word, encoded as phonetic vectors, to its
corresponding semantic representation in a
multilingual vector space. Then the processing network is
tested, without further training, on a language that
does not appear in the set of languages on which
the training has been performed. The performance
of the multilingual model is compared with the
results of (a) a monolingual model, trained and
tested on different subsets of a single language’s
vocabulary, and (b) a baseline model, where the
output vectors in the training are randomly
shuffled. After the quantitative evaluation of its
performance, a qualitative analysis is carried out on
the network’s most effective predictions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>In the present study, an LSTM-based Recurrent
Neural Network is trained to associate the
phonetic to the corresponding semantic representation
of a word. The semantic representations consist
in 300-dimensional word embeddings in a
multilingual vector space, whereas their
corresponding phonetic features are expressed as sequences
of phonetic vectors in 22 dimensions. The
experimental pipeline is summarized in the flowchart in
Figure 1.</p>
      <sec id="sec-2-1">
        <title>2.1 Semantic vectors</title>
        <p>
          The semantic representations included in the
model, provided by Facebook Research, consist
in multilingual word embeddings generated with
fastText from Wikipedia data
          <xref ref-type="bibr" rid="ref6">(Bojanowski
et al., 2017)</xref>
          and aligned in a common vector space
through a fully unsupervised methodology
          <xref ref-type="bibr" rid="ref9">(Conneau et al., 2017)</xref>
          1. The present study is
conducted on Italian, German, French, Vietnamese,
and Turkish embeddings.
        </p>
        <p>
          1Publicly available at https://github.com/
facebookresearch/MUSE
For each word in the embedding dataset, we
obtained its phonemic transcription with Epitran,
a Python library for transliterating orthographic
text in the International Phonetic Alphabet (IPA)
format. Then, we converted the IPA string into
a sequence of feature vectors in 22 dimensions
with PanPhon, a package that traduces IPA
segments into subsegmental articulatory features
          <xref ref-type="bibr" rid="ref20 ref3">(Mortensen et al., 2016)</xref>
          . It has been shown that
phonologically aware models built on the
linguistically motivated and information-rich
representations yielded by the Epitran-PanPhon pipeline
outperform the raw hot-encoding of character-based
models in different tasks
          <xref ref-type="bibr" rid="ref20 ref20 ref3 ref3">(Mortensen et al., 2016;
Bharadwaj et al., 2016)</xref>
          .
An LSTM-based Recurrent Neural Network is
trained to map the sequences of phonetic feature
vectors in input into semantic vectors in output.
The model is built with Keras, a deep learning
framework for Python
          <xref ref-type="bibr" rid="ref8">(Chollet et al., 2015)</xref>
          ; it
includes a single LSTM layer with 172 units, a
dropout of 0.2 and a recurrent dropout of 0.2.
Cosine similarity is used as both objective function
and metric, and the Adam optimization method is
employed for training
          <xref ref-type="bibr" rid="ref14 ref19">(Kingma and Ba, 2014)</xref>
          . We
adopted the tanh activation function for the output
layer since its codomain corresponds to the range
(-1, 1), in which the semantic vectors are defined.
The hyperparameters are set without tuning.
2.4
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Experimental conditions</title>
        <p>The experimental conditions are characterized by
different combinations of training and testing sets.
In the multilingual condition, the model is trained
for one epoch on the Italian, German, French, and
Vietnamese datasets, and then tested in Turkish.
Our unique concern in the language selection was
that none of the languages in the training set was
typologically close with the language presented in
the test set. Turkish has been chosen for the test
set since it is not considered to be related to any of
the languages presented in the training set, at least
within a reasonable time window. Indeed, Turkish
is a Turkic language, whereas Italian, German and
French are Indoeuropean, and Vietnamese belongs
to the Austroasiatic language family. To establish
a baseline for the evaluation of the model’s
performance, we trained a model randomly shuffling
the output vectors. We will refer to this
manipulation as the random condition. In the monolingual
condition, which defines the upper bound of the
network’s performance, the LSTM is trained and
tested on different subsets of the Italian dataset,
with a train-test split ratio of 0.2. In order to
compensate for the different dimensions of the training
set (roughly one fifth of the multilingual sample),
the monolingual model is trained for five epochs.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>Table 1 lists the test results for each of the models
described in Section 2.4. The number of lexical
items included in the training and in the test set are
reported in the Dimtrain and the Dimtest columns,
respectively. The last column of the table presents
the average cosine similarity between the target
semantic vector and the model’s prediction for every
word in the test set. As reported below, the
multilingual model outperforms the random baseline,
with a 0.0351 points higher average cosine
similarity. As expected, the monolingual performance
is stronger than the one achieved by the
multilingual model, with a difference of 0.0453 in the
metric. The relatively modest magnitude of the
difference between the monolingual and the
multilingual results should be attributed to the limited
size of the training set in the former condition:
increasing the number of epochs might have
partially compensated for the shortage in the training
data, but additional forward and back propagation
on the same data might not be as effective as
further training on unseen data, especially in terms
of generalization. The general pattern of results,
with the multilingual performance almost halfway
between the monolingual and the random results,
is in line with our predictions. The difference
between the multilingual and the random condition is
consistent with the hypothesis that a certain degree
of cross-linguistic correspondence between
phonetic and semantic representations is already
encoded in language; moreover, it shows that, with
sufficient training, this correspondence can be
efficiently captured by an LSTM network.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Qualitative analysis</title>
      <p>As previously mentioned, the LSTM network
trained on multilingual data showed the ability
to induce cross-linguistic regularities in
sound-tomeaning mappings, suggesting that linguistic data
Letter string
‘stagione’
fastText</p>
      <p>Phonemic
transcription</p>
      <p>/stad&gt;Zone/
Epitran
PanPhon</p>
      <p>Phonetic vector
sequence
[-1, -1, : : : 0, -1],
[-1, -1, : : : 0, -1],</p>
      <p>LSTM
alone contain the sufficient amount of information
to encode for phonosymbolic biases.</p>
      <p>
        In the light of the results presented above, a
natural question that arises is whether
phonosemantic information is uniformly distributed in the
lexicon, or some semantic areas tend to
incorporate stronger correspondences with their phonetic
counterparts. We hypothesized that some areas
of the semantic space might show a more
consistent mapping with their phonological
realization, but without any clear a priori expectation
on the regions that could reveal higher
phonosemantic transparency, we addressed this problem
through a data-driven qualitative analysis. We
extracted from the test results of the multilingual
condition 39,821 items (20% of the total),
selecting the words that the network had predicted with
the higher precision – that is, the words whose
vector prediction had the higher cosine
similarity with respect to the target. Then, we restricted
the analysis by excluding the items with low
frequency. We conjectured that it would be unlikely
for rare and unfamiliar terms to convey
phonosemantic relations without being etymologically
related to other languages. For instance, across
different disciplines, the technical jargon – whose
instances are typically infrequent in corpora – is
commonly derived from Greek and Latin roots.
We employed the Twitter-based Turkish frequency
estimates from the Worldlex dataset2, that has
been shown to outperform traditional frequency
estimates in predicting lexical decision reaction
times, thus exhibiting a higher cognitive validity
        <xref ref-type="bibr" rid="ref12 ref3">(Gimenes and New, 2016)</xref>
        . From the previously
extracted items, we excluded those that were not
in the list of the 20,000 most frequent words (that
is, the 1.21% of the words with higher frequency).
The resulting items were translated into English
with Googletrans, a Python library that
implements Google Translate API. The results of the
analysis are reported in Table 2, where the items
that satisfy the aforementioned constraints (from
now on, the quality subset) are grouped into four
intuitive categories according to their meaning and
their grammatical function.
      </p>
      <p>
        The most represented categories of words in
this subset of efficiently predicted items are proper
names and lexical borrowings, with the former
generally associated with a higher cosine
similarity between target and prediction. They are not
reported in Table 2, since their detailed analysis is
not relevant for the purposes of the study.
However, the predominance of proper names over
lexical borrowings is compatible with one of the basic
postulates of model-theoretic semantics. It is
generally assumed that proper names, unlike definite
descriptions and generalized quantifiers, directly
refer to entities in the world
        <xref ref-type="bibr" rid="ref10">(Delfitto and
Zamparelli, 2009)</xref>
        ; hence, they are expected to hold their
exact meaning across languages.
      </p>
      <p>The cross-linguistic consistencies in proper
names and lexical borrowings are clearly due to
contact between languages. Other word categories
strongly associated with the phonosemantic
feaavailable
at
http://worldlex.</p>
      <p>
        2Publicly
lexique.org
Internal states istedig˘imde (‘I want’),
du¨s¸u¨ncelerimi (‘my thoughts’),
isteyenlere (‘those who want’),
du¨s¸u¨nsenize (‘imagine’), du¨s¸u¨ncem
(‘I thought’), ac¸ıkc¸ası (‘frankly’),
as¸ıksın (‘you are in love’), kendimde
(‘in myself’)
Function words vee (‘and’), kendileri (‘themselves’),
onların (‘they’), gerektig˘inde
(‘when’), mıydın (‘did you’)
Interjections hee (‘ooh’), boku (‘shit’), himm
(‘uhm’)
Other yaklas¸ım (‘approach’), pog˘ac¸a
(‘pastry’), demis (‘said’), gerc¸ekmis¸
(‘real’), tabiiki (‘of course’),
uygulamaları (‘applications’), gani
(‘abundant’)
tures detected by the network are undoubtedly
more relevant in revealing lexical clusters with
privileged sound-to-meaning mappings. For
instance, a conspicuous portion of items in the
quality subset is semantically linked to different
internal states, with a predominance of concepts
related to mental processes. The quality subset
comprises also various function words (conjunctions,
pronouns, and one auxiliary verb). This result is
particularly informative since function words,
being a closed-class category, are not as numerous
as content words; therefore, their number of
instances in the training set was most likely
limited. An additional cluster in the quality set
comprises three interjections, including one
imprecation. Interjections express spontaneous feelings or
reactions
        <xref ref-type="bibr" rid="ref5">(Bloomfield, 1984)</xref>
        and can be closely
related to their natural manifestation
        <xref ref-type="bibr" rid="ref26">(Wharton,
2003)</xref>
        ; hence, it is not surprising to find a more
transparent link between their phonoarticulatory
expression and their meaning. Moreover, this
result is consistent with the findings of
Dingemanse et al. (2013), that show that the interjection
“Huh?” is a universal, found in roughly the same
form and function in spoken languages across the
globe.
      </p>
      <p>The present findings suggest that
phonosemantic information is not uniformly distributed in the
lexicon: the consistency of the mapping between
sound and meaning seems to be influenced by
semantic, syntactic, and pragmatic factors.
Indeed, the semantic neighbourhood linked to
internal states shows a privileged relationship between
sound and meaning, whereas on the syntactic side
function words seem to be favoured, if their
absolute prevalence in the lexicon is taken into account.
Moreover, interjections, which are characterized
by a strong pragmatical valence, stand among the
items predicted with the highest precision by the
model.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Limitations and further directions</title>
      <p>From a methodological standpoint, the reliability
of the present results could benefit from the
exclusion of lexical borrowings and proper names from
the training and the test sets. Excluding
etymologically related terms could further improve the
reliability of the results, but at the costs of raising
the difficulty of assessing the words’ relatedness
in different languages, with the subsequent need
of a proper metric.</p>
      <p>
        Another confound that we wish to address in
future research is the role played by
morphological factors in aiding the cross-linguistic feature
extraction performed by the network. FastText
vectors exploit information related to subword
character strings, and might therefore encode
regularities pertaining to recurrent morphemes in the
nonisolating languages in our dataset (Italian,
German, French, and Turkish). We acknowledge that
the network might have captured the recurrences
encoded in the semantic vectors comprising the
training set and their relationships with the
corresponding phonetic feature vectors; indeed, we
believe that this regularities might have played a
relevant role in the monolingual condition, where
the model might have learnt that morphologically
related words (i.e. in this context, words that are
similar at the character- and phoneme-level) tend
to be associated with close subregions of the
semantic space. Nonetheless, we do not see how this
information could have altered significantly the
performance in the multilingual condition. That
said, we leave for future research an assessment of
the algorithm’s performance on semantic vectors
which lack access to subword-related information,
such as word2vec
        <xref ref-type="bibr" rid="ref18">(Mikolov et al., 2013)</xref>
        , and
in languages with opaque orthography (e.g.
English and French) and non-concatenative
morphology (e.g. Chinese)3.
      </p>
      <p>3We gratefully thank an anonymous reviewer for drawing
our attention to this matter and suggesting the mentioned
options to address this confound.</p>
      <p>As for all the studies that employ artificial
neural networks to draw conclusions on human
cognition, it is mandatory to clarify some limitations
on the extent of the inferences that can
legitimately follow the presented results. The finding
that a neural network can succeed in a task
without the structural priors postulated in the human
mind does not necessarily imply that these
priors are not actually encoded in the brain: the
assumption of a functional equivalence between
artificial and biological processes needs to be
independently motivated. Moreover, it should be
noticed that the participants of the behavioural
studies presented in Section 1 were not necessarily
polyglots, whereas the promising cross-linguistic
performances described in the results have been
obtained with a multilingual model. In addition
to these intrinsic methodological limitations, an
account that does not assume any prior
specification for the linguistically encoded phonosymbolic
mappings would leave an open question
concerning their origin. Hence, the present study does not
claim to reject the multi-sensory integration
hypothesis presented in the Introduction. Its purpose
is simply to show that, in principle, linguistic
information alone could suffice for a generalization
in sound-to-meaning mappings.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Ekaterina</given-names>
            <surname>Abramova</surname>
          </string-name>
          , Raquel Ferna´ndez, and
          <string-name>
            <given-names>Federico</given-names>
            <surname>Sangati</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Automatic labeling of phonesthemic senses</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Brent</given-names>
            <surname>Berlin</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Evidence for pervasive synesthetic sound symbolism in ethnozoological nomenclature</article-title>
          , page
          <volume>76</volume>
          -
          <fpage>93</fpage>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Akash</given-names>
            <surname>Bharadwaj</surname>
          </string-name>
          , David Mortensen,
          <string-name>
            <given-names>Chris</given-names>
            <surname>Dyer</surname>
          </string-name>
          , and Jaime Carbonell.
          <year>2016</year>
          .
          <article-title>Phonologically aware neural model for named entity recognition in low resource transfer settings</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>1462</fpage>
          -
          <lpage>1472</lpage>
          , Austin, Texas. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Damia´n</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Blasi</surname>
          </string-name>
          , Søren Wichmann, Harald Hammarstro¨m,
          <string-name>
            <surname>Peter F. Stadler</surname>
          </string-name>
          , and
          <string-name>
            <surname>Morten</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Christiansen</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Sound-meaning association biases evidenced across thousands of languages</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          ,
          <volume>113</volume>
          (
          <issue>39</issue>
          ):
          <fpage>10818</fpage>
          -
          <lpage>10823</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Leonard</given-names>
            <surname>Bloomfield</surname>
          </string-name>
          .
          <year>1984</year>
          . Language. University of Chicago Press.
          <article-title>Google-Books-ID: 87BCDVsmFE4C.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Andrew J. Bremner</surname>
          </string-name>
          , Serge Caparos, Jules Davidoff, Jan de Fockert, Karina J.
          <string-name>
            <surname>Linnell</surname>
            , and
            <given-names>Charles</given-names>
          </string-name>
          <string-name>
            <surname>Spence</surname>
          </string-name>
          .
          <year>2013</year>
          . “
          <article-title>Bouba” and “Kiki” in Namibia? A remote culture make similar shape-sound matches, but different shape-taste matches to Westerners</article-title>
          . Cognition,
          <volume>126</volume>
          (
          <issue>2</issue>
          ):
          <fpage>165</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Franc¸ois Chollet</surname>
          </string-name>
          et al.
          <year>2015</year>
          . Keras. https:// keras.io.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Conneau</surname>
          </string-name>
          , Guillaume Lample,
          <string-name>
            <surname>Marc'Aurelio Ranzato</surname>
          </string-name>
          , Ludovic Denoyer, and Herve´ Je´gou.
          <year>2017</year>
          .
          <article-title>Word translation without parallel data</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Denis</given-names>
            <surname>Delfitto</surname>
          </string-name>
          and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Zamparelli</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Le strutture del significato</article-title>
          .
          <source>Itinerari Linguistica. Mulino</source>
          , Bologna. OCLC:
          <volume>695640183</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Mark</given-names>
            <surname>Dingemanse</surname>
          </string-name>
          , Francisco Torreira, and
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Enfield</surname>
          </string-name>
          .
          <year>2013</year>
          . Is “Huh?
          <article-title>” a Universal Word? Conversational Infrastructure and the Convergent Evolution of Linguistic Items</article-title>
          .
          <source>PLoS ONE</source>
          ,
          <volume>8</volume>
          (
          <issue>11</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Manuel</given-names>
            <surname>Gimenes</surname>
          </string-name>
          and Boris New.
          <year>2016</year>
          .
          <article-title>Worldlex: Twitter and blog word frequencies for 66 languages</article-title>
          .
          <source>Behavior Research Methods</source>
          ,
          <volume>48</volume>
          (
          <issue>3</issue>
          ):
          <fpage>963</fpage>
          -
          <lpage>972</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Dario</surname>
          </string-name>
          <article-title>Gutie´rrez, Roger Levy</article-title>
          , and
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bergen</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Finding non-arbitrary form-meaning systematicity using string-metric learning for kernel regression</article-title>
          .
          <source>In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>2379</fpage>
          -
          <lpage>2388</lpage>
          , Berlin, Germany. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Diederik P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jimmy</given-names>
            <surname>Ba</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Wolfgang</surname>
            <given-names>Ko¨hler. 1929. Gestalt</given-names>
          </string-name>
          <string-name>
            <surname>Psychology</surname>
          </string-name>
          . Liveright.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Asher</given-names>
            <surname>Koriat</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ilia</given-names>
            <surname>Levy</surname>
          </string-name>
          .
          <year>1977</year>
          .
          <article-title>The symbolic implications of vowels and of their orthographic representations in two natural languages</article-title>
          .
          <source>Journal of Psycholinguistic Research</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <fpage>93</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Daphne</given-names>
            <surname>Maurer</surname>
          </string-name>
          , Thanujeni Pathman, and
          <string-name>
            <given-names>Catherine J.</given-names>
            <surname>Mondloch</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The shape of boubas: sound-shape correspondences in toddlers and adults</article-title>
          .
          <source>Developmental Science</source>
          ,
          <volume>9</volume>
          (
          <issue>3</issue>
          ):
          <fpage>316</fpage>
          -
          <lpage>322</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Padraic</surname>
            <given-names>Monaghan</given-names>
          </string-name>
          , Richard Shillcock, Morten Christiansen, and
          <string-name>
            <given-names>Simon</given-names>
            <surname>Kirby</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>How arbitrary is language? Philosophical transactions of the Royal Society</article-title>
          of London. Series B,
          <string-name>
            <surname>Biological</surname>
            <given-names>sciences</given-names>
          </string-name>
          ,
          <volume>369</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>David R.</given-names>
            <surname>Mortensen</surname>
          </string-name>
          , Patrick Littell, Akash Bharadwaj, Kartik Goyal, Chris Dyer, and
          <string-name>
            <given-names>Lori</given-names>
            <surname>Levin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>PanPhon: A resource for mapping IPA segments to articulatory feature vectors</article-title>
          .
          <source>In Proceedings of COLING</source>
          <year>2016</year>
          ,
          <source>the 26th International Conference on Computational Linguistics: Technical Papers</source>
          , pages
          <fpage>3475</fpage>
          -
          <lpage>3484</lpage>
          , Osaka, Japan.
          <source>The COLING 2016 Organizing Committee.</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Plato. 1998.</given-names>
            <surname>Cratylus</surname>
          </string-name>
          . Hackett Publishing Company.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>V. S.</given-names>
            <surname>Ramachandran</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Hubbard</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Synaesthesia-a window into perception, thought and language</article-title>
          .
          <source>Journal of Consciousness Studies</source>
          ,
          <volume>8</volume>
          (
          <issue>12</issue>
          ):
          <fpage>3</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Eyal</given-names>
            <surname>Sagi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Katya</given-names>
            <surname>Otis</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Semantic glimmers: Phonaesthemes facilitate access to sentence meaning.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Sathian</surname>
          </string-name>
          and
          <string-name>
            <given-names>V. S.</given-names>
            <surname>Ramachandran</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Multisensory perception: from laboratory to clinic</article-title>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamariz</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Exploring systematicity between phonological and context-cooccurrence representations of the mental lexicon</article-title>
          .
          <source>The Mental Lexicon</source>
          ,
          <volume>3</volume>
          :
          <fpage>259</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Tim</given-names>
            <surname>Wharton</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Interjections, language, and the 'showing/saying' continuum</article-title>
          .
          <source>Pragmatics &amp; Cognition</source>
          ,
          <volume>11</volume>
          :
          <fpage>39</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Søren</given-names>
            <surname>Wichmann</surname>
          </string-name>
          , Eric Holman, and
          <string-name>
            <given-names>Cecil</given-names>
            <surname>Brown</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Sound symbolism in basic vocabulary</article-title>
          .
          <source>Entropy</source>
          ,
          <volume>12</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>