<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Gender Bias in French Literature</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>LaurineVianne</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yoann Dupont</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jean Barré</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>École nationale supérieure - Université PSL</institution>
          ,
          <addr-line>45 rue d'Ulm, Paris, 75005</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lattice (Langues</institution>
          ,
          <addr-line>Textes, Traitements informatiques, Cognition), 1 rue Maurice Arnoux, Montrouge, 92049</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Université Paris Cité</institution>
          ,
          <addr-line>85 boulevard Saint-Germain, Paris, 75006</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Université Sorbonne Nouvelle</institution>
          ,
          <addr-line>17 rue de la Sorbonne, Paris, 75230</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <fpage>247</fpage>
      <lpage>262</lpage>
      <abstract>
        <p>This study delves into the representation of gender in French literature from 1800 to the present, aiming to assess the prevalence of gender stereotypes in the description of 昀椀ctional characters. By employing an annotated corpus and statistical modeling techniques, the research explores how authors perpetuated gender biases while shaping characters and narratives. The 昀椀ndings reveal signi昀椀cant linguistic patterns that reinforce gender norms, with women being characterized by emotional and physical attributes, while men are associated with action and agency.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;french literature</kwd>
        <kwd>bias</kwd>
        <kwd>gender</kwd>
        <kwd>automatic classi昀椀cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Consequently, this research employs the masculine/feminine dichotomy to capture prevalent
gender constructs in French literature.</p>
      <p>Gender bias in 昀椀ctional characterization refers to the systemic tendency to portray female
characters in a derogatory or disadvantageous manner within works of 昀椀ction. This bias can
manifest through various means, such as reinforcing stereotypes, limiting agency or
complexity, or presenting characters in a demeaning or objectifying light based on their gender. In
the seminal study conducted by Jockers and Kirillo1昀[0], the authors explored these biases
by investigating the association between verbs and gendered pronouns and nouns in
characterizations. They concluded that gender serves as a confounding factor in character agency
and behavioral representations. Based on these 昀椀ndings, Underwood, Bamman, and Lee17[]
revealed a disproportionate representation of male characters by male authors, whereas female
authors tended to provide a more equal representation of both genders. The key 昀椀nding
indicated a decline in gender-based distinctiveness of 昀椀ctional characters over time, suggesting a
blurring of gender boundaries in contemporary 昀椀ction.</p>
      <p>
        Following this line of research, the aim of this study is to identify consistent patterns of
negative portrayals of female characters or the perpetuation of stereotypes, which may indicate
gender bias. In alignment with Naguib, Delaborde, Andrault, Bekolo, and Semin1c5k],[we will
concentrate on character space detection1[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] through coreference resolution5][. Our focus
will be on evaluating how adjectives and verbs are utilized di昀erently based on the gender of the
characters. Building upon Antoniak, Field, Mun, Walsh, Klein, and S1a]p, [which demonstrates
that detecting power relations, and agency in characterization can unveil gender bias, we aim
to elucidate how gender modi昀椀es portrayals in novels. This approach is anticipated to o昀er
a more precise and insightful understanding of how gender representation unfolds within the
context of literary works.
      </p>
      <p>The 昀椀rst part of this paper is about the creation of a large annotated corpus using the
predictions of a classi昀椀er. On this large annotated corpus, three main experiments were made
to determine whether characters are described di昀erently depending on their gender. First,
we focused on agency using linguistic features: we tried to 昀椀nd if women were more o昀琀en
grammatical objects of the sentence than men. The results are presented in sectio3.n2. The
second experiment, detailed in section3.1, was to use lemmatized words linked to a character
to predict its gender: if it works, it would mean that there are ’gendered’ words that help the
classi昀椀cation. Finally, we explored the vocabulary in more details to 昀椀nd out which words
were more associated to women or to men in sectio3n.3.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Materials and methods</title>
      <sec id="sec-2-1">
        <title>2.1. Corpus</title>
        <p>We used the Chapitres1 corpus for our analysis. It is comprised of 2942 French novels in
XMLTEI2 (Text Encoding Initiative) encoding from 1811 to 2020, put together in the context of the</p>
        <sec id="sec-2-1-1">
          <title>1Chapters in French.</title>
          <p>2TEI Consortium, eds. TEI P5: Guidelines for Electronic Text Encoding and Interchange. Version 1.0. TEI
Consortium. http://www.tei-c.org/Guidelines/P5/.
french ANR3 project ”Chapitres”4. An overview of the ”Chapitres” corpus is given in tabl1e.
Sentences and tokens were computed using the Spacy 8[] fr_dep_news_trf mode5l.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Annotation</title>
        <p>A 昀椀rst part of this work was to manually annotate the gender of the 10 main characters from
100 novels picked randomly in the corpus Chapitres. The four classes are:
• ’m’ for masculine characters
• ’f’ for feminine characters
• ’p’ for plural characters
• ’u’ (unde昀椀ned) when there were not enough clues</p>
        <p>A 1000 lines (one for each character) tsv 昀椀le was created. To each character, the most frequent
mentions, and related adjectives and verbs were associated. (see sectio2n.3 for the extraction
of related adjectives and verbs). A mention of a character is a linguistic item referring to this
given character, notably proper names and pronouns.</p>
        <p>
          Annotation was performed by three annotators who are native speakers. The gender was
attributed depending on the mentions (pronouns, names, ...), and the gender of the adjectives.
The latter are useful for 昀椀rst person narrators, as mentions do not carry gender marks. The
dataset was split evenly across all annotators to ensure that every character would be annotated
twice. This allowed us to compute an inter annotator agreement (IAA) using NLTK1’s3][
implementation of Cohen’s [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Overall on the corpus was0.7256. The annotations were
a昀琀erwards adjudicated to provide a gold standard. While considered ”substantial” by some
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], others recommend a highe r of 0.8 [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], using 0.67 as lowest acceptable value. We then
consider our IAA as acceptable. In tabl2eare presented four examples of annotations, one for
each class. In table3 are presented the results of the annotation.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Textual features</title>
        <p>For each character, the goal was to obtain all its mentions in the book, and all adjectives and
verbs having a dependency link with a mention of this character. A 昀椀rst step was to associate
3FNA, French National Agency
4Website of the the project:https://chapitres.hypotheses.org
5The model is described at the following URhLt:tps://spacy.io/models/fr#fr_dep_news_trf
each token with 2 informations: is this token a character, and what gender is this character. To
store these informations, we created a tsv format with the columns presented in ta4b.le</p>
        <p>
          By using an NLP pipeline speci昀椀cally tuned for novels, (fr-BookNLP, part of the multilingual
BookNLP project [
          <xref ref-type="bibr" rid="ref3 ref4">4, 3</xref>
          ]), we extracted literary characters along with all their mentions, thus
resolving their coreferenc6e.
        </p>
        <p>
          This NLP pipeline allowed us to associate column 7 (character id) to the 昀椀rst column. The 10
most important characters of each book were considered, thus column 7 contains character ids
from 1 to 10 when the token is a mention of a character. Then, all sentences from our corpus
were parsed with Spacy 8[
          <xref ref-type="bibr" rid="ref9">, 9</xref>
          ] to obtain columns 2 to 6. For the last column, we had to manually
annotate the gender (as described in section2.2), or we used the predictions (see section2.4)
for the unannotated part.
        </p>
        <sec id="sec-2-3-1">
          <title>6see appendix A.2 for the evaluation of Fr-BookNLP</title>
          <p>In table4, a sentence fromClaudine en ménage by Colette is presented. The character with
id 1 was annotated as a female and the one with id 4 as a male, as can be seen in the last two
columns. ’Renaud’ is the subject of the sentence. Character one’s 昀椀rst mention is ’me’, which
is the object of the sentence (see 6th column).</p>
          <p>In a second step, the information given by Spacy was used to de昀椀ne the lists of related
adjectives and verbs. For each character, we retrieve all tokens that are linked to a mention of
this character (each token whose index correspond to the head that governs the mention), if
these tokens are adjectives or verbs.</p>
          <p>To represent the features (the adjectives, verbs or mentions related to characters) as vectors,
we decided to use a Bag of Words (BoW) representation. Thus, each line corresponds to a
character, and each column represents a word. In each cell is found the number of occurrences
of the word corresponding to the column for the character given by the line, divided by the
total number of words associated to this character. This allows to have the relative frequency
for each word. Because of the over-representation of masculine characters in our corpus (see
table3), we decided to balance the classes.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Building a classifier for gender predictions to annotate the corpus</title>
        <p>
          We trained a classi昀椀er to predict the gender of characters and used these predictions as
annotations for our experiments described below. We employed a Random Forest classi7昀椀e]r, [
implemented in scikit-learn1[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] to predict character gender. We ran our model in a basic
5fold cross validation set up. Classes ’u’ (unde昀椀ned) and ’p’ (plural) were removed to make the
predictions easier.
        </p>
        <p>Using a Bag of Words representation of the most frequent character mentions as input, the
classi昀椀er proved e昀ective in predicting gender based on crucial words such as ’il’ (he) and ’elle’
(she), as depicted in Figure1.</p>
        <p>We then inferred gender across all our corpora, which encompasses a vast array of 29,490
characters, holds signi昀椀cant importance in unveiling broader patterns and trends in gender
representations within the literary landscape. By examining a wide range of characters from
various genres, time periods, and cultural backgrounds, we aim to discern recurrent patterns
and subtle nuances in character depictions, shedding light on the implicit gender constructs
prevalent throughout literary history.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Predicting gender without grammatical gender marks</title>
        <p>Following Underwood, Bamman, and Lee17[], we tried to reproduce on our corpus of French
novels some of their 昀椀ndings. For this experiment, we use a similar method to the one used to
annotate the corpus, described in2.4. We adapted the features used to predict gender: instead
of using mentions, we used adjectives and verbs.</p>
        <p>A major di昀erence between French and English that we had to take into account for this
experiment is the gender mark on certain adjectives and participles. For example, in the French
sentence below there are two signs that the speaker is feminine, while there are none in the
English equivalent:
”Je suis heureuse d’y être alleé”
”I’m glad I went there”</p>
        <p>If a classi昀椀er is able to correctly predict whether a character is feminine or masculine due to
these marks, there is no evidence that gender stereotypes helped the classi昀椀cation. Therefore,
we chose to use lemmas instead of word forms. Thus, in the example above, the gendered words
would become ’heureux’ and ’aller’. This way, if the classi昀椀er achieves an accuracy which is
signi昀椀cantly higher than the baseline, it will be because it used sociological features instead of
grammatical ones.</p>
        <p>m’ (me)
mmmee(M(mres))</p>
        <p>mon (my)
du (ø / some)
ses (his / her)
une (a / one)</p>
        <p>ma (my)
son (his / her)</p>
        <p>je (I)
sa (his / her)
madame (Madam / Missus)</p>
        <p>cette (this)
lui (him / her)
le (the)
la (the)
il (he)
elle (she)
0</p>
        <p>50
feature weight (×103)
100</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.6. Agency</title>
        <p>In a literary setting, character agency refers to the capacity of a character to take intentional and
autonomous actions within the narrative. It is the degree to which a character is portrayed as
an active agent who drives the events and makes choices that impact the story’s progression.
Characters with low agency may be more passive or reactive, in昀氀uenced by external forces
or events rather than actively driving the story. They might have limited control over their
circumstances or be more prone to being acted upon rather than taking independent action.
To evaluate a character’s agency, we compared the number of mentions of this character that
are subject of the sentence, and the ones that are object. To retrieve this information, we used
Spacy as dependency parser.</p>
        <p>To compare the degree of agency of di昀erent characters, we created a metric. Thaegency
score, between -1 and 1, corresponds to:</p>
        <p>With nsubj the number of occurrences of a character or group of characters as subject of the
sentence, and obj the number of occurrences of a character or group of characters as object of
the sentence.</p>
        <p>The higher the score, the more o昀琀en the character (or group of characters) holds the position
of subject of the phrase, and thus has an active role.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <sec id="sec-3-1">
        <title>3.1. Character gender prediction</title>
        <p>3.1.1. For the corpus annotation
Resulting accuracies range from 60.25% to 91.3%, as described in Tab5l.eThe classi昀椀er
demonstrated a good accuracy rate when using the mentions as features, achieving 91.3%. The errors
are mostly due to inaccuracies in the identi昀椀cation of characters by BookNLP: it happens
sometimes that a character has both feminine and masculine mentions because two di昀erent
characters were mixed up. However, we considered that 91.3% accuracy was enough to annotate
the whole corpus with this method.
3.1.2. For the bias analysis
In our 昀椀rst experiment (second line in the table), adjectives are used to predict the gender
of the characters they describe. We ran the same classi昀椀er using lemmas instead of words
for comparison (third line of the table). The accuracy using words is 0.782, which is higher
than the 0.602 accuracy obtained with lemmas This may indicate that the classi昀椀er relied on
grammatical gender marks when adjectives were not lemmatized.</p>
        <p>It is also interesting to notice that the results are better when using only adjectives when
word forms are used (0.782 with adjectives against 0.74 with adjectives and verbs), whereas
it is the contrary when using lemmas (0.602 with adjectives against 0.612 with adjectives and
verbs). A possible explanation is that adjectives bear a gender mark more than verbs. This
would mean that the classi昀椀er relies a lot on these gender marks when it can.</p>
        <p>Now, the results of the classi昀椀er using lemmas are interpretable. In tab6leis presented the
confusion matrix obtained with the BoW representation containing lemmatized adjectives and
verbs as input to a Random Forest. Because the two classems ale and female were balanced, the
baseline is 0.5. The classi昀椀er is able to predict the gender of a character given the adjectives and
verbs related to it with an accuracy of 0.612. The features that were most useful for the classi昀椀er
to classify the characters are presented in 昀椀gure2, and will be discussed later. Eventually, this
result indicates that male and female characters are not characterized in the same way.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Agency</title>
        <p>The results of our computations show that masculine characters tend to be a little bit more
’agent’ than feminine ones, but the di昀erence is not striking. Indeed, we can see in table7 that
men are more o昀琀en subjects than women (85% of the time against 82%) while they are less o昀琀en
objects, but their agency scores are close. However, the small di昀erence in the agency scores
might imply that the descriptions of male characters exhibit slightly higher agency than those
of female characters.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Gender biases in the vocabulary</title>
        <p>Using the method described in section2.3, we were able to obtain the verbs and adjectives
having a dependency link with a given character. From this vocabulary, our goal was to determine
whether some words were associated to one gender more than to the other.</p>
        <p>As seen in 昀椀gure 4, we can see that a common stereotype is to associate women to love
and passion, and men to action. We wanted to investigate about such well-known stereotypes,
to see if they also appear through words in the corpus. A simple and signi昀椀cant example is
the verb ’aimer’ (to love): we found that this verb occurs at least once for 44% of the female
characters, against 25% for the male characters.</p>
        <p>To visualize the results, plots were made with, for each verb, whether it is more associated
to women or to men in the corpus. More precisely, the values correspond to:
 −  
 +</p>
        <p>With masc the proportion of masculine characters to which the given verb is related in the
corpus, andfem the proportion of feminine characters. Thus, the more positive the values are,
the more the word is associated to men, and negative with women. To avoid imprecision due
to lack of data, verbs occurring for less than 1% of the characters were not considered.
3.3.1. Gendered di昀erentiation through action verbs
This experiment was, among others, used to have another approach of agency, more focused
on the lexicon than on the dependencies. Two kinds of verbs were used: action verbs (昀椀gure
3), and emotion verbs (昀椀gure 4). The lists are quite arbitrary: the categories ’action verb’ and
’emotion verb’ are not 昀椀xed. However, the list of action verbs contains verbs that show an
in昀氀uence on the narration. We considered that the more a character or a gender was related to
these verbs, the more he is an agent of the narration and plays an active role in it. Conversely,
emotion verbs are more static and focus on the mind.</p>
        <p>When comparing 昀椀gures 3 and 4, it appears that most of the ’action verbs’ are considered as
more masculine, and most of the ’emotion verbs’ occur more for feminine characters than for
masculine ones. Among these verbs conveying an emotion, those related to love are clearly
more feminine: ’aimer’ (to love), ’adorer’ (to really like, to adore).</p>
        <p>The mean of the proportion of masculine and feminine characters linked to a certain category
of verbs or adjectives was computed to compare more precisely the two genders. The results,
presented in table8, correspond (for each category of adjectives or verbs) to:
∑#= 1
%</p>
        <p>#</p>
        <p>Action verbs occur in average for 3.66% of masculine characters, against 3.45% of feminine
characters. The di昀erence is slight; however, it is bigger when looking at the emotion verbs:
9.11% of men are, in average, linked to an emotion verb, against 12.26% of women.
0.2</p>
        <p>0
−0.2
3.3.2. Other comparisons
We created a non-exhaustive speech verbs list to see which speech verbs are more associated
to a gender than to the other. Figure5 presents the results. Some interesting observations
can be made: 昀椀rst, the verb ’crier’ (to shout) is more masculine, as opposed to ’murmurer’ (to
whisper), which is more feminine. This seems to refer to the stereotype associating women to
quietness and men to loudness.</p>
        <p>Similarly, gender-related adjectives also contribute to reinforcing stereotypes in French
literature. In 昀椀gure 6, adjectives such as ”joli” (pretty) and ”doux” (gentle) are commonly used
to describe female characters, emphasizing their appearance and nurturing qualities. While
adjectives such as ”vieux” (old) and ”gros” (fat) are more frequently applied to male characters,
their antonyms ”jeune” (young) and ”昀椀n” (thin) are more associated to women. This opposition
seems to refer to the valorization of youth and thinness for women.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
      <sec id="sec-4-1">
        <title>4.1. Analysis of the discriminant features</title>
        <p>Considering the model that uses lemmas to predict the gender (third model in ta5b) lseo it
cannot use gender evidence from word, we want to discuss here the possibility that these
features show gender biases. The discriminant features of our model, presented in 昀椀gu2r,eare
the features that helped the most our classi昀椀er to predict the gender of a character.</p>
        <p>Some of these lemmas are particularly associated to women or to men, as showed in section
3.3. On the 昀椀rst hand, the verb ’aimer’ (to love) is much more associated to feminine characters,
jolib(eplr(ebtetdayuo)utixfu(ls)wlaeiedt)(puâglleyp)(eptaitle(s)mdujaerlul()nhear(dy)oufning)g(mtrhaainnidg)r(eta(sllk)ingnryo)sv(ifeautx)(old)
and is also the fourth most useful feature to determine if a character is feminine or masculine
(see 昀椀gure 2). Verbs commonly associated with female characters include ”pleurer” (to cry),
”reprocher” (to reproach), ”murmurer” (to murmur), and ”rire” (to laugh). These verbs tend to
emphasize emotional expressions and portray women as more emotive and sensitive.</p>
        <p>On the other hand, verbs associated with male characters include ”tirer” (to pull), ”découvrir”
(to discover), and ”expliquer” (to explain). These verbs o昀琀en depict actions and intellectual
pursuits, implying that male characters are more active and analytical in their roles.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Limitations of the approach</title>
        <p>Our approach is subject to the inherent accuracy limitations of many NLP and machine learning
algorithms used, including fr-BookNLP, Spacy, and our Random Forest, all of which are prone
to making errors.</p>
        <p>As mentioned, the classi昀椀er used to predict the gender of the 29,490 characters achieves a
0.91 accuracy on the set of 1000 annotated characters. This means that approximately 9% of
the characters on which the experiments were made do not have the correct label.</p>
        <p>Most of the time, the errors are certainly due to the imprecision of the retrieved characters.
Indeed, when doing the manual annotation, 9% of the characters were classi昀椀ed as ’unde昀椀ned’
because the mentions referred in fact to di昀erent characters, as showed in tabl2e.</p>
        <p>A limitation of this study lies in the fact that the temporal dri昀琀 of gender bias has not been
explored. Our intention was to develop methods to delve into this question using a large corpus
and identify the associated challenges. Future research will speci昀椀cally focus on this aspect,
aiming to assess the resistance of these stereotypes to temporal shi昀琀s and explore their
evolution within 昀椀ction.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In conclusion, this article focuses on the representation of gender in French 昀椀ction, from 1800 to
the present day. By examining gender bias prevalent in French literature, the study sheds light
on how authors perpetuated gender stereotypes while shaping their characters and narratives.
The study highlights the importance of critically examining such representations to foster a
more inclusive and diverse literary landscape.</p>
      <p>Our statistical modeling reach from 61% to 91% accuracy, depending on the features given
to the classi昀椀er (with or without obvious gender marks). We employed it to infer the gender of
29490 characters, on which we can run some experiments. This allowed us to study how each
gender is portrayed in French literature.</p>
      <p>The analysis of French literature revealed certain gender stereotypes that are re昀氀ected in
the language used to describe female and male characters. These linguistic patterns in the
portrayal of characters in French literature re昀氀ect underlying gender bias and perpetuate
traditional gender norms and roles. For example, we found that women get characterized with
far more physical and emotional adjectives and verbs than men. These last were more linked
with action verbs, indicating a stronger agency for male characters.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was funded in part by the French government under management of Agence
Nationale de la Recherche as part of the “Investissements d’avenir” program, reference
ANR-19P3IA0001 (PRAIRIE 3IA Institute). Jean Barré’s PhD is supported by the EUR (Ecole
Universitaire de Recherche) Translitterae (programme “Investissements d’avenir” ANR-10-
IDEX-000102 PSL and ANR-17-EURE-0025).</p>
    </sec>
    <sec id="sec-7">
      <title>A. Appendix</title>
      <sec id="sec-7-1">
        <title>A.1. Translation of table 2 A.2. Corpus distribution</title>
        <p>Figure 7 shows the distribution of novels over time broken out by author’s gender. The
Chapitres corpus spans from 1810 to 2020, with a notable concentration of novels in the
latter half of the 19th century. The proportion of female writers remains between 15 and 20% of
the literary production, with a slight increase towards the end of the corpus, reaching about
30%. The corpus includes 714 authors, with 149 (21%) being female writers, contributing to a
total of 626 novels (27%).
250
200
150
100
50
0
male writers
female writers</p>
      </sec>
      <sec id="sec-7-2">
        <title>A.3. Fr-BookNLP evaluation</title>
        <p>We evaluated Fr-BookNLP coreference pipeline using three metrics used in state of the art
coreference evaluation. All three evaluate the correctness of coreference chains and how well</p>
        <p>None of them are really relevant for coreference at the novel’s scale, so we follo3w]ed [
coreference evaluation, averaging those three metrics. This version of Fr-BookNLP is ten points
behind English evaluation.
tion
Evaluation of the fr-BookNLP pipeline, scores for character mentions detection and coreference
resolu</p>
        <p>Metrics
43,5</p>
        <p>Coreference resolution is still hard for long texts. An issue is that characters can get
duplicated, that is, the same character is detected in multiple coreference chains. While problematic
at 昀椀rst glance, our primary objective is not to retrieve unique and distinct characters but a
proxy for characterization as a whole. The focus is more on character’s gender representation
and the overall impact of these characters on the literary landscape, rather than identifying
separate and non-repeating characters.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Antoniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Field</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Walsh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Klein</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Sap</surname>
          </string-name>
          . “Riveter:
          <article-title>Measuring Power and Social Dynamics Between Entities”</article-title>
          .
          <source>In:Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>3</volume>
          :
          <string-name>
            <given-names>System</given-names>
            <surname>Demonstrations</surname>
          </string-name>
          <article-title>)</article-title>
          . Toronto, Canada: Association for Computational Linguistics,
          <year>2023</year>
          , pp.
          <fpage>377</fpage>
          -
          <lpage>388</lpage>
          .
          <year>do1i0</year>
          :.
          <volume>18653</volume>
          /v1 /
          <year>2023</year>
          .acl-demo.
          <volume>36</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Artstein</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Poesio</surname>
          </string-name>
          . “
          <article-title>Inter-coder agreement for computational linguistics”</article-title>
          .
          <source>In: Computational linguistics 34.4</source>
          (
          <issue>2008</issue>
          ), pp.
          <fpage>555</fpage>
          -
          <lpage>596</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bamman</surname>
          </string-name>
          .
          <source>BookNLP</source>
          .
          <year>2021</year>
          . url: https://github.com/booknlp/booknl.p
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bamman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Underwood</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Smith. “A Bayesian Mixed</surname>
          </string-name>
          <article-title>E昀ects Model of Literary Character”</article-title>
          . In:
          <article-title>Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</article-title>
          .
          <source>Acl</source>
          <year>2014</year>
          . Baltimore, Maryland: Association for Computational Linguistics,
          <year>2014</year>
          , pp.
          <fpage>370</fpage>
          -
          <lpage>379</lpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .3115/v1/
          <fpage>P14</fpage>
          -1035.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Barré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cabrera Ramírez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mélanie</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Galleron.</surname>
          </string-name>
          “
          <article-title>Pour une détection automatique de l'espace textuel des personnages romanesques”</article-title>
          .
          <source>InH:</source>
          umanistica
          <year>2023</year>
          .
          <article-title>Corpus. Association francophone des humanités numériques</article-title>
          . Genève, Switzerland,
          <year>2023</year>
          , pp.
          <fpage>56</fpage>
          -
          <lpage>61</lpage>
          . url: https://hal.science/hal-04105537.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <article-title>“A coe昀케cient of agreement for nominal scales”</article-title>
          .
          <source>In:Educational and psychological measurement 20.1</source>
          (
          <issue>1960</issue>
          ), pp.
          <fpage>37</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T. K.</given-names>
            <surname>Ho</surname>
          </string-name>
          . “
          <article-title>Random decision forests”. InP:roceedings of 3rd international conference on document analysis and recognition</article-title>
          . Vol.
          <volume>1</volume>
          .
          <string-name>
            <surname>Ieee</surname>
          </string-name>
          .
          <year>1995</year>
          , pp.
          <fpage>278</fpage>
          -
          <lpage>282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Honnibal</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Johnson</surname>
          </string-name>
          . “
          <article-title>An improved non-monotonic transition system for dependency parsing”</article-title>
          .
          <source>In:Proceedings of the 2015 conference on empirical methods in natural language processing</source>
          .
          <year>2015</year>
          , pp.
          <fpage>1373</fpage>
          -
          <lpage>1378</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Honnibal</surname>
          </string-name>
          , I. Montani,
          <string-name>
            <surname>S. Van Landeghem</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Boyds</surname>
          </string-name>
          .paCy:
          <string-name>
            <surname>Industrial-strength Natural</surname>
          </string-name>
          Language Processing in Python.
          <year>2020</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.1212303.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Jockers</surname>
          </string-name>
          and
          <string-name>
            <surname>G. Kirillo昀.</surname>
          </string-name>
          “
          <article-title>Understanding Gender and Character Agency in the 19th Century Novel”</article-title>
          .
          <source>In:Journal of Cultural Analytics</source>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .22148/16.010.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Koolen</surname>
          </string-name>
          . “
          <article-title>Reading beyond the female: the relationship between perception of author gender and literary quality”</article-title>
          .
          <source>PhD thesis</source>
          . University of Amsterdam - Universiteit van Amsterdam,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Landis</surname>
          </string-name>
          and
          <string-name>
            <surname>G. G. Koch. “</surname>
          </string-name>
          <article-title>The measurement of observer agreement for categorical data”</article-title>
          . In: biometrics (
          <year>1977</year>
          ), pp.
          <fpage>159</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Loper</surname>
          </string-name>
          and
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>BirdN.ltk: The natural language toolkit</article-title>
          .
          <year>2002</year>
          . arXiv:cs/0205028.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Moretti</surname>
          </string-name>
          . “Conjectures on world literature”.
          <source>NIne:</source>
          w le昀琀 review (
          <year>2000</year>
          ), pp.
          <fpage>54</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Naguib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Delaborde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Andrault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bekolo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Seminck</surname>
          </string-name>
          . “
          <article-title>Romanciers et romancières du XIXème siècle : une étude automatique du genre sur le corpus GIRLS (Male and female novelists : an automatic study of gender of authors and their characters)”</article-title>
          . In:
          <article-title>Actes de la 29e Conférence sur le Traitement Automatique des Langues Naturelles. Atelier TAL et Humanités Numériques (TAL-HN)</article-title>
          . Avignon, France: Atala,
          <year>2022</year>
          , pp.
          <fpage>66</fpage>
          -
          <lpage>77</lpage>
          . url: https://aclanthology.org/
          <year>2022</year>
          .jeptalnrecital-humanu
          <source>m..8</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Duchesnay.</surname>
          </string-name>
          “
          <article-title>Scikit-learn: Machine Learning in Python”</article-title>
          .
          <source>IJno:urnal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          ), pp.
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>T.</given-names>
            <surname>Underwood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bamman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          . “
          <article-title>The Transformation of Gender in EnglishLanguage Fiction”</article-title>
          .
          <source>In:Journal of Cultural Analytics</source>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .22148/16.019.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Woloch</surname>
          </string-name>
          .
          <article-title>The One vs</article-title>
          .
          <source>the Many</source>
          . Princeton University Press,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>