<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Summarising scRNAseq expression data in FlyBase</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Damien Goutte-Gattat</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Physiology, Development and Neuroscience, University of Cambridge</institution>
          ,
          <addr-line>Cambridge CB2 1TN</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Single-cell RNA sequencing has proved an invaluable tool in biomedical research. The ability to survey the transcriptome of individual cells offers many opportunities and has already paved the way to many discoveries in both basic and clinical research. With the increasing amount of single-cell transcriptomic data available, including whole-organism single-cell transcriptomic atlases, biological databases face a challenge to integrate these data and make them easily accessible to their users. In this Short Paper, we describe how the fruit fly-specific database FlyBase is making use of single-cell RNA sequencing datasets to let fly researchers quickly know in which cell types genes are known to be expressed.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Gene expression</kwd>
        <kwd>single-cell RNA sequencing</kwd>
        <kwd>cell types</kwd>
        <kwd>biocuration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>First introduced in 2009, single-cell RNA
sequencing (scRNAseq) has become a powerful tool
to investigate cell states and functions [1]. Since
2017, many studies have leveraged the technique in
the fruit fly Drosophila melanogaster, leading to
various new insights [2], and the number of
scRNAseq datasets from flies is only expected to
grow quickly in the coming years. Recently, a
consortium of 40 fly laboratories, the Fly Cell Atlas
project, completed and reported the first single-cell
transcriptomic atlas of the entire adult fruit fly,
sequencing more than half a million cells of 250
different types across 17 tissues [3].</p>
      <p>FlyBase is the Model Organism Database (MOD)
for all data related to Drosophila melanogaster [4]. It
provides access to a wide range of scientific
information either manually curated from the
published literature or from high-throughput research
projects. Expression data are a particularly important
subset, and we strive to give access to as many
resources as possible to allow fly researchers to find
out in which tissues and at which developmental
stages their gene of interest is known to be
expressed. To that end, we already make use of
several high-throughput projects such as
modENCODE [5], FlyAtlas [6], and FlyAtlas 2 [7].
We now want to exploit available scRNAseq</p>
    </sec>
    <sec id="sec-2">
      <title>2. Aims</title>
      <sec id="sec-2-1">
        <title>We want to help FlyBase users to: (a) discover</title>
        <p>the available Drosophila scRNAseq datasets; (b) get
some information (metadata) about these datasets;
(c) get a quick overview of the expression data from
those datasets.</p>
        <p>In this paper, we will focus on (c), and explain
how we leverage the expression data provided by
scRNAseq experiments to obtain and present a
percell type summary of the expression data, so that
users can gen an “immediate” answer (straight from
the “Gene Report” page) to the following questions:
What are the cell types in which a given gene is
expressed? What is the proportion of cells of a given
type in which the gene is expressed? What is the
average expression level of the gene across all cells
of that type?</p>
        <p>An important design decision is that no
scRNAseq dataset will ever be stored within FlyBase
itself. Instead, we will only store some metadata as
well as a simplified version of the gene expression
data – only what is needed to cover the questions
listed above. Users wanting to inspect the full data
will be directed to a pre-existing external data store
where the data will already be available.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Data pipeline</title>
    </sec>
    <sec id="sec-4">
      <title>3.1. Data source</title>
      <sec id="sec-4-1">
        <title>All the scRNAseq data used in FlyBase are</title>
        <p>obtained through the Single Cell Expression
Atlas (SCEA), the EMBL-EBI resource for gene
expression data at the single cell level [8]. We
thus have a single data provider that is
independent of all the individual research
projects. The SCEA data curators process the
original raw data, as deposited by the original
authors in common data stores such as the Gene
Expression Omnibus or Array Express (Figure 1,
steps 1–2), according to a standard pipeline with
common parameters that allows for cross-dataset
comparisons. As an added benefit, the SCEA
provides the processed data in a common format,
independently of the actual scRNAseq method
originally used (e.g., Smart-seq2 or 10⨉
sequencing).</p>
      </sec>
      <sec id="sec-4-2">
        <title>1. the authors may have used a “common”</title>
        <p>cell type name, not knowing that this cell type
is formally known under another name in the
DAO (an example is “astrocyte”, for which
the actual term in the DAO is “astrocyte-like
glial cell”);
2. an annotation may refer to a cell state
rather than, or in addition to, a cell type (an
example is “plasmatocyte-prolif”, which was
used in a dataset to annotate a cluster of
proliferating plasmatocytes; the appropriate
ontology term in that case is “plasmatocyte”,
leaving the cell state aside);
3. an annotation may incorrectly refer to an
organ or a tissue rather than to the cell types
that make up this organ or tissue (an example
was a cluster annotated as “dorsal vessel”,
where the correct cell type name should have
“cardial cell”);
4. an annotation may refer to several cell
types (an example was “OPN innervating
DA1, VA1 or DC3 glomerulus”; the correct
term in such case is the closest cell type that
encompasses all cell types in the cluster, in
this instance “olfactory projection neuron”).</p>
        <p>Lastly, there is one case where the original
annotation cannot be converted to an ontology
term: when the cell type is either not clearly
identified (e.g. “btl-GAL4 positive cell, likely to
be ovary cell”), or identified only for what it is
not (e.g. “non-hemocyte”). Such clusters will be
excluded from the expression data summarisation
described below.</p>
        <p>Of note, this “conversion” of the original
annotations is not destructive: the original
annotations from the upstream authors are at all
time preserved. FlyBase curators merely add a
new set of ontology-compliant annotations
alongside the original annotations. Both sets of
annotations are ultimately available on the Single
Cell Expression Atlas.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>3.3. Ontology updates</title>
      <p>In the process of converting the original
annotations to DAO-compliant annotations, the
need to update the DAO itself may arise, for
mostly two reasons: a) a dataset reports a new
cell type, which must be added to the ontology
(an example was “adipohemocyte”, a subtype of
plasmatocytes found in a scRNAseq analysis of
the lymph gland [10]); b) the ontology is found to
lack some “grouping terms”, useful to offer a
better qualification of some clusters (an example
was the addition of “endocrine cell of the ring
gland”, to encompass the various types of
endocrine cells found in the ring gland). FlyBase
data curators then work alongside FlyBase
ontology editors to update the DAO as needed
and the newly added terms are used to annotate
the scRNAseq clusters.</p>
      <p>In some cases, this process of updating the
DAO to better annotate a dataset may actually
start before the dataset is even published, through
a direct collaboration between FlyBase and the
dataset’s authors. This was notably the case of
the Fly Cell Atlas dataset [3].</p>
    </sec>
    <sec id="sec-6">
      <title>3.4. Summarising the expression data</title>
      <sec id="sec-6-1">
        <title>Once the raw sequencing data have been</title>
        <p>processed by the SCEA standard pipeline, the
most important output is the Normalised
Expression Matrix: for any single dataset, this is
a large matrix of n cells by m genes, with n the
total number of cells in the dataset and m the total
number of identified genes; each matrix cell
contains the normalised reads count for a
particular gene in a particular cell. This matrix is
accompanied by the Experiment Design Table,
which contains annotations for each single cell
(notably the cell type annotations, both as
submitted by the original authors and as corrected
by the FlyBase curators as described above).</p>
        <p>From those two files, we proceed to what we
call the per-cell type summarisation of the gene
expression data (Figure 1, step 7). For each gene
and for each identified cell type, we extract three
measures: a) the cluster size, that is the number
of cells annotated with that cell type; b) the
spread or extent of expression, the proportion of
cells in the cluster in which the gene is expressed
(the number of cells in the cluster with a non-zero
reads count for that gene, divided by the cluster
size); and c) the average expression of the gene
in positive cells (the total number of reads across
all cells in the cluster, divided by the number of
cells with a non-zero count). Those measures are
then stored in the Chado database at the heart of
FlyBase, from where they can be exploited on the
FlyBase website.</p>
        <p>The summarised data are also slated to be
used by Virtual Fly Brain [11]. To that effect, a
subset of the data, corresponding to the cell types
present in the adult brain only and to genes with
an extent of expression of more than 0.9 (genes
that are expressed in more than 90% of cells of a
given type), is loaded from the Chado database to
the Neo4j graph database (Figure 1, step 8) that
powers the Virtual Fly Brain website.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>4. User-visible outcomes</title>
      <sec id="sec-7-1">
        <title>The summarised expression data will be</title>
        <p>exploited on the FlyBase website in several
incremental steps. The first step was realised for
the 2022_03 release of FlyBase in June 2022. In
this release, a cell type ribbon was added to the
Gene Report page (Figure 2A). The ribbon lists
22 high-level cell types, each of those being
coloured according to the extent of expression of
the current gene in cells of that type. Because
several more precise cell types are agglomerated
into a single high-level cell type, a list of all the
precise subtypes, with their corresponding extent
of expression, is given in a tooltip box shown
when the mouse pointer is left hovering over a
ribbon cell (Figure 2B). To avoid cluttering that
tooltip, cell types where the extent of expression
is less than 1% are filtered out.</p>
        <p>At present, the cell type ribbon is only fed
with the data from the Fly Cell Atlas project [3].
In a second step and as more datasets will be
accumulated into FlyBase, we will update the
ribbon to display consolidated data from all the
available scRNAseq datasets.</p>
        <p>Lastly, in a future release we will add a new
type of graph to the “high-throughput expression</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>6. References</title>
      <p>[1] Haque, A., Engel, J., Teichmann, S.A.,</p>
      <p>Lönnberg, T., A practical guide to
singlecell RNA-sequencing for biomedical
research and clinical applications. Genome</p>
      <p>Med 2017, 9, 1–12.
[2] Li, H., Single-cell RNA sequencing in</p>
      <p>Drosophila : Technologies and applications.</p>
      <p>WIREs Dev Biol 2021, 10.
[3] Li, H., Janssens, J., De Waegeneer, M.,</p>
      <p>Kolluru, S.S., et al., Fly Cell Atlas: A
single-nucleus transcriptomic atlas of the
adult fruit fly. Science 2022, 375,
Figure 2: Graphical display of summarised [4] eLaabrkki2n4,3A2.., Marygold, S.J., Antonazzo, G.,
expression data on FlyBase. (A) The cell type Attrill, H., et al., FlyBase: updates to the
ribbon. Each cell in the ribbon corresponds to Drosophila melanogaster knowledge base.
one of the main Drosophila cell types. Each cell is Nucleic Acids Res 2020, 49, D899–D907.
coloured depending on the fraction of cells of [5] Graveley, B.R., Brooks, A.N., Carlson,
that type in which the current gene is expressed J.W., Duff, M.O., et al., The developmental
according to the Fly Cell Atlas dataset. (B) When transcriptome of Drosophila melanogaster.
the mouse pointer is hovering over one cell, a Nature 2011, 471, 473–479.
tooltip shows the list of cell subtypes in which [6] Chintapalli, V.R., Wang, J., Dow, J.A.T.,
the current gene is expressed, along with the Using FlyAtlas to identify better Drosophila
proportion of cells expressing the gene in each melanogaster models of human disease. Nat
subtype. (C) Mockup of a more complete Genet 2007, 39, 715–720.
graphical representation of summarised [7] Leader, D.P., Krause, S.A., Pandit, A.,
expression data that is planned for a future Davies, S.A., Dow, J.A.T., FlyAtlas 2: a
release of FlyBase. For each cell type, this graph new version of the Drosophila melanogaster
will display both the proportion of cells of that expression atlas with RNA-Seq,
miRNASeq and sex-specific data. Nucleic Acids
type in which the gene is expressed (left) and the Research 2018, 46, D809–D815.
average expression level in all cells that do [8] Papatheodorou, I., Moreno, P., Manning, J.,
express the gene (right). Fuentes, A.M.-P., et al., Expression Atlas
data” section of the Gene Report page. The exact update: from tissues to single cells. Nucleic
modalities of this new graph are yet to be Acids Research 2019, gkz947.
determined, but il will display both the extent of [9] Costa, M., Reeve, S., Grumbling, G.,
expression and the average expression level for a Osumi-Sutherland, D., The Drosophila
given gene across high-level cell types (Figure anatomy ontology. Journal of Biomedical
2C). Semantics 2013, 4, 32.
[10] Cho, B., Yoon, S.-H., Lee, D., Koranteng,</p>
      <p>F., et al., Single-cell transcriptome maps of
5. Acknowledgements myeloid blood cell lineages in Drosophila.
Nat Commun 2020, 11, 4483.</p>
      <p>This work was supported by joint grants [11] Milyaev, N., Osumi-Sutherland, D., Reeve,
BB/T014008/1 (FlyBase) and BB/T014563/1 S., Burton, N., et al., The Virtual Fly Brain
(EMBL-EBI) from the UK Biotechnology and browser and query interface. Bioinformatics
2012, 28, 411–415.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>