<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Sound and Repeatable Approach to Building Integrated Repositories of Genomic Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anna Bernasconi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano</institution>
          ,
          <addr-line>20133, Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <fpage>20</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>PhD Thesis Award (Extended Abstract). The integration of genomic data and of their describing metadata is, at the same time, an important, dificult, and well-recognized challenge. It is important because a wealth of public data repositories is available to drive biological and clinical research; combining information from various heterogeneous and widely dispersed sources is paramount to a number of biological discoveries. It is dificult complex and there is no agreement among the various data formats, data models, and metadata definitions, which refer to diferent vocabularies and ontologies. It is bioinformatics community because, in the common practice, repositories are accessed oneby-one, learning their specific metadata definitions as result of long and tedious eforts, and such practice is error-prone; moreover, downloaded datasets need considerable eforts prior to insertion in analysis pipelines.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>multi-ontology knowledge base, that allows to locate relevant genomic datasets, which can
be then analyzed with of-the-shelf bioinformatics tools. This interface has been evaluated by
running an extended empirical study with participants knowledgeable in both Biology and
Computer Science, collecting many insights on the practices of diferent user profiles and on
their understanding of procedures for extracting datasets relevant for research.</p>
      <p>The models, frameworks and tools that are described in this thesis are already included in
follow-up projects; they can be exploited to provide biologists and clinicians with a complete
data extraction/analysis environment, equipped by a ‘marketplace’ of ready-to-use best practices.
The process may be guided by a conversational interface, which breaks down the technological
barriers that currently slow down the practical adoption of our systems. Our commitment is to
continue the inclusion of relevant data sources for bioinformatics tertiary analysis, continuously
improving the process from a data quality and interoperability point of view.</p>
      <p>Inspired by our work on genomic data integration, during the outbreak of the COVID-19
pandemic we searched for efective ways to help mitigate its efects; in this direction, we
successfully re-applied the model-build-search paradigm used for human genomics.</p>
      <p>Even if the domain of viral genomics is completely new, it presents many analogies with our
previous challenges. In this new context, we model viral nucleotide sequences as strings of
letters, with corresponding sub-sequences – the genes – that encode for proteins composed of
amino acids. To highlight diferences with previously considered data, we have designed the
Viral Conceptual Model [4] which accounts for their technological, biological and organizational
aspects, in addition to computed annotations and mutations on both nucleotides and amino acid
sequences. We then integrate sequences with their metadata from a variety of diferent sources
and propose the powerful search interface ViruSurf [5] (http://www.gmql.eu/virusurf/), able to
quickly extract sequences based on their combined mutations, to compare diferent conditions,
and to build interesting populations for downstream analysis. When applied to SARS-CoV-2,
the virus responsible for COVID-19, complex conceptual queries upon our system are able to
replicate the search results of recent articles, hence demonstrating considerable potential in
supporting virology research.</p>
      <p>This work has been realized during the first spread of the SARS-CoV-2 pandemic (March–
December 2020); after setting the first milestones, we are now moving forward, considering the
next challenges of this new domain with growing interest. These include the development of a
requirements elicitation technique for emergency times, the extension of ViruSurf to other types
of data (e.g., relevant for vaccine design), and the provision of visual and statistical support to
the integrated data.</p>
      <p>The results on this thesis are part of a broad vision: the availability of conceptual models,
related databases, and search systems for both human and viral genomics will provide important
opportunities for genomic and clinical research, especially if virus data will be connected to its
host, the human being, who is the provider of genomic and phenotype information.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <source>Proceedings International Conference ER 2017</source>
          , Springer, pp.
          <fpage>325</fpage>
          -
          <lpage>339</lpage>
          . [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <source>IEEE/ACM Trans. on Comput. Biol. and Bioinf</source>
          .
          <volume>19</volume>
          (
          <year>2022</year>
          )
          <fpage>543</fpage>
          -
          <lpage>557</lpage>
          . [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Canakoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <source>Database</source>
          , Volume
          <volume>2019</volume>
          ,
          <year>baz132</year>
          . [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <source>Proceedings International Conference ER 2020</source>
          , Springer, pp.
          <fpage>388</fpage>
          -
          <lpage>402</lpage>
          . [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Canakoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pinoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <source>Nucleic Acids Research</source>
          <volume>49</volume>
          (
          <year>2021</year>
          )
          <fpage>D817</fpage>
          -
          <lpage>D824</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>