<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Document-to-document relevance assessment for TREC Genomics Track 2005</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga Giraldo</string-name>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>María Fernanda Cadena</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Robayo-Gama</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dhwani Solanki</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tim Fellerhoff</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lukas Geist</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rohitha Ravinder</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muhammad Talha</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dietrich Rebholz-Schuhmann</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leyla Jael Castro</string-name>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bonn-Aachen International Centre for Information Technology (B-IT), University of Bonn</institution>
          ,
          <addr-line>Friedrich-Hirzebruch-Allee 6, Bonn, 53115</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Facultad de Farmacia y Bioquímica, Universidad de Buenos Aires</institution>
          ,
          <addr-line>Junín 956,Buenos Aires, C1113AAD</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Heinrich-Heine University Düsseldorf</institution>
          ,
          <addr-line>Universitätsstraße 1, Düsseldorf, 40225</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Hochschule Bonn-Rhein-Sieg</institution>
          ,
          <addr-line>Grantham-Allee 20, Sankt Augustin, 53757</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Institute of Molecular Medicine and Cell Research, University of Freiburg</institution>
          ,
          <addr-line>Stefan-Meier-Str. 17, Freiburg im Breisgau, 79104</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Cologne</institution>
          ,
          <addr-line>Albertus-Magnus-Platz, Cologne, 50923</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>ZB MED Information Centre for Life Sciences</institution>
          ,
          <addr-line>Gleueler Str. 60, Cologne, 50931</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Here we present a doc-2-doc relevance assessment performed on a subset of the TREC Genomics Track 2005 collection. Our approach includes an experimental set up to manually assess doc-2-doc relevance and the corresponding analysis done on the results obtained from this experiment. The experiment takes one document as a reference and assesses a second document regarding its relevance to the reference one. The consistency of the assessments done by 4 domain experts was evaluated. The lack of agreement between annotators may be due to: i) The abstract lacks key information and/or ii) Lack of experience of the annotators in the evaluation of some topics.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Relevance assessment</kwd>
        <kwd>literature manual curation</kwd>
        <kwd>document similarity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The TREC Genomics Track 2005 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] provides a collection of document-to-topic relevance
assessments for Medline abstracts. This collection has commonly been used also for
document-to-document related tasks [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2,3,4</xref>
        ]; however, to the best of our knowledge, no analysis has
been done regarding the suitability of such collection for such tasks, e.g., similarity or relevance
assessment between a pair of documents. Our doc-2-doc relevance analysis aims at filling this gap. We
take one document as a “reference article” while a second one is evaluated wrt its relevance to the
referenced document. In this experiment the user is engaged in the evaluation process as a way to
achieve better results.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <p>Selection of topics and documents. TREC 2005 Genomics track included a total of 50 topics. For the
document-to-document relevance assessment we wanted to have 15 documents to assess per each
reference document, at least 10 documents judged as definitely relevant wrt the TREC topic, and no
more than 80 relevant articles (either definitely or partially relevant). The goal was having a sample
covering about 10% of the relevant articles for the document-to-document assessment task. This gave
us a total of 16 topics which were further reduced to 8 topics due to time constraints and expertise of
the annotators on the different TREC topics. These topics contained a total of 42 reference documents
and 630 documents to be assessed.</p>
      <p>
        Development of an in-house annotation tool and a corpus of documents. The tool [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] presents a
corpus of documents organized by topics. The corpus is based on the TREC 2005 Genomics track
with some pre-processing to obtain the reference documents and the ones to be assessed against them.
The annotation workflow is shown on Figure 1.
Training sessions. Virtual sessions were organized with the participants in the evaluation of
documents to train them in the use of the tool, and to solve doubts. The meetings were carried out by
using Zoom. The participants were 4 domain experts with expertise in life science and/or
bioinformatics.
      </p>
      <p>Relevance assessment by domain experts. In the tool, the documents were organized by topics.
Where each topic includes “N” number of reference articles. Each reference article includes 15
documents (or evaluation articles), to be assessed. Only title and abstract are available. The relevance
assessment possible values are as follows: i) Relevant to the reference article, meaning “Yes, the user
wants to get a hold of the full-text as it is definitely relevant to their research”. ii) Partially relevant to
the reference article, meaning “Looks promising but not sure yet. The user will keep the PMID just in
case, as a maybe”. iii) Non-relevant to the reference article, meaning “Not worth giving it a second
look at all”.</p>
      <p>Analysis of results. Here the consistency of the assessments done by the 4 domain experts was
evaluated. We focused on inter-annotator agreement.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>
        Documents assessed. A total of 630 “evaluation articles” classified in 8 topics were assessed by 4
annotators. The evaluation articles are distributed into 42 reference articles (15 documents per
reference article). The full data is available online [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Inter-annotator agreement. Annotators rated the documents into three categories (2 definitely
relevant, 1 partially relevant, 0 non-relevant). We observed that the four annotators all rated as
“definitely relevant” 35 of the 630 assessed documents evaluated (5.56%). Similarly, the four
annotators agreed on the rating as “partially relevant” on 6 documents (0.95%); and others 123
documents (19.52%) were rated as “non-relevant”, giving us a total agreement among the four
annotators for 164 articles (26.03%). The Fleiss Kappa results are distributed into three levels of
agreement: “Poor”, with values from -0.1708 to 0.1885. “Fair”, with values from 0.2214 to 0.375; and
“Moderate” with values from 0.4564 to 0.5328. From the 42 reference articles, 24 got a Fleiss Kappa
corresponding to “Poor”, 14 corresponding to “Fair”, and 5 corresponding to “Moderate”. A table
summarizing the results about the inter-annotator agreement is available online [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion and conclusions</title>
      <p>This work is about the analysis done to the results obtained in an experiment focused on
evaluating the relevance between two articles. The experiment takes one document as a reference and
assesses a second document regarding its relevance to the reference one. The methodological aspects
involved the participation of four domain experts, who used an in-house annotation tool tailored to the
initial TREC data and the task at hand. The lack of agreement between annotators may be due to: i)
The abstract lacks key information. For example, the objective, main results or conclusions. In this
case, the reader has to search the entire document for the missing information. ii) Lack of experience
of the annotators in the evaluation of some topics. In this case, the reader must search for more
information on the web on a topic to better understand the document to be evaluated. Both
implications are time consuming. In order to overcome those limitations and improve the results, we
propose as a future work to extend the time required in the evaluation tasks and/or extend the number
of annotators to cover the lack of experience in some topics.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Acknowledgements</title>
      <p>This work is part of the STELLA project funded by DFG (project no. 407518790). This work was
supported by the BMBF-funded de.NBI Cloud within the German Network for Bioinformatics
Infrastructure (de.NBI) (031A532B, 031A533A, 031A533B, 031A534A, 031A535A, 031A537A,
031A537B, 031A537C, 031A537D, 031A538A).</p>
    </sec>
    <sec id="sec-6">
      <title>6. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Hersh</surname>
            <given-names>W</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhupatiraju</surname>
            <given-names>RT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hearst</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>TREC 2005 Genomics Track Overview</article-title>
          . :
          <volume>26</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Lin</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilbur</surname>
            <given-names>WJ</given-names>
          </string-name>
          .
          <article-title>PubMed related articles: a probabilistic topic-based model for content similarity</article-title>
          .
          <source>BMC Bioinformatics</source>
          .
          <year>2007</year>
          ;
          <volume>8</volume>
          :
          <fpage>423</fpage>
          . doi:
          <volume>10</volume>
          .1186/
          <fpage>1471</fpage>
          -2105-8-423
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Garcia</given-names>
            <surname>Castro</surname>
          </string-name>
          <string-name>
            <given-names>LJ</given-names>
            ,
            <surname>Berlanga</surname>
          </string-name>
          <string-name>
            <surname>R</surname>
          </string-name>
          ,
          <article-title>Garcia A. In the pursuit of a semantic similarity metric based on UMLS annotations for articles in PubMed Central Open Access</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          .
          <year>2015</year>
          ;
          <volume>57</volume>
          :
          <fpage>204</fpage>
          -
          <lpage>218</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.jbi.
          <year>2015</year>
          .
          <volume>07</volume>
          .015
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Wei</surname>
            <given-names>W</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marmor</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuo</surname>
            <given-names>TT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            <given-names>CN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohno-Machado L. Finding Related</surname>
          </string-name>
          <article-title>Publications: Extending the Set of Terms Used to Assess Article Similarity</article-title>
          .
          <source>AMIA Jt Summits Transl Sci Proc. 2016 Jul</source>
          <volume>20</volume>
          ;
          <year>2016</year>
          :
          <fpage>225</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Talha</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Geist</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fellerhoff</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravinder</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giraldo</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            <given-names>D</given-names>
          </string-name>
          , et al.
          <article-title>TREC-doc-2-doc-relevance assessment interface</article-title>
          .
          <source>Zenodo; 2022. doi:10</source>
          .5281/zenodo.7341391
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Giraldo</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solanki</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cadena</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robayo-Gama</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castro</surname>
            <given-names>LJ</given-names>
          </string-name>
          .
          <article-title>Document-to-document relevant assessment for TREC Genomics Track 2005</article-title>
          . Zenodo;
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.7324822
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Giraldo</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solanki</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castro</surname>
            <given-names>LJ</given-names>
          </string-name>
          .
          <article-title>Fleiss kappa for doc-2-doc relevance assessment</article-title>
          . Zenodo;
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.7338056
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>