=Paper= {{Paper |id=Vol-1276/MedIR-SIGIR2014-04 |storemode=property |title=Why Assessing Relevance in Medical IR is Demanding |pdfUrl=https://ceur-ws.org/Vol-1276/MedIR-SIGIR2014-04.pdf |volume=Vol-1276 |dblpUrl=https://dblp.org/rec/conf/sigir/KoopmanZ14b }} ==Why Assessing Relevance in Medical IR is Demanding== https://ceur-ws.org/Vol-1276/MedIR-SIGIR2014-04.pdf
      Why Assessing Relevance in Medical IR is Demanding

                          Bevan Koopman                                                    Guido Zuccon
         Australian e-Health Research Centre, CSIRO                             Queensland University of Technology
                      Brisbane, Australia                                               Brisbane, Australia
                  bevan.koopman@csiro.au                                              g.zuccon@qut.edu.au



ABSTRACT                                                                    Given the familiarity that medical experts have with med-
This study investigates if and why assessing relevance of                ical documents, one may posit that the task of assessing
clinical records for a clinical retrieval task is cognitively de-        relevance in these documents is not demanding for experts.
manding. Previous research has highlighted the challenges                On the contrary, our quantitative and qualitative analysis of
and issues information retrieval systems are faced with when             a relevance assessment exercise, performed by four experts,
determining the relevance of documents in this domain, e.g.,             revealed that assessing relevance in the medical domain is
the vocabulary mismatch problem. Determining if this as-                 often demanding: assessments required substantial time to
sessment imposes cognitive load on human assessors, and                  be formed, implying a substantial cognitive load on the as-
why this is the case, may shed lights on what are the (cogni-            sessors. Given this result, we explore and validate a number
tive) processes that assessors use for determining document              of factors associated to both queries and documents that
relevance (in this domain). High cognitive load may impair               contribute to the difficulty of the assessment task, revealing
the ability of the user to make accurate relevance judgements            why this task is demanding.
and hence the design of IR mechanisms may need to take
this into account in order to reduce the load.                           2.   EXPERIMENTAL DESIGN
                                                                            We used data gathered from a previous relevance assess-
Categories and Subject Descriptors: H.3 [Information                     ment task [1]. In this previous study, four medical profes-
Storage and Retrieval]: H.3.3 Information Search and Re-                 sionals were asked to judge clinical documents taken from
trieval                                                                  the TREC MedTrack collection [5]. As we used data from
General Terms: Experimentation.                                          an existing study not explicitly designed to fully answer the
                                                                         research questions of this paper, we are constrained by the
1. INTRODUCTION                                                          data captured in the previous study. Nevertheless, a num-
                                                                         ber of insights into how demanding assessment are can be
   The collection of relevance assessments is important for
                                                                         derived.
information retrieval (IR) systems evaluation. Relevance is
                                                                            The original TREC MedTrack queries were used and a
a complex notion: subjective to the person performing the
                                                                         total of 1030 documents were assessed.1 To collect assess-
assessment, dependent on contextual factors and often act-
                                                                         ments, the Relevation! judging system was used [2]. Queries
ing on multiple dimensions (i.e., factors like opinion, read-
                                                                         were divided between the four assessors with each query be-
ability and trustworthiness may influence a relevance judge-
                                                                         ing fully judged by only one assessor. Each assessor also
ment) [3]. To the best of our knowledge, however, there
                                                                         completed two control queries to familiarise themselves with
has been little or no work that investigates if and why it is
                                                                         the task. As all assessors completed the same control queries,
cognitively demanding for assessors to judge relevance.
                                                                         these were used to determine inter-coder agreement. The
   In this paper, we aim to determine: (i) if assessing doc-
                                                                         test queries were divided so that each assessor judged, in
ument relevance is demanding; if so (ii) what are the indi-
                                                                         total, roughly an equal number of documents. For each doc-
cators of a demanding assessment; and (iii) what are the
                                                                         ument, judges were asked to mark the document as “highly
reasons behind an assessment being demanding or not. To-
                                                                         relevant”, “somewhat relevant” or “not relevant” with respect
ward these aims, we focus on medical IR, and more specifi-
                                                                         to that query (as per TREC MedTrack guidelines). In ad-
cally on the task of finding patients suitable to clinical trials,
                                                                         dition, using Relevation!, assessors could provide a free-text
i.e., the task modelled in the TREC Medical Records Track
                                                                         comment regarding their decision. On completion of judg-
(MedTrack) [5]. It has been shown that this is, in general,
                                                                         ing all documents for a query, the assessor was also asked to
a difficult task for IR systems due to factors like vocabu-
                                                                         answer the following questions about the query:       1) “How
lary and granularity mismatch, conceptual implication, and
                                                                         difficult was this query to judge?”. Choices: “Very difficult”,
inferences of similarity [1]. However, no previous work has
                                                                         “Moderately difficult” or “Easy”. 2) “How would you rate the
explored whether this also applies for humans, and whether
                                                                         quality of the assessments you have provided for this query?”
assessing the relevance of health records for this task is cog-
                                                                         Choices: “High quality”, “Average in quality” or “Poor qual-
nitively demanding (indeed, difficult) for expert assessors.
                                                                         ity”. 3) “Other comments?” Here judges could provide qual-
                                                                         itative comments regarding the particular query.
Copyright is held by the author/owner(s).                                1 4 queries were excluded from the original 85 TREC MedTrack
MedIR July 11, 2014, Gold Coast, Australia.
ACM SIGIR.                                                               queries as no relevance assessments were collected for these.




                                                                    16
                                                                                                                Difficulty                                                      Quality
   As Relevation! is a web-based system, the HTTP access




                                                                                           20 40 60 80




                                                                                                                                                           20 40 60 80
                                                                       Number of queries




                                                                                                                                       Number of queries
log was used to capture the interaction assessors had with
the system. This included which queries and documents
they viewed, when documents were judged and, importantly,
the timestamps for these events. These timestamps were




                                                                                           0




                                                                                                                                                           0
used to extract the amount of time each assessor spent in
                                                                                                         Easy   Mod. Hard Very Hard                                       Low    Avg.     High
judging individual documents.2 The difference in time be-
tween two consecutive HTTP POSTs was used as the mea-                  Figure 1: Judges’ qualitative feedback on difficulty
sure of time it took to judge that document. On manual                 and quality of their assessments.
review, any time periods greater than 2500 seconds (42 min-
utes) was indicated as a break (e.g., lunch or coffee) and                                  Difficulty                           #Queries                                   Median sec./doc
these timings were excluded. Note that qualitative feedback
from assessors (e.g., difficulty and quality) were collected                                 Easy                                       44                                 130sec
at query level, while quantitative statistics such as time to                              Moderately Difficult                         36                                 207sec (+59%)
perform a judgement were collected both at query and doc-                                  Very Hard                                  1                                  219sec (+68%)
ument level.
   A total of 58 hours (14.5 hours per assessor) of judging                                              Table 1: Timing results by difficulty.
was required to complete the 942 documents.3 The average
time spent per document was 3.7 minutes. Using the control             ilarly, the longer it took on average to judge documents for
queries, inter-coder agreement was found to be 0.85, in-line           a query, the more demanding that query.
with an inter-coder agreement of 0.8 found by the TREC                    The use of time as an indicator of assessment demand is
MedTrack organisers.4 Control queries also contained doc-              confirmed by the results of Table 1 that shows the judges’
uments already judged by TREC assessors; therefore, if the             qualitative feedback about query difficulty along with the
TREC assessor is added as a fifth assessor, then agreement              median document judging time for each difficulty level. This
between all five assessors was 0.80.                                    analysis shows that queries judged as moderately difficult
                                                                       took 59% longer to judge than those marked easy, endorsing
3. IS ASSESSING RELEVANCE                                              the intuition that time is a (fine grain) indicator of assess-
   DEMANDING?                                                          ment demand.
  To determine if and why assessing relevance is demanding
we analysed: (i) qualitative feedbacks given by assessors in           4.                          WHAT INDUCES COGNITIVE LOAD?
relation to the assessment difficulty of each query; and (ii)
the amount of time required to judge documents.                        4.1 Are longer documents harder to judge?
                                                                         Smucker & Clarke found that in web search, judging time
3.1 Did assessors find judging difficult?                                was mainly influenced by document length [4]. Document
   Assessors rated each query according to how difficult it              length was, therefore, used as the main indicator for their
was to judge and further provided a self-assessment of the             time-biased evaluation measure [4].
quality of their judgements. Results are shown in Figure 1.              In our study, if document length was also a measure of de-
Assessors stated that about half of the queries were easy to           mand, then the Easy/Mod/Hard label assigned by assessors
assess, with the remaining half being of moderate difficulty.            would simply relate to short, moderate and long documents
Only one query was considered very difficult to judge.5 Nev-             respectively. By extension, shorter documents would be less
ertheless, the assessors believed the judgements they pro-             demanding to judge. However, this was not found to be
vided were of average or high quality. (No queries were                the case: there was no correlation between time to judge a
marked as low quality.)                                                document and the length of the document (p = −0.0132).
   While these qualitative assessments are ultimately subjec-
tive (the self-perception of difficulty and quality may vary             4.2 Are documents with discharge summaries
between assessors), it is clear how a significant number of                 easier to judge?
assessments was perceived to be more demanding than oth-
                                                                         Many of the clinical documents used in our collection con-
ers.
                                                                       tained a discharge summary section.6 Assessors commented
3.2 Time as indicator of demand                                        that they often skimmed the document looking for a dis-
                                                                       charge summary section to read first rather than reading the
  Beside examining the qualitative feedback of the difficulty            document from top to bottom. Sometimes the relevance of
in assessing documents, we also consider time as an indica-            a document could be determined from reading the discharge
tor of judging demand. The intuition is that documents that            summary alone.7 Based on these comments, we formed the
required more time for assessment are more demanding; sim-             hypothesis that documents containing a discharge summary
2 The HTTP log is available online at:                                 would be quicker and less demanding to judge. However,
https://github.com/ielab/MedIR2014-RelanceAssessment                   our results show the contrary: the median time to judge a
3 This number excludes documents from the control queries and          document with a discharge summary was 184 sec., vs. 118
those which took more than 2500 seconds to judge (i.e., where          sec. for documents without a discharge summary.
the assessors was deemed to have taken a break).
4 Based on personal communication with Bill Hersh, TREC Med-           6 A discharge summary is a narrative produced when a patient
Track organiser, 29 May 2013.                                          is discharged from hospital. Discharge summaries provide an
5 Query 149: “Patients with delirium hypertension and tachycar-        overview of the patient’s entire stay in hospital.
dia”.                                                                  7 Note that not all documents contained discharge summaries.




                                                                  17
                              Documents                         Time to judge (seconds)                              of relevance was clear and explicit according to the assessor;
                                                               mean stddev max min                                   (ii) “temporal”, where relevance was strongly dependent on
                                                                                                                     temporal aspects (of query or documents); (iii) “interpreta-
                              non-relevant                       219               191       1614           5        tion”, where the interpretation of the query was subjective
                              relevant                           224               221       2092          26        and the assessor had to decide on a particular interpreta-
                               highly-relevant                   167               209       2092          26        tion; and (iv) “dependent aspects”, where there were two or
                               somewhat-relevant                 289               217       1314          60        more conditions specified in the query — often dependent on
                                                                                                                     each other — that had to be met. Queries not exhibiting any
                              Table 2: Timing results by relevance grade.                                            of the aforementioned characteristics were characteristed as
                                                                                                                     “none”. Note, these characteristics were derived from the
                                                                                                                     relevance criteria, as stated in the assessor’s comments, and
4.3 Is the grade of relevance related to cogni-                                                                      not according to the query keywords. Queries were grouped
    tive load?                                                                                                       according to these characteristics and we analysed the aver-
   Does the relevance grade of a document (i.e. highly rel-                                                          age time to judge the documents for queries with that char-
evant, relevant, not relevant) affect how demanding it is to                                                          acteristic. This is done to understand if some characteristics
judge? Table 2 shows the time it takes to judge documents                                                            — and therefore some queries — were more demanding than
according to the relevance grade. When considering only bi-                                                          others. The average time to judge according to each char-
nary relevance (i.e., relevant vs. non-relevant), the average                                                        acteristic is shown in Figure 2.
time to judge relevant and non-relevant documents does not                                                              Those queries identified as “none” (n=34, 60%) required,
differ significantly, although the time to judge relevant doc-                                                         on average, the least assessment time and were the least
uments varies more (stddev); both the maximum and min-                                                               demanding. Queries identified as “objective” (n=12, 21%)
imum judging time are greater for relevant documents. In                                                             were marginally more demanding, as the assessor had a clear
contrast, when graded relevance is considered, some impor-                                                           criteria to identify relevance and all that was required was
tant differences are revealed: highly relevant documents are                                                          to assert if that criteria applied to the particular document.
the least demanding to judge, whereas somewhat-relevant
documents are the most demanding to judge. This finding                                                               5.1 The effect of temporality on relevance
suggests that clear cases of relevance (highly relevant or non-                                                         For “temporal” queries (n=10, 18%), the assessors specifi-
relevant) are less demanding. What is demanding is judging                                                           cally cited temporality as an important factor in determining
documents where relevance is less certain: cases where rel-                                                          relevance. The most common situation was when informa-
evance is subjective or where the evidence for relevance is                                                          tion pertaining to the query was found in the patient’s past
implicit and needs to be inferred. We explore more of these                                                          medical history section. Assessors had to decide whether
situations in the following section by analysing the assessors                                                       the information was still valid: some conditions are ongoing
qualitative feedback.                                                                                                (e.g., query 162, Table 3), while others are temporal and
                                                                                                                     are unlikely to still be valid (e.g., query 127). In certain
5. WHY IS ASSESSMENT DEMANDING?                                                                                      cases, assessors consulted the actual dates of the past med-
   On completion of judging a query, assessors could option-                                                         ical history information to determine how recent the infor-
ally provide free-text, qualitative comments regarding their                                                         mation was and whether it might still apply. In other cases,
judging of the particular query. Assessors provided these                                                            the query was interpreted according to a temporal defini-
comments for 57 out of 81 (70%) queries. We analysed their                                                           tion (e.g., query 111, where the assessor defined ‘chronic
comments to gain a greater insight into their rational for                                                           back pain’ as a condition persisting for at least 3 months).
assessment and to determine why it might be demanding.                                                               Queries exhibiting temporality tended to be the most de-
Table 3 contains a selection of assessor’s comments which                                                            manding as assessors had to locate and reason with dates
will be referred to throughout this section. The assessors’                                                          found in the documents.
comments were used to identify queries exhibiting the fol-
lowing characteristics: (i) “objective”, where the indicator                                                         5.2 Judging was highly subjective
                                                                                                                        For “interpretation” queries, assessors, at times, discussed
                                                                                                                     their decisions regarding relevance. Although confident in
Avg. sec. to judge document




                                        ●                                                                            their assessments, they stated that the interpretation of the
                              400
                                                                                                                     query was subjective and often required careful considera-
                              300
                                                                                                                     tion regarding different possible interpretations. For exam-
                                                                                                                     ple, for query 101, assessors debated whether a patient born
                              200                                                                                    deaf could be considered as exhibiting hearing loss. (Tech-
                                                                                                                     nically, if they never had any hearing, then they never had a
                              100                                                                                    loss of hearing.) One assessor thought such a document was
                                                                                                                     relevant, while another assessor thought the document was
                                0     n=34     n=12                    n=10          n=14           n=35             not relevant. A medical encyclopaedia was consulted and
                                                                                                                     the assessor decided to include patients born deaf as rele-
                                       None




                                                   Objective




                                                                        Temporal



                                                                                     Interpre−




                                                                                                      Aspects
                                                                                         tation


                                                                                                    Dependent




                                                                                                                     vant. Queries requiring subjective interpretation showed a
                                                                                                                     higher level of demand compared to other queries.
                                                                                                                        The task description given to assessors (recruitment of
Figure 2: Average time to judge the documents for                                                                    patients matching a certain inclusion criteria for clinical
queries with different characteristics. Queries re-                                                                   trials [5]) also affected their decisions regarding relevance.
quiring some “interpretation” on the part of asses-                                                                  Certain documents described patients who had hearing loss
sors were the most demanding.




                                                                                                                18
 Query                                                         Assessors’ Comment

 101    Patients with hearing loss                             It was not clear whether you wanted someone with current hearing loss or
                                                               someone who had experienced reversible hearing loss due to an infection.
 102    Patients with complicated GERD who receive en-         Complicated GERD is a rather ambiguous term - could use clarification to
        doscopy                                                yield better results (ex. stage a/b/c). Endoscopy is a blanket term for visu-
                                                               alisation of a hollow organ - therefore some search results included patients
                                                               who have had colonoscopies, but not upper endoscopies relevant to GERD.
 103    Hospitalized patients treated for methicillin resis-   Treatment of MRSA is the same no matter where it is in the body. Could
        tant Staphylococcus aureus MRSA endocarditis           have picked up a lot of documents because of the treatment regime or MRSA.
 111    Patients with chronic back pain who receive an         The definition of chronic back pain used for these judgements was “greater
        intraspinal pain medicine pump                         than 3 months”
 127    Patients admitted with morbid obesity and sec-         Without dates, it was difficult to ascertain whether or not hypertension and
        ondary diseases of diabetes and or hypertension        diabetes were secondary to patients’ obesity, as is suggested by the query.
 162    Patients with hypertension on antihypertensive         Once diagnosed with hypertension, you are generally considered to have it
        medication                                             for the rest of your life ...
 171    Patients with thyrotoxicosis treated with beta         A lot of hits for beta blockers and very few for any thyroid dysfunction.
        blockers
 182    Patients with Ischemic Vascular Disease                Straightforward to look at past medical history for coronary artery disease,
                                                               bypass grafts or stents.

       Table 3: Assessors’ qualitative comments regarding their experience judging the particular query.

on admission but the hearing loss was treated and resolved                 between the dependent aspects. Doing so required longer
by discharge. In this case, assessors decided these patients               judging times and was, therefore, more demanding.
would not be eligible for the clinical trial and, therefore, not
relevant to the query. For other tasks (for example, finding                6.   CONCLUSION
how hearing loss is treated) these documents may have been                    Assessing relevance in medical IR is sometimes cognitively
highly relevant. These cases highlight the complex and of-                 demanding and that demand differs depending on queries.
ten subjective nature of information need in this domain and               Contrary to intuition and previous studies in other domains
that there are often implicit factors in the information need              [4], this study found that document length does not influence
that do not transpire in the query. This further adds to the               demand. On the other hand, the grade of relevance is related
demand of relevance assessment for these types of queries.                 with cognitive load (somewhat relevant documents were the
                                                                           most demanding to judge). Characteristics of queries that
5.3 Queries with dependent aspects                                         did increase demand included: temporality, subjectiveness
   Queries with multiple “dependent aspects” received more                 of interpretation and the presence of multiple dependent as-
debate by assessors and were also among the most demand-                   pects in the query.
ing and those with the highest variance in judging time.                      A by-product of this study on what makes a relevance
The high variance in time to judge a document is due to                    decision demanding, is the identification of some of the as-
the fact that queries with dependent aspects were either:                  pects that influence a relevance decision (for example, the
(i) simple to judge, because the assessor just had to ascer-               role of temporality). Future work would, therefore, consider
tain that a document met all aspects; or (ii) demanding to                 the actual features of the document (for example, temporal
judge, because the assessor had to determine the interaction               ranges or chronic vs. acute conditions) that identify these
between the required aspects. Query 171 is an example of                   different aspects affecting relevance.
the former, simple case. Query 102 is an example of the                       Data used in this study, including the HTTP interaction
latter case: GERD8 is a common condition and is therefore                  log, assessors’ comments and qrels, is provided at:
                                                                           http://github.com/ielab/MedIR2014-RelanceAssessment.
found in many patients’ records. The difficulty in interpret-
ing this query was whether the endoscopy was performed                     Acknowledgements. The authors are grateful to Peter Bruza
because of the GERD or for some other, unrelated condi-                    for his continued mentorship. The relevance assessments were
tion. There were a number of documents where patients                      conducted by Timothy Sladden, Warren Brown, Digvijay Khangarot
                                                                           and Thomas Souchen, from the University of Queensland.
had GERD but received the endoscopy for another reason;
these were marked as not relevant. A similar query was                     7.   REFERENCES
103, where endocarditis and MRSA were mentioned in the                     [1] Bevan Koopman. Semantic Search as Inference:
same document, but the cause of the endocarditis was not                       Applications in Health Informatics. PhD thesis, Queensland
the MRSA. Again, these documents were marked as not rel-                       University of Technology, 2014.
evant. These queries all have multiple dependent aspects                   [2] B. Koopman and G. Zuccon. Relevation!: An open source
to the query; even if both aspects are present in a docu-                      system for information retrieval relevance assessment. In
                                                                               SIGIR Demo, Gold Coast, Australia, July 2014.
ment, that document may still not be relevant unless the
                                                                           [3] S. Mizzaro. Relevance: The whole history. JASIST,
dependence between them can be determined. Determining                         48(9):810–832, 1997.
the dependence often required the assessors to exhaustively                [4] M. D. Smucker and C. L. Clarke. Time-based calibration of
search through the document to identify the relationships                      effectiveness measures. In Proc. of SIGIR, pages 95–104,
                                                                               Portland, U.S.A, 2012.
8 Gastroesophageal reflux disease (GERD) is caused when stom-               [5] E. M. Voorhees and W. Hersh. Overview of the trec 2012
ach acid comes up from the stomach into the esophagus.                         medical records track. In Proc. of TREC, 2012.




                                                                      19