<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <article-id pub-id-type="doi">10.1145/1963405.1963424</article-id>
      <title-group>
        <article-title>Learning Profile-Based Recommendations for Medical Search Auto-Complete</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guusje Boomgaard</string-name>
          <email>g.boomgaard@student.vu.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Selene Báez Santamaría</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilaria Tiddi</string-name>
          <email>i.tiddi@vu.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Jan Sips</string-name>
          <email>r.sips@mytomorrows.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zoltán Szlávik</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>In A. Martin, K. Hinkelmann</institution>
          ,
          <addr-line>H.-G. Fill, A. Gerber, D. Lenat, R. Stolle, F. van Harmelen (Eds.)</addr-line>
          ,
          <institution>Proceedings of the AAAI 2021 Spring Symposium on Combining Machine Learning and Knowledge Engineering (AAAI-MAKE 2021) - Stanford University</institution>
          ,
          <addr-line>Palo Alto, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vrije Universiteit Amsterdam</institution>
          ,
          <addr-line>De Boelelaan 1111, 1081 HN, Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>myTomorrows</institution>
          ,
          <addr-line>Anthony Fokkerweg 61, 1059 CP, Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <abstract>
        <p>Query popularity is a main feature in web-search auto-completion. Several personalization features have been proposed to support specific users' searches, but often do not meet the privacy requirements of a medical environment (e.g. clinical trial search). Furthermore, in such specialized domains, the diferences in user expertise and the domain-specific language users employ are far more widespread than in web-search. We propose a query auto-completion method based on diferent relevancy and diversity features, which can appropriately meet diferent user needs. Our method incorporates indirect popularity measures, along with graph topology and semantic features. An evolutionary algorithm optimizes relevance, diversity, and coverage to return a top-k list of query completions to the user. We evaluated our approach quantitatively and qualitatively using query log data from a clinical trial search engine, comparing the efects of diferent relevancy and diversity settings using domain experts. We found that syntax-based diversity has more impact on efectiveness and eficiency, graph-based diversity shows a more compact list of results, and relevancy the most efect on indicated preferences.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Query Auto-Completion</kwd>
        <kwd>Medical Information Retrieval</kwd>
        <kwd>Knowledge Graphs</kwd>
        <kwd>Professional Search</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>context, location) may be unwanted, especially when dealing with sensitive information or
when unbiased suggestions are required. Secondly, diferences in user expertise are far more
widespread than in web-search, i.e. a system often has to answer needs of patients, healthcare
providers and pharmaceutical professionals at the same time. Thirdly, the language of diferent
users is also diverse, thus raising problems with processing the user input as well. For example,
medical specialists use diferent language than laymen [7], similarly to native vs. non-native
English speakers [8]. These diferences in user population come with diferent requirements,
which may be hard to tackle with a single solution.</p>
      <p>This research aims to improve the disease auto-completion process of a clinical trial search
engine at the e-Health company myTomorrows1. The existing QAC method incorporates string
similarity, Pubmed and clinical trial statistics, and string length. However, the provided
suggestions sufer from redundancy; and due to the vast amount of matches to any short prefix, there
is a need for an intelligent selection and ranking of suggested terms. In this paper, we propose
to improve the QAC method using a graph-based taxonomy of medical conditions. This
structured knowledge source combines semantic information, corpus statistics and graph topology,
allowing us to study how diferent types of relevancy and diversity may aid in avoiding the
common problem of suggestion redundancy [9] and to support diferent user profile needs. We
evaluate the efectiveness (recall) and eficiency (tokens saved rate) of each method in finding
the intended suggestion with the goal of understanding how to support diferent user profiles.
We also evaluate the set of suggestions presented to the user, both in set length and coverage.
As a result, we provide alternative approaches for frequency-based and personalized methods,
and recommending diferent versions based on the requirements for various user profiles.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>In this section, we present generic literature on web and professional user profiling, then
focusing on the medical domain. We then present methods to improve auto-complete suggestions,
i.e. combining relevancy with diversity, individual tokens in queries, and knowledge graphs.
Web vs. professional search. With the increasing usage of the web, a large body of work
has focused on the analysis of query logs both to profile single users and cohorts [10, 11]
in web search contexts, revealing that (a) additional external knowledge about users and the
search corpus can be relevant for personalization, and (b) diferent techniques can be relevant
for diferent target type. Several studies [12, 13] have investigated the search practices and
preferences in diferent specific domains (e.g. legal, recruitment, academia, healthcare
professionals), showing that challenges such as boolean query formulation, the need of knowledge
management and sharing across searches, and the ambivalence of relevance ranking are
common despite the domain diferences.</p>
      <p>Various studies have investigated how to automatically identify medical experts and
laypeople using query log data. White et al. [14] developed a general model to predict whether users
were domain experts in four diferent domains, namely medicine, finance, law, and computer
science. Knowledge and usage of Pubmed was identified as a salient feature for medical experts.
Similarly, Palotti et al. [15] estimates medical expertise by using two query log sources aimed
1https://search.mytomorrows.com/search
20 terms</p>
      <p>Evolutionary algorithm
Optimize on:
1) Relevance
2) Diversity
3) Coverage</p>
      <p>Top-k disease terms</p>
      <p>III. Optimization
at diferent audiences to diferentiate between medical experts and laypeople. Pang et al. [16]
also show that people searching for health-related topics often engage in a more exploratory
type of search (e.g. visiting multiple websites) in an efort to fill knowledge gaps and receive
hints for correct spelling.</p>
      <p>Improving recommendations. Relevance and diversity have been identified as
requirements for high-quality recommendations [17]. Relevancy in QAC is often defined as an item
being popular, since the historical frequency is often a good predictor of the likelihood that
the item will be searched again in the future, making MPC [1] a widely used approach.
However, MPC tends to overlook long-tail queries and causes redundancy in the list of results [18].
Diversity can be introduced as the counterpart to relevancy, following a Goldilocks principle.
This equilibrium is closely related to the balance achieved between exploration, which
typically favors long and diverse lists, and exploitation, which is heavily relevance focused. Some
studies investigate approaches to combine these bi-criteria [9], a more efective tri-criteria
approach is presented by Zhong et al. [19], where local diversity (i.e. the dissimilarity between
top-k returned items) is distinguished from global diversity (i.e. how many diferent relevant
non-returned items are similar to at least one of the returned items).</p>
      <p>Term importance. Given a user’s input prefix (e.g. di-), individual tokens (e.g. a single
word like disease) may be relevant to a diferent degree (e.g. looking for diabetes, as opposed
to pulmonary disease). In Groza and Verspoor [20], term importance is applied to improve
biomedical concept recognition in texts, using concepts from the UMLS Metathesaurus to
create a representation of a document2. This document is then used to determine the information
gain of specific terms by calculating a Divergence From Randomness score for individual terms.
Semantics. Structured knowledge in the form of (knowledge) graphs allow to identify
complex semantic relations, such as concept similarity [21], along with the relational knowledge
represented in traditional databases. Concept similarity can serve as an important feature for
diversifying query results. Additionally, graphs can accommodate search by synonyms [22].</p>
    </sec>
    <sec id="sec-3">
      <title>3. Approach</title>
      <p>Our approach produces auto-complete suggestions optimized across three dimensions:
relevancy, local diversity, and global diversity (or coverage). In order to produce a top-k list of
results, the auto-complete suggestions are matched, ranked, and selected in a three-phased
2A collection of various source vocabularies such as ICD10, MeSH and SNOMED CT.</p>
      <p>Retrieval
1. Breast carcinoma
2. Breast cancer stage II
3. Recurrent breast cancer
4. Stage II breast carcinoma AJCC V7
5. Invasive breast carcinoma
6. Hereditary male breast carcinoma
7. Terminal cancer
8. Female breast cancer</p>
      <p>Filtering
1. Breast carcinoma
2. Female breast cancer
3. Breast cancer stage II
4. Recurrent breast cancer
5. Terminal cancer
6. Stage II breast carcinoma AJCC V7</p>
      <p>
        Optimization
1. Female breast cancer
2. Breast carcinoma
3. Recurrent breast cancer
4. Breast cancer stage II
process (Figure 1). First, in the retrieval phase (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), all disease terms that match the prefix are
retrieved from the database. Subsequently, in the filtering phase (2), this set is reduced by
taking the 20 most relevant items in order to limit the computational load. In the final
optimization phase (3), the optimal subset, according to all three dimensions, is selected using an
evolutionary algorithm.
      </p>
      <p>Let us consider the user input: Breast ca-. In the retrieval phase, the retrieved keywords
would include {Breast, Cancer, Carcinoma}, while the returned disease names are shown in
Figure 2. Some disease terms (e.g. Breast carcinoma) are more likely to be searched for (i.e.
more relevant) than others (e.g. Hereditary male breast carcinoma). During the filtering
phase, less relevant terms would be ranked at a lower position, below the filtering threshold.
Note that Breast cancer stage II and Stage II breast carcinoma AJCC V7 are related
terms in the UMLS taxonomy, the latter being a subtype (i.e. a ‘child’) of the former.</p>
      <p>In the optimisation phase, the system identifies this and decide to return only Breast cancer
stage II, as it implicitly covers its children, too. Consider now the case in which only higher
level candidates such as Breast carci- noma and Terminal cancer are kept. The results
would not provide diversity high enough, causing specific but desired disease terms (e.g. Male
breast cancer) to be omitted. Our optimisation step aims to balance such considerations,
resulting in a list of relevant, diverse, and high-coverage terms.</p>
      <sec id="sec-3-1">
        <title>3.1. Data Preprocessing</title>
        <p>
          In order to implement all dimensions in our algorithm and to enable fast retrieval of items, the
graph data needs to be pre-processed. Features such
as TF-IDF, trial, and paper counts are generated,
normalized, and combined into a various scores. Then,
keywords are created for every disease node (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) to
ensure word-order independent matching of the prefix to
a disease term, and (2) to calculate TF-IDF scores for
each disease keyword.
        </p>
        <p>Knowledge Graph. Our knowledge graph (stored Figure 3: Knowledge graph schema.
both as a SQL and a Neo4J database) contains
information about disease concepts derived from the diferent medical vocabularies provided by
UMLS. A disease concept consists of a preferred disease term and its synonyms. The graph
includes three types of nodes (see Figure 3). Disease nodes are connected through has_child
relationships (as per UMLS standards), and Disease_name nodes are connected to their
corresponding Disease nodes through has_alternative_name relationships. After the creation
of the keywords, we link each Keyword node to their corresponding Disease nodes using the
has_keyword relationship. Keywords are generated for each disease concept by tokenizing the
preferred disease term and its synonyms. Tokens were normalized to American English.
Features. The features to implement diferent dimensions can be divided into individual
features (i.e. related to individual disease concepts) and group features (i.e. an aggregated score
of a group of disease concepts). We use [R], [D], and [C] to indicate whether they refer to
relevancy, diversity, or coverage, respectively.
∙ Clinical trial count (TC) [R] refers to the total count of clinical trials related to a disease
concept, standardized using a sigmoid transformation around the median. The NER system
QuickUMLS3 was used to detect disease concepts from the title, keywords and conditions
sections of clinical trials, collected from the oficial repositories from the Unites States and
Europe4.
∙ Paper count (PC) [R] is computed per disease concept by processing title and abstract text
from Pubmed articles and standardized using the sigmoid transformation.
∙ TF-IDF [R] was calculated for each has_keyword relationship using tf-idf = tfn ∗  ( /df),
where tfn is the normalized term frequency, N the number of disease concepts, and df the
document frequency of a keyword. The score is also standardized as above.
∙ Children count (Ch) [R] represents the number of child nodes connected to a disease node,
divided by 100, where a limit is set to a maximum value of 1.
∙ Graph-depth [D] is computed per disease node as the median of the graph depths of
its originating sources. Recall that our graph is an aggregate of multiple sources, so disease
concepts can have multiple graph-depth values according to their taxonomy of origin.
∙ Concept similarity [D] is measured using the average graph distance, calculated by taking
the average of pairwise shortest-path distances of a set of disease nodes, as in [21].
∙ Covered items [C] exploits the hierarchical relationships in the graph to compute the number
of children of a returned item. Given a set of relevant items, we compute the amount of
relevant non-returned items covered by the returned items.</p>
        <p>The above features contribute to one of the following score variants (ranging [0,1]):
1. Basic Relevance5 [R], which combines the TC and PC features as Basic =   +   .
2
2. TF-IDF Relevance [R], which combines the former two with the TF-IDF feature as
TF-IDF = (  +   )∗(    +1) ).</p>
        <p>4
3https://github.com/Georgetown-IR-Lab/QuickUMLS
4https://clinicaltrials.gov/ and https://www.clinicaltrialsregister.eu/
5 All aggregated relevance scores were determined by summing the scores while applying a discount for higher
raarenkosf:t e n not looke=d∑at6=b1 y t1he ∗us er [6].  . The relevance scores of items below rank 6 were not counted as these
4. Distance Diversity [D], is calculated by taking the average of all pairwise distances, where
 +1) .</p>
        <p>distance is measured in number of edges, as in [21].
5. Specificity Diversity [D], through graph depth is rewarded by calculating the diversity by
dividing the unique depth values by the total number of depth values: ℎ
, as in [21].
6. Coverage [C], We use the Covered items feature as the coverage score of a set of items by


counting the number of items covered by the returned set, divided by 10.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Query Auto-Complete</title>
        <p>Prefixes are matched to disease nodes through one or more keyword nodes (fig. 2). The retrieval
algorithm matches the prefix with keywords, which are linked to disease nodes. In the filtering
phase, this set is reduced to a smaller subset in order to limit the computational load. This is
done by taking the 20 most relevant items.</p>
        <p>
          In the optimization phase, a set of
disease terms (from now on called items)
is selected, such that they are (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
relevant, (2) diverse, and (3) cover-relevant.
        </p>
        <sec id="sec-3-2-1">
          <title>The problem is approached as a multi</title>
          <p>objective optimization task.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Evolu</title>
          <p>tionary algorithms have shown to
perform well at similar query
recommendation problems when the search space
is large, such as the selection of topical
queries [23]. Therefore, a genetic
algo</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Evaluation</title>
      <sec id="sec-4-1">
        <title>4.1. Experimental Set-Up</title>
        <p>Experimental settings.</p>
        <p>We carried out experiments with historical query data to evaluate
the eficiency and efectiveness of our algorithm. To investigate the role of diferent feature
seSpecificity  ’.
tups, we compare three variants of relevancy and two of diversity: ‘Basic - Distance ’, ‘TF-IDF

- Distance ’, ‘Semantic - Distance ’, ‘Basic - Specificity  ’, ‘TF-IDF - Specificity  ’, ‘Semantic
Baseline. As baseline, we take the current disease QAC employed in myTomorrows’s clinical
trial search engine. This depends on three resources: (a) a mapping of n-gram prefixes to
disease name tokens, (b) a mapping of disease name tokens to disease names, and (c) a concept
relevance score per disease node in UMLS. The baseline consists of four phases:
1. Individual string matching: First, each n-gram in the user input (as separated by spaces)
is matched to a list of tokens, which is associated to a list of disease name candidates. We
reward user-input identified as a valid disease name tokens (e.g. BRE, standing for Benign
rolandic epilepsy), and penalize them otherwise. The score is then ind_score = 

(
(
)
) ∗
 , where  is 1.05 for rewards and 0.7 for penalties.
2. Aggregated string matching: Since diferent input n-grams may lead to the same
disease name candidate (i.e. Breast and Canc are both associated with Breast Cancer), an
aggregation step is needed to provide with a unique list of disease name candidates. The
aggregated score is calculated as agg_score = log
(
)+1(∑(
only candidate names scoring above a certain threshold (set at 0.05) are further processed.
_

) + 1). At this point,
3. Semantic relevance: The list of disease names is queried against UMLS to retrieve
Concept Unique Identifiers (CUI) and their relevance. The specifics of the relevance metric fall
beyond the scope of this paper, but they are similar to the Basic Relevance score (Section
3.1). To ensure unique concepts are shown to the user, we remove any duplicate concepts,
while keeping the disease name with the highest string matching score.
4. Ranking: The list of suggestions are ranked according to the average between the string
match score and the concept’s relevance i.e. score = 
items are shown to the user6.
_

+
2
_

. Finally, the top-10
Data. For the quantitative evaluation, myTomorrows provided the 500 most popular queries,
consisting of anonymized searched and selected disease terms. The list contained 409 distinct
terms. For the qualitative evaluation, the 18 most queried terms from our query log were used.
For each query, the corresponding selected term was treated as the intended search term.
Evaluation.</p>
        <p>Via two quantitative experiments, we compare recall, precision and eficiency
scores. In Experiment 1, eficiency was measured as Tokens Saved Rates (TSRs) by
increasing the number of characters of the query entered into the QAC until the clicked term was
included in the results. To evaluate how diferent methods ranked the intended item,
Experiment 2 compares the item’s rank after input of diferent query lengths (2, 4, 6, 8, 10 characters).
Results were compared by performing pairwise t-tests, with Bonferroni correction applied to
accommodate for multiple testing.</p>
        <p>
          Experiment 3 consisted of an ofline user-based evaluation. Due to accessibility, target users
(experts and non-experts in the medical domain) were simulated by myTomorrows employees
6Note that this QAC method was created through informal experimentation, and its behavior has not been
thoroughly studied, hence motivating the current work.
with a similar split in profiles. Each participant performed 10 tasks per round, and on
average completed 22.5 comparison tasks per person. In each task, participants were shown two
images with auto-complete results (Figure 4), and were asked to (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) indicate their method
preference, and (2) briefly motivate their decision. Forcing participants to make a choice between
two shown lists allowed us to make preferences more explicit, and method diferences more
detectable, similarly to how pairwise preference elicitation works for recommender systems [24].
        </p>
        <p>For each query, a prefix was constructed varying the lengths between queries. In the first 6
conditions, each method is compared to the baseline. Additionally, to evaluate the diferences
within our methods, another 9 combinations were compared: 6 to compare three relevance
scores within each diversity score and 3 to measure the efect of each diversity score within
each relevance method.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Experimental Results</title>
        <p>TF-IDF</p>
        <p>Experiment 1.
TSR results show that Distance
performs similarly to the
baseline, while other methods are
showing a significantly lower TSR.</p>
        <sec id="sec-4-2-1">
          <title>Regarding the rank of items at the</title>
          <p>Baseline
Dist
Bas
TF-IDF
Sem
Spec
Bas</p>
          <p>TF-IDF
Sem
moment they were returned, the TF-IDF and Semantic -Distance methods show a
significantly higher rank than the baseline, while for the Basic and the three Specificity  methods
no significant diference was found. Additionally, we studied the moment that the intended
items were covered by another item for the first time (i.e. the intended item is a child of an
item from the results). A complementary efect is found: where TSR improves, more keystrokes
are required before an item is covered. All our methods improved in this aspect compared to
the baseline, with Semantic requiring the least keystrokes. The mean number of suggestions
returned by each method are shown in the fourth column (Terms). All methods show a
significant decrease in this aspect compared to the baseline. We also found that although some of
the items were not found, they were covered by other items that were returned. The amounts
of non-returned but covered items (NRCI) are shown in the last column of Table 2, where the
last number indicates the amount of non-returned items.</p>
          <p>Experiment 3. In this experiment, participants were asked to indicate their preferences.
Each method received a score of either 1 or 0 per query, based on the majority of votes.
Comparing the baseline to our methods, for a majority of the queries the baseline was preferred
by users (ranging from 60% to 100% of the queries). In terms of relevance, user preferences
are as follows: TF-IDF is the most preferred, followed by Semantic , and finally Basic
(FigExperiment 2. While
Experiment 1 looked at if
and when items get
returned, Experiment 2
focuses on the rank at which
relevant items get returned
over various prefix lengths.</p>
          <p>Our methods improve on
the baseline in terms of
ranking, however, not on
recall (Figure 5).
Additionally, complementary
efects are shown on
ranking and recall as the
baseline initially ranks items
high in the list, while
starting with a relatively
low recall, but these both
reverse later on. Overall,
Distance shows to
consistently have the lowest
average rank from length
4 and higher, whereas
it seemingly counterpart
Specificity  shows to have
the highest recall.</p>
          <p>3.0
Experiment 3. User agreement is calculated using Fleiss  . Names are shortened for readability.</p>
          <p>D
R
D
R
R
D</p>
          <p>TF-IDF</p>
          <p>60%
TF-IDF
82%


Dist</p>
          <p>62%</p>
          <p>Bas
40%
Bas
18%
Spec
38%</p>
          <p>Sem
50%
Sem
67%
Dist
62%</p>
          <p>Dist (k=.57)</p>
          <p>Dist (k=.48)</p>
          <p>Dist (k=.44)
Spec (k=.35)</p>
          <p>Spec (k=.49)</p>
          <p>Spec (k=.38)
Bas
50%
Bas
33%
Spec
38%</p>
          <p>TF-IDF</p>
          <p>56%
TF-IDF
57%


Dist</p>
          <p>44%</p>
          <p>Sem
44%
Sem
43%
Spec</p>
          <p>56%
Bas (k=.21)</p>
          <p>TF-IDF (k=.19)</p>
          <p>Sem (k=.55)
over Specificity 
ure 3, top two tables). In terms of diversity settings, users seemed to mostly prefer Distance
, but we note here that the inter-user agreement is low for both conditions.
Specificity  is instead preferred when combined with Semantic (Table 3, bottom table).</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Lesson Learnt</title>
      <p>Baseline vs. Proposed Method.</p>
      <p>Our results show that our newly introduced methods are
precision-oriented, and, when returning the relevant item, they also rank correct results higher
than the baseline; while the baseline is recall-oriented, and tends to return longer result lists.
Our initial assumptions were that users want to be pointed to a particular item quickly, as
opposed to being shown a longer set of alternatives from which it may take more time to
choose the intended item. Contrary to this, our users found the baseline preferable to any of
our methods. This shows that QAC users in our use-case may be heavily recall and
rankingfocused, and they do not prefer short, focused lists of suggestions, as opposed to what findings
for web search indicate [6].</p>
      <p>There may be several reasons for users preferring the baseline. Primarily, most of our users
have already been exposed to the baseline auto-complete method, and they may have expressed
their preference towards something familiar to them. To confirm this, we are planning
experiments with users who see the search system, and any auto-complete method we want to
evaluate, for the first time. Secondly, user feedback is highly dependent on the user
experience design, and since our methods provide additional information to the user compared to
the “plain” baseline, this angle should also be considered when comparing various methods.</p>
      <p>Beyond this, using evaluation data from logs where the baseline method was in operation
could be a source of bias (e.g. for recall), and hence our findings are potentially more insightful
when comparing our methods to one another.</p>
      <p>Diversity</p>
      <p>Given a set of 20 most relevant items, Specificity 
than Distance . On one hand, those items tend to be closely located in the graph - thus, not
diverse according to Distance ; on the other hand, they show high variance in graph depth
- therefore, high diversity according to Distance . If a user profile requires (a) only a few
keystrokes before having the intended term suggested, and (b) there is a preference towards
more suggestions, then Specificity  would be the most adequate to use in the QAC system (see
will possibly select more items
stacle for them [16], and it could be important that slight spelling variations between diferent
concepts are brought to their awareness by showing more suggestions. However, since
Experiment 3 did not show convincing preferences for either Diversity
are needed to confirm this. Experiment 2 showed that
 method, further experiments
overall returns relatively
higher-ranked items in more concise lists. Therefore, user profiles that require fast typing and
Distance

quick investigating of suggestions would most likely gain more from using Distance .</p>
      <sec id="sec-5-1">
        <title>Experiment 3, we found that users preferred TF-IDF</title>
        <p>Relevancy</p>
        <p>Experiment 1 showed that relevance variations have the most impact on how
quickly an item is covered as users type. From Basic to TF-IDF and Semantic the ability to
cover items after a few keystrokes shows to increase. Given that TSR and the number of
returned suggestions both decrease between these settings, suggestions seem to be more abstract
for TF-IDF and, even more, for Semantic when compared to both Basic and the baseline. In
 over both Basic and Semantic , which
TF-IDF , with an observed user-preference for TF-IDF .
could indicate that they prefer the level of specificity of items selected by</p>
      </sec>
      <sec id="sec-5-2">
        <title>TF-IDF (i.e.more</title>
        <p>specific than</p>
        <p>Semantic , but more abstract than Basic ). As mentioned before, laypeople might
benefit from additional support in concept disambiguation. Therefore, given a user profile

where there is a need for awareness of diferences between subtypes to be inspected with care,
grouped items should be suggested first, while further refinements could be provided after
the query is submitted. This type of behaviour may be achieved through either Semantic or</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>Often, complex problems cannot have a ‘one size fits all’ solution. In the context of Query
Auto-Completion for medical search, we have found that no one solution fits all users’ needs.
However, we showed that recommendations could be learned for diferent user profiles, such
as, but not limited to, medical experts (i.e. healthcare providers, pharmaceutical
professionals) and laypeople (i.e. patients). We have proposed and investigated a graph-based method
that aimed to outperform a currently implemented QAC system. Our method has shown to
achieve this in terms of ranking and covering items. We experienced many benefits of using
a graph over a traditional database, such as, handling complex queries more time-eficiently
and easing the process of tracing descendants while calculating graph distance. Future work
will focus on improving recall as well. Furthermore, we will extend our user-based evaluation
to assess all returned items’ relevance. Ultimately, our work aspires to be used as a step
towards accommodating both laypeople and experts and improving the accessibility of health
information7.</p>
      <p>7Code and sample dataset are publicly available at https://research.mytomorrows.com/
[2] F. Su, M. Somaiya, S. Mishra, R. Mukherjee, EXOS: Expansion on session for enhancing
efectiveness of query auto-completion, in: 2015 IEEE International Conference on Big
Data (Big Data), 2015, pp. 1154–1163. doi:10.1109/BigData.2015.7363869.
[3] S. Whiting, J. M. Jose, Recent and robust query auto-completion, in: Proceedings of the
23rd international conference on World wide web - WWW ’14, ACM Press, Seoul, Korea,
2014, pp. 971–982. doi:10.1145/2566486.2568009.
[4] B. Rieder, G. Sire, Conflicts of interest and incentives to bias: A microeconomic critique of
Google’s tangled position on the Web, New Media &amp; Society 16 (2014) 195–211. doi:10.
1177/1461444813481195.
[5] Z. Szlávik, W. Kowalczyk, M. Schut, Diversity measurement of recommender systems
under diferent user choice models, in: Fifth International AAAI Conference on Weblogs
and Social Media, 2011.
[6] E. Cutrell, Z. Guan, What are you looking for? an eye-tracking study of information usage
in web search, in: Proceedings of the SIGCHI conference on Human factors in computing
systems, 2007, pp. 407–416. doi:10.1145/1240624.1240690.
[7] A. Rotegård, L. Slaughter, C. Ruland, Mapping nurses’ natural language to oncology
patients’ symptom expressions, Studies in health technology and informatics 122 (2006)
987–8.
[8] M. Dahm, Coming to terms with medical terms – exploring insights from native and
nonnative english speakers in patient-physician communication, HERMES - Journal of
Language and Communication in Business 25 (2017) 79–98. doi:10.7146/hjlcb.v25i49.
97739.
[9] J. Carbonell, J. Goldstein, The use of MMR, diversity-based reranking for reordering
documents and producing summaries, in: Proceedings of the 21st annual international
ACM SIGIR conference on Research and development in information retrieval - SIGIR ’98,
ACM Press, Melbourne, Australia, 1998, pp. 335–336. doi:10.1145/290941.291025.
[10] A. Alhindi, U. Kruschwitz, C. Fox, M.-D. Albakour, Profile-based summarisation for web
site navigation, ACM Transactions on Information Systems (TOIS) 33 (2015) 1–39. doi:10.
1145/2699661.
[11] J. Yan, W. Chu, R. W. White, Cohort modeling for enhanced personalized search, in:
Proceedings of the 37th international ACM SIGIR conference on Research &amp; development
in information retrieval, 2014, pp. 505–514. doi:10.1145/2600428.2609617.
[12] J. List, The name of the game: Information seeking in a professional context, in:
Proceedings of the Integrating IR Technologies for Professional Search Workshop, Moscow,
Russia (March 24, 2013), 2013.
[13] S. Verberne, J. He, G. Wiggers, T. Russell-Rose, U. Kruschwitz, A. P. de Vries, Information
search in a professional context-exploring a collection of professional search tasks, arXiv
preprint arXiv:1905.04577 abs/1905.04577 (2019). arXiv:1905.04577.
[14] R. W. White, S. T. Dumais, J. Teevan, Characterizing the influence of domain expertise
on web search behavior, in: Proceedings of the Second ACM International Conference
on Web Search and Data Mining - WSDM ’09, ACM Press, Barcelona, Spain, 2009, p. 132.
doi:10.1145/1498759.1498819.
[15] J. Palotti, A. Hanbury, H. Müller, C. E. Kahn, How users search and what they
search for in the medical domain, Inf Retrieval J 19 (2016) 189–224. doi:10.1007/
s10791-015-9269-8.
[16] P. C.-I. Pang, K. Verspoor, S. Chang, J. Pearce, Conceptualising health information seeking
behaviours and exploratory search: result of a qualitative study, Health Technol. 5 (2015)
45–55. doi:10.1007/s12553-015-0096-0.
[17] P. Castells, N. J. Hurley, S. Vargas, Novelty and diversity in recommender
systems, in: Recommender systems handbook, Springer, 2015, pp. 881–918. doi:10.1007/
978-1-4899-7637-6_26.
[18] Z. Huang, B. Cautis, R. Cheng, Y. Zheng, N. Mamoulis, J. Yan, Entity-Based Query
Recommendation for Long-Tail Queries, ACM Transactions on Knowledge Discovery from
Data 12 (2018) 1–24. doi:10.1145/3233186.
[19] M. Zhong, H. Cheng, Y. Wang, Y. Zhu, T. Qian, J. Li, Towards both Local and Global
Query Result Diversification, in: G. Li, J. Yang, J. Gama, J. Natwichai, Y. Tong (Eds.),
Database Systems for Advanced Applications, volume 11447, Springer International
Publishing, Cham, 2019, pp. 464–481. doi:10.1007/978-3-030-18579-4_28.
[20] T. Groza, K. Verspoor, Assessing the Impact of Case Sensitivity and Term Information
Gain on Biomedical Concept Recognition, PLOS ONE 10 (2015) e0119091. doi:10.1371/
journal.pone.0119091.
[21] B. Sathiya, T. V. Geetha, A review on semantic similarity measures for ontology, Journal
of Intelligent &amp; Fuzzy Systems 36 (2019) 3045–3059. doi:10.3233/JIFS-18120.
[22] A. Jaech, M. Ostendorf, Personalized language model for query auto-completion, in:
Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics
(Volume 2: Short Papers), Association for Computational Linguistics, Melbourne,
Australia, 2018, pp. 700–705. doi:10.18653/v1/P18-2111.
[23] R. L. Cecchini, C. M. Lorenzetti, A. G. Maguitman, I. Ponzoni, Topic relevance and
diversity in information retrieval from large datasets: A multi-objective evolutionary
algorithm approach, Applied Soft Computing 69 (2018) 749 – 770. doi:10.1016/j.asoc.
2017.11.016.
[24] S. Kalloori, F. Ricci, R. Gennari, Eliciting pairwise preferences in recommender
systems, in: Proceedings of the 12th ACM Conference on Recommender Systems,
RecSys ’18, Association for Computing Machinery, New York, NY, USA, 2018, p. 329–337.
doi:10.1145/3240323.3240364.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bar-Yossef</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kraus</surname>
          </string-name>
          ,
          <article-title>Context-sensitive query auto-completion</article-title>
          ,
          <source>in: Proceedings of the 20th international conference on World wide web - WWW '11</source>
          , ACM Press, Hyderabad,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>