<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Survey Data System: a Framework to Enhance Response Rate in Clinical Studies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giulio De Sabbata</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>PhD Sudent, ICT Doctorate at DBGroup, University of Modena and Reggio Emilia</institution>
          ,
          <addr-line>Modena</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>Clinical studies involve typically surveys of patients to gather data. Coaching aims at identifying the least engaged patients to enhance the response rate. To support the coaches we have developed a platform that enables ongoing data pipelines to elicit insights during the study from these patient-generated data. The most challenging aspect derives from the attempt of including the content of the individual responses. Exploiting this information occurs to the detriment of the possibility to generalize the results, preventing any kind of inference across diferent studies. The framework designs two AI solutions for the analysis of survey data. The first method is a forecast of the response rate. The second consists in an implementation of a descriptive module that generates an user-friendly interface displaying patterns in the response rate behaviour. The coach leverages this visualization tool especially if the predictive module falls short, taking advantage of a tidy view for optimizing the patient engagement strategies during the study. A first approach to the stated problem and an innovative employment of the tree technique are the main contributions of this paper.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data Mining</kwd>
        <kwd>Questionnaire Processing</kwd>
        <kwd>Tree Technique</kwd>
        <kwd>Data Visualization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        For medical research, surveys are the standard to collect data from patients participating
to clinical studies such as trials or observational studies. A survey research is essentially a
collection of information from a sample of individuals through their responses to questions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
and we refer, from now on, to them as surveys or studies. One fundamental requirement for
the study success is the patient engagement [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which is typically measured by the survey
response rate. To increase the response rate our approach is to exploit questionnaire compiling
information (i.e., log information and actual responses to the questions) during the study for
defining questionnaire administration (i.e., policy for reminders or scheduling time) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This
approach entails processing datasets that pose significant challenges for a predictive task, such
as limited observations number and a high number of explanatory variables.
      </p>
      <sec id="sec-1-1">
        <title>Survey Data Platform for Clinical studies</title>
        <p>In this paper we design the implementation of a tool in a digital platform for gathering and
processing survey data. This tool is developed to meet the specific needs of the coach, the
enduser of the platform that is an health worker committed to patients monitoring. The coaching
activity aims at increasing response rate by contacting the patients to maximize the response
rate. The objective of the work is to elicit from survey data relevant insights for the coaching.</p>
        <p>To achieve that, we propose a data-driven approach with the implementation of two
independent AI solutions. The first functionality enables a forecast of the response rate levels allowing
to straightforwardly identify the least engaged patients. Alternatively, the coach could resort to
a second module that ofers a tidy view of the relevant data afecting mostly the response rate
behaviour. The coach, a domain expert with no knowledge about data, identifies supported by
an highly interpretable interface which patients to contact.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Contributions</title>
        <p>The novelty of the work consists in a new approach to increase the response rate by eliciting
insights from survey data while the study is still running and by including the responses.
Furthermore, to our knowledge there are no works which exploit the tree technique in such a
manner.</p>
        <p>The remainder of the paper is organized as follows: Section 2 describes the studies features
and the relative implications for the analysis; Section 3 deals with the system managing the data
survey platform; Sections 4 and 5 introduce the methods designed to solve the problem resorting
to the gradient boosting algorithm for the predictive task and to an innovative implementation of
the tree technique for the descriptive interface; Section 6 describes the results of the experiments
and the challenges involved in the field-test of the method; in Section 7 are argued the main
contributions; lastly, Section 8 draws the conclusions and links them to future improvements.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Preliminaries</title>
      <p>
        In the introduction are mentioned two features colliding with predictive task: dataset with a
limited sample size and a high number of explanatory variables. The large number of explanatory
variables derives from considering each question as an input feature in the framework and,
hence, the forecast could sufer from the curse of dimensionality. Whereas, we can identify at
least two main reasons that pose the framework in a scenario of narrow data: small surveys are
frequent and their processing occurs in the ongoing period. Typically, survey management is
expensive and it is common for small and medium research group to aim at keeping their size
as small as possible due to limited resources [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. For the same reason the questionnaires are
extremely heterogeneous and characterized by traits that limit analysis possibilities [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In the analyzed studies those aspects are actually present (i) restricted amount of patients
(ranges from 30 to 100 patients) (ii) large amount of questions for each study (the count can
reach 100 questions) (iii) time span of studies difers (lengths of the studies range from one
month to one year).</p>
      <p>Some may argue that combining studies could overcome the limited size of the dataset.
However, the description of the studies suggests that they difer significantly preventing us to
join them. Additionally, the lack of consistency in the questions across the studies precludes us
from combining the studies and meanwhile exploiting the responses.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Survey Data Management System</title>
      <p>Patients
PostgreSQL
Gathering
Extraction
1. Pre-processing
2. Processing
End-users
Visualization
Prediction</p>
      <sec id="sec-3-1">
        <title>3.1. Architecture Overview</title>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Survey Data Pre-Processing Module</title>
        <p>
          This module organizes data so that the processing phase runs optimally; the actions performed
in the module pertain mostly to the structure or the organization of the dataset. Granularity
refers to the number of entities represented by a record in a dataset [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The granularity of
the data in the raw stage is too fine with each record representing a scheduled question to a
patient in a certain date. To identify meaningful patterns, user records are aggregated over a
time period, called cycle, such that each row represents the patient behaviour in that cycle. To
minimize the loss of information caused by the aggregation on each cycle, it is important to
balance the need for shorter cycles and the identification of underlying patterns. The optimal
length of the cycle, and the number of cycles, may depend on the length of the study and on the
frequency of the questionnaires. Once the time frame is defined, switching from the original
indexing to the new one requires to pivot the diferent questions on the columns, setting them
as input features. In each cycle, the value for each question is computed as the average of the
corresponding non-missing responses in that specific cycle. In conclusion, it is measured the
response rate as the proportion of compiled questionnaires on the total. To get flexible insights
we derive a new key indicator, the spread of the response rate, which measures the diference
in response rate over consecutive cycles. The coach interactively sets the response rate or the
spread as the target variable.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Response Rate Forecast</title>
      <p>
        The predictive module gives in output the response rate values of the future cycle, fitted
with a Gradient Boosting algorithm. This particular model belongs to a highly efective and
extensively used family of machine-learning techniques [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Boosting algorithm follows an
ensemble learning approach where one of the most performing is the stump-based. The boosting
approach generates an additive model that re-weights the data at each iteration to focus on
the most dificult examples to predict and, thus, to reduce the risk of overfitting. To optimize
the model performance these are the main parameters to tune (i) number of iterations
controls the number of weak models; (ii) maximum tree depth controls the size of each
estimator, if passed one each weak model is a tree stump; (iii) learning rate controls the
amount by which the contribution of each tree is reduced.
      </p>
      <p>The module contains an evaluation modality as well as the predictive one. In the field
of predictive modeling, it is crucial to evaluate the dependability of the fitted values. The
assessment of the model’s performance on the current data provides an indication of its ability
to perform well in future predictions. The evaluation modality generates two distinct metrics
the score and the progressive recall. The score evaluates the goodness of fit using a
re-examined coeficient of determination, which is measured on a scale from -1 to 1. When
the score is 0, the model performance is equivalent to the baseline model, and, hence, if the
score is negative, the model performance falls shorter than fitting all values with the mean.
The progressive recall measures the percentage of patients correctly identified by the
predictive model as part of the least responding patients.
5. Descriptive Trees: a Method to Discriminate Patient</p>
      <p>
        Behaviour
Alternatively to the predictive module, the coach could resort to a tree interface displaying
descriptive rules for discriminating patient behaviour based on their response rate. The task
consists in clustering patients not only by response rate, but also by the factors that have the
greatest impact on it. This visualization may suggest the coach which are the most efective
actions for increasing patient engagement. To perform this task a decision tree method is
adopted for two main reasons. Firstly, it can handle numerous explanatory variables and
perform the task of variable selection [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], overcoming the curse of dimensionality generated
from the large number of pivoted questions. Secondly, it is easily displayed and interpretable [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
generating highly intuitive insights. Overall, this technique allows to select relevant variables
and displays higly intuitive results.
      </p>
      <p>The work exploits the tree technique in an innovative way since we are not performing
any predictive task. In particular, we do not split data in training and test set but we just run
the model on the whole dataset. This choice follows the intuition that the method needs just
describing/summarising data.</p>
      <p>It was possible to process data indiferently either with a regression tree or with a decision
tree. A decision tree is adopted since it requires to discretize the dependent variable into a
qualitative one, making it more intuitive with the definition of two categories of patients. The
mean of the response rate is the default choice to set the threshold to categorize the response
rate. However the coach may customize the tree by setting a diferent threshold according to
her budget. For what concerns (Figure 2) (i) the two patient categories are distinguished by
diferent colours (i.e., the most engaged patients are the green) (ii) it is reported the text of the
input feature that generates the split.
6. Experiments and Open Challenges of Descriptive Trees</p>
      <p>Field-Test</p>
      <sec id="sec-4-1">
        <title>6.1. Experiments</title>
        <p>The framework has been tested on 5 diferent studies. Initially, diferent set of explanatory
variables are tested to evaluate how the predictor behaves. The results indicate that all available
explanatory variables should be included, suggesting that the algorithm can efectively handle
the curse of dimensionality. The metrics of each study are generated at an intermediate and a
conclusive time point and collected in Table 1.</p>
        <p>In general, the performance improves according to the length of the study and to its progress.
This result suggests that the size of the study is a critical point and lack of data represents an
issue.</p>
      </sec>
      <sec id="sec-4-2">
        <title>6.2. Open Challenges of Descriptive Trees Field-Test</title>
        <p>The evaluation of the efectiveness of the descriptive tree method is yet to be tested through a
ifeld-test in an ongoing study, and thus, there are no actual validations available. Due to the
experimental nature of the method, a well-structured plan is required for the implementation
of its field-test. The first challenge is the identification of appropriate metrics to measure the
efectiveness of the method. One possible approach could involve soliciting feedback from
Study
Study1
Study1
Study2
Study2
Study3
Study3
Study4
Study4
Study5
Study5</p>
        <p>Time
Intermediate
End
Intermediate
End
Intermediate
End
Intermediate
End
Intermediate
End
coaches who have utilized this method, such as through a survey. Secondly, further
improvements are necessary to ensure that the resulting clusters are meaningful. The meaningfulness of
a cluster depends on the relevance of the insights it provides, where the explanatory variables
involved in the splits can suggest the reasons for the cluster behavior and guide the appropriate
actions. Thus, the inclusion of domain-specific knowledge or expert input could address the
interpretability of the visualization. For instance, the coach could undertake a feature selection
task by retaining the most interesting questions related to the treatment. Alternatively, the
coach could re-weight the features to give more importance to the most promising ones. The
third point, which is closely related to the second, is the balance between a user-friendly tool
and its efectiveness. Simplifying the tool usability occurs to the detriment of its capability
to provide relevant summaries. For example, achieving the desired level of usability requires
careful consideration of the design parameters with particular regard to their number and
complexity of the settings. The issue of balancing the size and the accuracy of the displayed
tree is an additional example of this challenge. In general, a cost-complexity parameter controls
this trade-of, which, in a predictive scenario, trees are pruned according to preventing the tree
from overfitting the data. In this framework the parameter regulation depends on other criteria
such as the screen size of the coach’s smart device or the lack of enough explanatory power of
the data itself.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7. Related Works</title>
      <p>
        In the literature there is a plethora of research focusing on the processing phase when the studies
have been concluded yet [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The novelty of this work lies in the idea of enhancing the response
rate directly in the administration phase by processing the questionnaires while the study is still
running. Moreover, to our knowledge in the literature they just exploit contextual information
and ignore the responses of the questionnaires. To provide for the potential criticisms in the
predictive task, a descriptive module supports the end-user decision-making. Despite this
module performs a task similar to clustering, the labels, generated by the levels of the response
rate, of the data are known and clusters are generated with respect to these labels. An additional
feature distinguishing the task from clustering pertains the grouping operation that selects the
factors afecting most the response rate behaviour.
      </p>
    </sec>
    <sec id="sec-6">
      <title>8. Conclusion and Future work</title>
      <p>This framework aims at eliciting insights from survey data to increase the response rate. In
addition to its predictive functionality, the framework includes a descriptive module which
ofers an alternative solution when the predictive task falls short. The descriptive module
employs a highly interpretable tree visualization that groups patient based on their response
rate behaviour. However, further improvements and a field-test are necessary to validate the
efectiveness of the descriptive trees, given their experimental nature.</p>
      <p>In the future, mapping at least the more general questions across diferent studies and
standardizing the size and length of each study could enable the aggregation of survey data
from multiple sources. This approach may facilitate the identification of patterns to infer across
diferent studies.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Check</surname>
          </string-name>
          , R. K. Schutt, Research methods in education,
          <source>Sage Publications</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Dufett</surname>
          </string-name>
          ,
          <article-title>Patient engagement: What partnering with patient in research is all about</article-title>
          ,
          <source>Thrombosis Research</source>
          <volume>150</volume>
          (
          <year>2017</year>
          )
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          . URL: https://www.sciencedirect.com/science/ article/pii/S0049384816306089. doi:https://doi.org/10.1016/j.thromres.
          <year>2016</year>
          .
          <volume>10</volume>
          .029.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dörnyei</surname>
          </string-name>
          , T. Taguchi, Questionnaires in second language research: Construction, administration, and processing, Routledge,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <article-title>Sample size for a clinical trial: why do some trials need only 100 patients and others 1000 patients or more?</article-title>
          ,
          <source>Ophthalmology</source>
          <volume>113</volume>
          (
          <year>2006</year>
          )
          <fpage>721</fpage>
          -
          <lpage>722</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>I. S.</given-names>
            <surname>Sjetne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. A.</given-names>
            <surname>Bjertnaes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Olsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. H.</given-names>
            <surname>Iversen</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Bukholm,</surname>
          </string-name>
          <article-title>The generic short patient experiences questionnaire (gs-peq): identification of core items from a survey in norway</article-title>
          ,
          <source>BMC Health Services Research</source>
          <volume>11</volume>
          (
          <year>2011</year>
          )
          <article-title>88</article-title>
          . URL: https://doi.org/10.1186/
          <fpage>1472</fpage>
          -6963-11-
          <lpage>88</lpage>
          . doi:
          <volume>10</volume>
          .1186/
          <fpage>1472</fpage>
          -6963-11-88.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Rattenbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Hellerstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kandel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Carreras</surname>
          </string-name>
          ,
          <article-title>Principles of data wrangling: Practical techniques for data preparation, "</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <source>Inc."</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Natekin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Knoll</surname>
          </string-name>
          ,
          <article-title>Gradient boosting machines, a tutorial</article-title>
          ,
          <source>Front Neurorobot</source>
          <volume>7</volume>
          (
          <year>2013</year>
          )
          <fpage>21</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W.</given-names>
            <surname>Loh</surname>
          </string-name>
          ,
          <article-title>Variable Selection for Classification and Regression in Large p, Small n Problems, Probability Approximations and Beyond (</article-title>
          <year>2012</year>
          )
          <fpage>135</fpage>
          -
          <lpage>159</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Elmachtoub</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C. N.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>McNellis</surname>
          </string-name>
          ,
          <article-title>Decision trees for decision-making under the predict-then-optimize framework</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2858</fpage>
          -
          <lpage>2867</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Saleh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bista</surname>
          </string-name>
          ,
          <article-title>Examining factors impacting online survey response rates in educational research: Perceptions of graduate students</article-title>
          .,
          <source>Online Submission</source>
          <volume>13</volume>
          (
          <year>2017</year>
          )
          <fpage>63</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>