<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Decision Support Architecture for Primary Studies Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vilmar Nepomuceno</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Informatics Center (CIn) Federal University of Pernambuco Recife - PE</institution>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <fpage>10</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>Background. A systematic literature review is a process in which all relevant available research about a research question is identified, evaluated, and interpreted through individual studies. The workload required for this process may bias the evaluation of the studies, affecting the result. Aim. Creating a decision support architecture to assist participants of a systematic review in the selection process of the individual studies and quality assessment of these studies, possibly improving the execution time and reducing the evaluation bias. Method. Improving the primary studies selection and quality assessment processes by using text mining techniques and ontologies to construct a decision support architecture. We will also conduct experiments to evaluate the proposed architecture. Contribution. Improve the primary studies selection and quality assessment processes, reducing its workload, and lowering the evaluation bias in systematic literature reviews.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Systematic Literature Review</kwd>
        <kwd>Text Mining</kwd>
        <kwd>Ontology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The process of primary studies selection is a free interpretation of
the criteria of whom is leading the SLR, hence, the large number
of papers retrieved during the search process and the poor quality
of their abstracts [2] makes the completion of the selection
process a hard task and sometimes inaccurate. After the selection
process is carried out to evaluate the quality of selected papers to
increase the reliability and importance of SLR results, and to
perform this task, there are several guidelines, which are usually
not properly followed and the use, in general, is not justified by
the authors [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ]. These two processes of primary studies
evaluation, study selection and quality assessment are hard tasks
and time consuming [
        <xref ref-type="bibr" rid="ref3">4</xref>
        ], much of this time due to the primary
study reading procedure.
      </p>
      <p>
        There are studies that talk about automatic selection of primary
studies using text mining techniques, in which the primary studies
are classified by the similarity of texts [
        <xref ref-type="bibr" rid="ref4">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ]. However, in these
studies, the conductor does not assist the process, which can
generate a set of papers that is significantly different than would
be generated by the manual process. Besides, according to
Kitchenham et al. [
        <xref ref-type="bibr" rid="ref6">7</xref>
        ], the process of quality assessment for
primary studies is essential for SLRs. Our research seeks (i) to
identify a way to semi-automatically support the primary studies
quality assessment, with the use of text mining techniques and
ontologies to describe prior knowledge about the SLR, as well as
(ii) identifies, semi-automatically, the criteria for inclusion/
exclusion inside the text of the studies returned by the search, also
through text mining techniques, supporting the selection process
of the primary studies.
      </p>
      <p>
        The semi-automatic evaluation of the quality and the
inclusion/exclusion criteria semi-automatic search should support
the evaluation of primary studies in SLR process. With this, we
aim to reduce the execution time of the SLR, due to the primary
study reading procedure, decreasing the primary studies
evaluation bias that may occur due to the evaluation process
subjectivity, increasing the assurance that the outcome of the SLR
is not being compromised [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ]. In addition, you can increase the
studies search space, which today is limited because of the effort
spent during the selection process and subsequent quality
evaluation of these studies.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. RELATED WORKS</title>
    </sec>
    <sec id="sec-3">
      <title>2.1 Quality Assessment</title>
      <p>
        There are several guidelines available for quality assessment of
primary studies, such as [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ] which put forward eleven evaluation
criteria based on CASP [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ]. In Kitchenham et al. [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ] was used a
checklist for quality assessment, in order to specify an appropriate
process for evaluating quality. The study concluded that at least
two evaluators are required to improve the reliability and the
quality assessment should be represented by the sum of the
criteria used. However, Dieste et al. [
        <xref ref-type="bibr" rid="ref11">12</xref>
        ] identified trends that
should not exist a plethora of items into instruments of quality
control on SLRs, and that this assessment should be careful about
the limits of this process with respect to aspects of internal
validity [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ].
      </p>
      <p>There is still not a standard process for primary studies quality
assessment in software engineering, several authors have
suggested several ways to estimate the validity of the studies.
Probably due to this lack of standardization, in our research were
not found tools that automate the process, or part of this process.</p>
    </sec>
    <sec id="sec-4">
      <title>2.2 Primary Studies Selection</title>
      <p>
        The selection of primary studies is the most time-consuming task
for an SLR, which can be affected by the titles and abstracts that
Copyright © 2015 for this paper by its authors. Copying permitted for private and academic purposes
do not reflect well the content of the work [
        <xref ref-type="bibr" rid="ref3">4</xref>
        ]. Additionally, time
constraints may lead the research conductors to reduce the search
space. Hence, automate the process, or part of it can help
overcome these barriers.
      </p>
      <p>
        Automatic classification of primary studies, indicating its
inclusion/exclusion can be found in some works [
        <xref ref-type="bibr" rid="ref4">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ], in other
words, in these works, the study selection process is done in an
automated way without the intervention of the researcher. With
this approach can dramatically reduce the effort spent on that task,
however, methods that do this kind of selection may include
studies that did not help in the research, because of their low
quality, or the algorithm low accuracy, as well as, they may
exclude studies with low textual similarity, but with good quality,
which could be used by the researcher in some way.
      </p>
      <p>Therefore, this research proposes a semi-automated approach,
leaving to the researcher the final decision of the inclusion or not
of these studies. It is likely that when comparing our methodology
and the automatic classifiers, there will be a loss of time taken to
perform the task by using the technique proposed in this work.</p>
      <p>However, it is possible that the final result of the selection
process, using our proposal, is more satisfactory and significantly
faster than the manual approach.</p>
    </sec>
    <sec id="sec-5">
      <title>3. PROPOSAL</title>
      <p>According to the aim of this study, our guiding research question
(RQ) is “How can we improve the process of evaluating primary
studies by automating parts of the process?”. This question can
be decomposed into:
 RQ1. How can we improve the selection process by
automating parts of the process?
 RQ2. How can we improve the quality assessment
process by automating parts of the process?
An initial version of the proposed Decision Support Architecture
(DS Architecture) is presented in Figure 1.</p>
      <p>
        The knowledge about the SLR will be represented by an ontology
(1) based on the protocol designed for the execution of the SLR
and possible research conductor's refinements. This approach is
proposed by Biolchini et al [
        <xref ref-type="bibr" rid="ref12">13</xref>
        ], that state it is possible to
improve the results obtained during the SLR through
standardization of the terminology for the concepts involved using
an ontology. It is still undefined in what form the creation of the
ontology will be held, but there are free available tools such as
protégé [
        <xref ref-type="bibr" rid="ref13">14</xref>
        ] and Jena [
        <xref ref-type="bibr" rid="ref14">15</xref>
        ] and the GATE (General Architecture
for Text Engineering) framework [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ], that was already used in
[
        <xref ref-type="bibr" rid="ref16">17</xref>
        ] for this purpose.
      </p>
      <p>
        These criteria should be used as input to a text mining algorithm,
which should be available as a tool component (DSTool) and will
be able to identify whether the selection criteria have been met
and which quality criteria can be identified inside the primary
study text. The interest points of the text, where the tool will be
based to respond, will be shown to the research conductor, which
will decide whether the criteria were actually achieved. The
inclusion/exclusion criteria will be provided by the conductor in
the protocol creation, and the quality criteria will come from one
of the guidelines available in the literature, after to perform a
systematic literature review that should indicate the best guideline
to be implemented by the text mining algorithm. The main idea to
provide a previous guideline is to improve the robustness and
accuracy of the algorithm, which does not prevent that other
criteria, which are not present in the selected guideline, may be
used. Another input to the text mining algorithm is the ontology
built upon the protocol (3), many text mining techniques use
ontologies as a knowledge base, one of which is the
ontologybased question answering system [
        <xref ref-type="bibr" rid="ref17">18</xref>
        ], which It is the starting
point in the tool development. Still will be evaluated, ways to
present the results from the tool (4).
      </p>
    </sec>
    <sec id="sec-6">
      <title>3.1 Decision Support Algorithm</title>
      <p>
        Based on the system architecture proposed by Bo and Yunqing
[
        <xref ref-type="bibr" rid="ref17">18</xref>
        ], a question answer architecture is being proposed to the
decision support algorithm (Figure 2).
      </p>
      <p>
        The questions that will serve as input to the algorithm will be
provided by the protocol, at this time the criteria for
inclusion/exclusion and the quality criteria will be provided one
by one to the segmentation process (1). At this point, the question
will be broken in terms, and the keywords will be extracted (2).
As, for example, the following quality criteria: “Is there a clear
statement (definition) of the aims (goals, purposes, problems,
motivations, objectives, questions) of the research?” [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ], where
we can draw the following set of terms T = {statement, definition,
goals, purposes, problems, motivations, objectives, questions,
research}. The keywords found in the criteria will be expanded
with the help of the ontology (3). Thus, we can create a larger set
of terms that should improve the algorithm accuracy. These three
points can be called as the search query creation process.
After creating the search query, we look for the answers in the
primary studies (4), at this point we use text mining algorithms to
find possible answers to the criteria and return the points in the
text where the answers can be found, as well as the text of the
response (5).
      </p>
    </sec>
    <sec id="sec-7">
      <title>4. EVALUATION PROPOSAL</title>
      <p>
        To evaluate the effectiveness of the Decision Support Architecture
(DSArch) a controlled experiment [
        <xref ref-type="bibr" rid="ref18">19</xref>
        ] will be performed. The
study aims to answer the following research questions:



      </p>
      <p>RQ1. Does DSArch decrease the selection time of
primary studies?
RQ2. Does DSArch increase consensus between
individuals of the same pair including/excluding
studies?
RQ3. Does DSArch decrease the quality assessment
time of the selected primary studies?</p>
    </sec>
    <sec id="sec-8">
      <title>4.1 Hypotheses, Variables, and Parameters</title>
      <p>The null hypotheses are presented below:
 H0,RQ1. There is no difference in the execution time in
the selection process with or without the use of DSArch.
 H0,RQ2. There is no difference in the consensus among
peers with or without the use of DSArch.
 H0,RQ3. There is no difference in the quality assessment
time of the selected primary studies with or without the
use of DSArch.</p>
      <p>To examine the hypotheses the following dependent variables will
be used:
</p>
      <p>Selection Time. The time to complete the process for
selecting primary studies.
 Consensus. Measured by inter-rater agreement
coefficient. Cohen´s Kappa will be used, due to the
decision be taken by pairs.
 Quality Time. The quality assessment time of the
selected primary studies.</p>
      <p>The experiment factor is the process for evaluating the primary
studies, by selecting these studies and quality assessment of the
selected primary studies. The factor alternatives (treatments) are
executing the inclusion/exclusion process and quality assessment
using DSArch and the other is using only the protocol criteria
manually, without tool support for the alternatives.</p>
    </sec>
    <sec id="sec-9">
      <title>4.2 Material, Tool, and Training</title>
      <p>To perform the experiment a subset of primary studies drawn from
an SLR will be provided to participants. The protocol created for
the SLR will be provided and both the search process and the
visualization of the primary studies will be conducted by
REviewER1. It belongs to our research group and it is a tool that
gives support to search process in some databases (ACM library,
Engineering Village, IEEExplorer, Science Direct, Scopus, and
Springer Link) in a automated way and gives support to the
primary studies selection process. A version of REviewER will be
developed, containing the DS Architecture, to be used in the
evaluation. However, only the DSArchictecture should be
assessed.
1 http://sites.google.com/site/eseportal/tools/reviewer</p>
      <p>A training will be conducted with the participants to present the
reviewer tool and how one should analyze the DSArch. For this
training some primary studies, chosen from a subset of the
primary studies set obtained from the SLR search process, will be
selected and participants should evaluate them using the factor
alternatives.</p>
    </sec>
    <sec id="sec-10">
      <title>4.3 Task and Data</title>
      <p>To measure the dependent variables two tasks will be realized,
which requires no prior knowledge on SLR from the subjects,
hence, we aim to facilitate the process of selection of these
participants. The tasks to be performed in the experiment are the
selection of primary studies from a subset of the primary studies
set obtained from the SLR search process and the quality
assessment of these selected studies. No other SLR procedures
will be performed.</p>
      <p>At the end of the execution, the participants must provide a list of
accepted primary studies, the time taken for completion of the
selection process and the time taken for completion of the quality
assessment process. A questionnaire will be performed after the
experiment to evaluate the experiment itself and what the
participants thought about the proposed architecture.</p>
    </sec>
    <sec id="sec-11">
      <title>4.4 Execution</title>
      <p>
        The experiment will take place in a lab with the presence of all
participants at the same time. Participants will be divided into
pairs following the proposed by Kitchenham et al [
        <xref ref-type="bibr" rid="ref6">7</xref>
        ], such choice
will be at random.
      </p>
    </sec>
    <sec id="sec-12">
      <title>4.5 Data Analysis</title>
      <p>Latin squares have an analysis procedure very similar to the
factorial experiments (multiple factors). In the case of factorial
experiments, we can consider the lock variable as a factor. For the
Latin square, we have two lock variables and a factor of interest.
Also, there is another complicating factor: the experiment design
has multiple replications. This leads to a replicated Latin square
design with equal columns (processes) and different lines
(participants).</p>
      <p>
        A possible statistical test for the analysis is the ANOVA, as
proposed in [
        <xref ref-type="bibr" rid="ref18">19</xref>
        ], however, for being a parametric test, some
preconditions should be evaluated, and if one of them is violated a
equivalent non-parametric test can be used.
      </p>
    </sec>
    <sec id="sec-13">
      <title>4.6 Threats to Validity</title>
      <p>Some possible threats to validity of the experiment are already
being assessed.</p>
      <sec id="sec-13-1">
        <title>4.6.1 Internal Validity</title>
        <p>The completion of the training can generate an apprenticeship in
relation to DSArch, which may influence the evaluation process
of the primary studies when participants are not using the
DSArch. However, if there is influence, it will be in favor of the
null hypothesis.</p>
      </sec>
      <sec id="sec-13-2">
        <title>4.6.2 External Validity</title>
        <p>The results may not be generalizable to all the researches that
perform SLRs, because we do not sample from the population of
SLRs researchers, but we intend to make a satisfactory outcome
that will bring evidence of the effectiveness of DSArch.</p>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>5. CONTRIBUTIONS, FUTURE WORK</title>
    </sec>
    <sec id="sec-15">
      <title>AND ADVICES</title>
      <p>The use of the DSArch might reduce the effort to complete SLR,
which is one of the major problems encountered in conducting
this type of research. Another problem to be addressed by using
DSArch is the primary studies subjective selection, which is
biased, by automating part of the process. Another gain by
automating part of the process is decreasing the number of
conflicts generated during the evaluation of studies within the
pairs. An expectation for the proposed algorithm is that it can be
used to assist in the SLR data analysis process, since, the question
answer principle can also be used at this stage, however, this
analysis is out of the context of this thesis.</p>
      <p>So far, it was conducted an ad-hoc literature review on the
proposed theme, as well a first design of the controlled
experiment that will be conducted to evaluate the DSArch. The
planned next steps are:
 Execute a systematic literature review of techniques for
assessing the quality of primary studies, aiming at
choosing a technique to be semi-automated;
 Select and implement an ontology to represent the prior
knowledge on the SLR;
 Development of the DSArch;
 Refine the plan and execute the controlled experiment
to assess the DSArch;
 Analyze the obtained results;
 Write the Thesis.</p>
      <p>The main points where advices are needed:


</p>
      <p>The proposed architecture is consistent with the
problem found?
The experimental design can evaluate the proposed
architecture?
The statistical test, ANOVA can evaluate the data
generated by the experiment?</p>
    </sec>
    <sec id="sec-16">
      <title>6. REFERENCES</title>
      <p>[1] Kitchenham B. A. and Charters S.. Guidelines for performing
systematic literature reviews in software engineering,
Technical Report: 2007.
[2] Brereton P., Kitchenham B. A., Budgen D., Turner M. and
Khalil M.. Lessons from applying the systematic literature
review process within the software engineering domain.
at</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>Journal of Systems and Software</source>
          , vol.
          <volume>80</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>571</fpage>
          -
          <lpage>583</lpage>
          ,
          <year>2007</year>
          . DOI:
          <volume>10</volume>
          .1016/j.jss.
          <year>2006</year>
          .
          <volume>07</volume>
          .009.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Zhou</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Babar</surname>
            <given-names>M. A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tang</surname>
            <given-names>H..</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Quality assessment of systematic reviews in software engineering: a tertiary study</article-title>
          .
          <source>In Proceedings of the 19th International Conference on Evaluation and Assessment in Software Engineering (EASE '15)</source>
          . ACM, New York, NY, USA, Article
          <volume>14</volume>
          , 14 pages.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Carver</surname>
            <given-names>J. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hassler</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hernandes</surname>
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kraft</surname>
            <given-names>N. A..</given-names>
          </string-name>
          <article-title>Identifying barriers to the systematic literature review process</article-title>
          .
          <source>In Empirical Software Engineering and Measurement</source>
          , 2013 ACM/IEEE International Symposium on, pages
          <fpage>203</fpage>
          -
          <lpage>212</lpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Cohen</surname>
            <given-names>A. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hersh</surname>
            <given-names>W. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peterson</surname>
            <given-names>K.</given-names>
          </string-name>
          , and Yen P. Y..
          <article-title>Reducing workload in systematic review preparation using automated citation classification</article-title>
          .
          <source>Journal of the American Medical Informatics Association: JAMIA</source>
          ,
          <volume>13</volume>
          (
          <issue>2</issue>
          ):
          <fpage>206</fpage>
          -
          <lpage>219</lpage>
          ,
          <year>2006</year>
          . ISSN 1067-5027.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Tomassetti</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rizzo</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vetro</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ardito</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torchiano</surname>
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Morisio</surname>
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Linked Data approach for selection process automation in Systematic Reviews</article-title>
          .
          <source>In Proceedings of the 15th Annual Conference on EASE</source>
          , pp.
          <fpage>31</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Kitchenham</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brereton</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turner</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niazi</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Linkman</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pretorius</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Budgen</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <article-title>Refining the systematic literature review process - two observer participant case studies</article-title>
          ,
          <source>Empirical Software Engineering</source>
          <volume>15</volume>
          (
          <issue>6</issue>
          ) (
          <year>2010</year>
          )
          <fpage>619</fpage>
          -
          <lpage>653</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Kitchenham</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sjøberg</surname>
            <given-names>D. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brereton</surname>
            <given-names>O. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Budgen</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dybå</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Höst</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>and Runeson P..</surname>
          </string-name>
          <article-title>Trends in the quality of human-intensive software engineering experiments: a quasiexperiment</article-title>
          . ieee,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Dyba</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dingsøyr</surname>
            <given-names>T</given-names>
          </string-name>
          ..
          <article-title>Strength of evidence in systematic reviews in software engineering</article-title>
          .
          <source>Proceedings of the Second ACM-IEEE international symposium on Empirical software engineering and measurement</source>
          .
          <source>2008</source>
          , p.
          <fpage>178</fpage>
          -
          <lpage>187</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <surname>P. H. R</surname>
          </string-name>
          . U. in Oxford. Critical appraisal skills programme. http://www.casp-uk.net/,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Kitchenham</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sjøberg</surname>
            <given-names>D. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brereton</surname>
            <given-names>O. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Budgen</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dybå</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Höst</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>and Runeson P..</surname>
          </string-name>
          <article-title>Can we evaluate the quality of software engineering experiments</article-title>
          ?
          <source>In Proceedings of the 2010 ACM-IEEE International Symposium on Empirical Software Engineering and Measurement</source>
          , ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Dieste</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grimán</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juristo</surname>
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Saxena</surname>
            <given-names>H.. “</given-names>
          </string-name>
          <article-title>Quantitative Determination of the Relationship between Internal Validity and Bias in Software Engineering: Consequences for Systematic Literature Reviews,”</article-title>
          <source>Proc. Int'l Symp. Empirical Software Eng. and Metrics</source>
          , pp.
          <fpage>285</fpage>
          -
          <lpage>288</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Biolchini</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mian</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Natali</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Conte</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Travassos</surname>
            <given-names>G</given-names>
          </string-name>
          ..
          <article-title>Scientific research ontology to support systematic review in software engineering</article-title>
          .
          <source>Advanced Engineering Informatics</source>
          , vol.
          <volume>21</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>133</fpage>
          -
          <lpage>151</lpage>
          ,
          <year>2007</year>
          . DOI:
          <volume>10</volume>
          .1016/j.aei.
          <year>2006</year>
          .
          <volume>11</volume>
          .006.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Knaublock</surname>
          </string-name>
          , H..
          <source>Protégé-OWL</source>
          .
          <year>2003</year>
          . Avaiable http://protege.stanford.edu. Accessed on
          <volume>07</volume>
          /03/
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Labs</surname>
            ,
            <given-names>H..</given-names>
          </string-name>
          <article-title>Jena: A free and open source Java framework for building Semantic Web and Linked Data applications</article-title>
          .
          <year>2010</year>
          . Avaiable at http://jena.apache.org. Accessed on
          <volume>07</volume>
          /03/
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maynard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bontcheva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tablan</surname>
          </string-name>
          , V.:
          <article-title>GATE: A framework and graphical development environment for robust NLP tools and applications</article-title>
          .
          <source>In: Proc. of the 40th Anniversary Meeting of the ACL</source>
          . (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Witte</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>Q.</given-names>
          </string-name>
          , Zhang Y. and
          <string-name>
            <surname>Rilling</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <source>Ontological Text Mining of Software Documents, 12th International Conference on Applications of Natural Language to Information Systems (NLDB</source>
          <year>2007</year>
          , Paris, France, June 27- 29,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Bo</surname>
            <given-names>W.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Yunqing</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <article-title>"Research on the Design of the Ontology-Based Automatic Question Answering System,"</article-title>
          <source>Computer Science and Software Engineering</source>
          , 2008 International Conference on , vol.
          <volume>5</volume>
          , no., pp.
          <volume>871</volume>
          ,
          <issue>874</issue>
          ,
          <fpage>12</fpage>
          -
          <lpage>14</lpage>
          Dec.
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Juristo</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreno</surname>
            <given-names>A. M.</given-names>
          </string-name>
          ,
          <source>Basics of Software Engineering Experimentation</source>
          , Springer Publishing Company, Incorporated,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>