<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Technologies used in the Data Analysis Process: Measuring Customer Satisfaction in context</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Clemens Forster</string-name>
          <email>cforster@a1.net</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Vienna University of Economics and Business</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The objective is to transform indicator-coded variables extracted from open-ended questions by using semantic technologies rather than manual coding, in order to compare two forecast models and evaluate which yields the best results. Experiments are carried out within the domain of customer satisfaction. For this purpose, a leading Austrian financial institution provides a large sample of survey data. In the questionnaire, respondents are asked explicitly about their overall customer satisfaction. This dependent variable is to be predicted by use of statistical analysis techniques. For the process of model building, a procedure is required for generating coding categories in order to enable the use of unstructured verbatim data on critical incidents. One approach is based on traditional methods, while the other uses ontologies and reasoning.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic Technologies</kwd>
        <kwd>Customer Satisfaction</kwd>
        <kwd>Forecast Models</kwd>
        <kwd>Survey Data</kwd>
        <kwd>Verbatim</kwd>
        <kwd>Model Building</kwd>
        <kwd>Ontologies and Reasoning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Semantic technologies enable new forms of collaboration beyond the boundaries of
applications. Through this inventive kind of technology, computers are able to put
information to a wider variety of uses and provide a novel basis for reasoning, and
this results in a direct effect on quality.</p>
      <p>
        Quantitative research methods can be supplemented with semantic technologies to
make use of unstructured data from qualitative research. Companies might use the
integration of such procedures in knowledge management systems as a strategic
advantage [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This can result, for example, in better understanding of customer
perspective as well as enhancing controlling-correlated key performance indicators
tied to balanced scorecards.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Problem</title>
      <p>The coding of open-ended responses to survey questions by applying semantic
technologies to utilize the information in the data analysis process is a new, and
interdisciplinary area of application. Hardly any explicit literature referring to this
topic exists. Combining “data and text mining” in business forecasting is also
something which has only relatively recently been demonstrated experimentally.
There have so far been no documented or published evaluations on the application of
semantic technology in the coding of open questions, or its integrated usage within
the scope of data mining utilized in building predictive modeling methods. Although
applications for such combinations have been presented by well-known software
companies, as solutions with high potential for the near future, the scientific
evaluation of the integration of semantic technologies in forecast analytics is missing.</p>
    </sec>
    <sec id="sec-3">
      <title>State of the Field</title>
      <p>Terms of reference and state of the field vary between the borders of standard
quantitative and qualitative research procedures.</p>
      <sec id="sec-3-1">
        <title>Quantitative research methods in context</title>
        <p>
          In the field of quantitative research methods, there is a wide variety of statistical and
mathematical analysis procedures to choose from. The concept of “data mining” [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ],
is deployed for the purpose of pattern recognition i.e. in gathering new levels of
understanding to make forecasts [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], frequently in connection with algorithms which
are used on large databases and show the most efficient possible asymptotic
consumption of computer resources.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Connection with qualitative research methods</title>
        <p>
          In analytics and interpretation within qualitative research procedures, computer-aided
techniques are becoming increasingly more used in order to analyse and interpret
methodically evaluated data such as text files [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Despite the fact that critics warn of
a possible adverse affect to survey design, creativity [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and the procedure of
exhaustive analysis, their use is nonetheless becoming more and more commonplace.
        </p>
        <p>
          Transcribed text files from qualitative interviews, such as expert interviews, are
becoming more commonly analysed through QDA (qualitative data analysis)
applications in correlation with concepts like “the stages of open, axial and selective
coding in the grounded theory”, “thematic framework” or the “application of the most
essential technique in objective hermeneutics” [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>Responses to open questions determined by a specific survey design are regularly
coded (founded on similar concepts used in QDA applications) to take advantage of
statistical analysis. ESOMAR (European Society for Opinion and Marketing</p>
        <p>Semantic Technologies used in the Data Analysis Process 3
Research) defines the process of coding as: “The organising of responses into
categories and the assignment of a unique numerical code to each response prior to
data entry.”</p>
        <p>This transformation into quantitative components is subsequently used for analysis
and can lead, in principle, based on the resultant additional information, to an
improvement in predictive modelling. Manual coding efforts are nevertheless time
consuming and subject to human error when coders develop certain rules for
classifying ambiguous cases. Part of the conceived work is to evaluate a new method
for improving this coding procedure through the application of semantic technologies.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Proposed Approach</title>
      <p>The aim of the thesis is to develop and deploy an application by assembling
opensource components capable of handling both the mathematical and statistical
evaluation of analytical methods, as well as the classification of answers to
openended questions. This will be carried out by the use of semantic technologies, in
contrast to the use of manual classification and coding. In order to demonstrate a
comparison between these experimental results, the quantification of any differences
will be shown empirically.</p>
      <sec id="sec-4-1">
        <title>Mathematical and statistical procedures</title>
        <p>
          Most procedures for automatic and quantitative content analysis track the occurrence
of specific words in n-dimensional space [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], or rather the co-occurrences [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] of
words, and then use the information on frequencies for further clustering.
        </p>
        <p>Rather than solely use these mathematical and statistical procedures, this paper will
focus on the extended possibilities offered by semantic technologies, such as the
classification of open answers to “meaningful higher-level unities” through reasoning
by rules and ontologies in combination with Natural Language Processing methods.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Pre-existing knowledge in the form of upper-level ontologies</title>
        <p>
          Automatic word sense disambiguation using a “light-weight ontology” like GermaNet
[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] together with opinion mining and sentiment analysis methods [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] similar to the
sentiment polarity classification will be applied.
        </p>
        <p>
          The classification process should be extended to pre-existing knowledge in the
form of upper-level ontologies like e.g. DOLCE, PROTON, SUMO [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] or DBpedia
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] to improve identifying key concepts for the further statistical analysis.
The procedural model is shown in the following figure:
Quality assurance (with respect to the code-plan, inter-coder reliability and
intracoder reliability) will be implemented in the manual coding process. The forecasting
method chosen for the experiment will be specified and applied in accordance with
the ceteris paribus clause. The other coding method is based on ontologies and
reasoning. The results of the experiment will yield data to determine which approach
offers greater accuracy in predicting the overall satisfaction variable.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Methodology</title>
      <p>For this purpose, a leading Austrian financial institution will provide a large sample
of survey data. An extensive dataset with 21,146 interviews from 35 customer
satisfaction projects/waves are available for analysis. The critical incident technique
was applied using questionnaires. This resulted in 8,242 positive comments and 9,854
negative comments, which can potentially be used for the coding.</p>
      <p>In the questionnaire, respondents were asked explicitly about their overall level of
customer satisfaction. This dependent variable is to be predicted through statistical
analysis techniques. The objective of this research is the theoretical foundation,
analysis and development of a coding process for forecast models with semantic
methodologies.</p>
    </sec>
    <sec id="sec-6">
      <title>Results</title>
      <p>These experiments have not been performed as of yet, and the results therefore remain
to be determined. The concept for the work has, however, been developed. The means
of obtaining the data have been established and the authorisation granted - a
confidential agreement has been signed and several discussions have been held
regarding the data-quality, as well as the procedure itself.</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions and future work</title>
      <p>The next steps will involve: development of the code-plan; performing the coding
process with several coders; and applying quality assurance. Data understanding, data
preparation and performing the modelling tasks are to follow subsequently. The
work’s conclusion will provide an answer to the research question of whether or not it
is possible to substitute and/or enhance manual coding through the use of semantic
technologies.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Dietz</surname>
            ,
            <given-names>J.L.G.</given-names>
          </string-name>
          :
          <article-title>Enterprise ontology: theory and methodology</article-title>
          . Springer, Berlin; NY (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Berry</surname>
            ,
            <given-names>M.J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Linoff</surname>
          </string-name>
          , G.:
          <article-title>Mastering data mining: The art and science of customer relationship management</article-title>
          . Wiley, New York, u.a. (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mertens</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albers</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Prognoserechnung. Physica-Verl.,
          <string-name>
            <surname>Heidelberg</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kuckartz</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Einführung in die computergestützte Analyse qualitativer Daten</article-title>
          . VS, Verl. für Sozialwiss.,
          <string-name>
            <surname>Wiesbaden</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lindsay</surname>
          </string-name>
          , V.J.:
          <article-title>Handbook of qualitative research methods for international business</article-title>
          . In: Marschan-Piekkari,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Welch</surname>
          </string-name>
          , C. (eds.). Elgar, Cheltenham [u.a.] (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Wernet</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Einführung in die Interpretationstechnik der objektiven Hermeneutik</article-title>
          . VS Verlag (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lourenço</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carreira</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glez-Peña</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Méndez</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carneiro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rocha</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Díaz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>E.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rocha</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fdez-Riverola</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rocha</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>BioDR: Semantic indexing networks for biomedical document retrieval</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>37</volume>
          (
          <issue>4</issue>
          ):
          <fpage>3444</fpage>
          -
          <lpage>3453</lpage>
          , (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mazanec</surname>
          </string-name>
          , J.:
          <article-title>Eine "Landkarte der Werbeforschung": Schlagwortvisualisierung am Beispiel der Zeitschrift "transfer - Werbeforschung &amp; Praxis"</article-title>
          . In: Strebinger,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Kurz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Mayerhofer</surname>
          </string-name>
          , W. (eds.): In: Werbe- und
          <string-name>
            <surname>Markenforschung</surname>
          </string-name>
          :
          <article-title>Meilensteine - State of the Art - Perspektiven</article-title>
          . Gabler, (Hrsg.)
          <string-name>
            <surname>Wiesbaden</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Finthammer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cramer</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Exploring and navigating: Tools for germanet</article-title>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Opinion mining and sentiment analysis</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          -2):
          <fpage>1</fpage>
          -
          <lpage>135</lpage>
          , (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mascardi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cordì</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lemnitzer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wunsch</surname>
          </string-name>
          , H.:
          <article-title>A comparison of upper ontologies. Acquisition of a New Type of Lexical-Semantic Relation from German Corpora</article-title>
          .: WOA07, Genova, Italy. Citeseer, IOS Press (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ayers</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raimond</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Interlinking open data on the web</article-title>
          .
          <source>4th European Semantic Web Conference</source>
          . (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>