<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Content analyses for the Psychological Purposes: Requirements to Software Supporting Tools</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>N. Almayev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Psychology, Russian academy of sciences</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>249</fpage>
      <lpage>254</lpage>
      <abstract>
        <p>This paper deals with the problem of adequacy of software tools for content analyses within the scope of psychology. Understanding of text by humans and processing of it by the software is considered. Existing practices are criticized and the conditions under which software tools may support human analyses are outlined. Two types of the psychological content analyses tasks are described with the summarization of requirements for both of them.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>Key words</kwd>
        <kwd>words</kwd>
        <kwd>judgments</kwd>
        <kwd>parser</kwd>
        <kwd>dictionary</kwd>
        <kwd>human expert</kwd>
        <kwd>computer learning</kwd>
        <kwd>knowledge base</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Attempts of texts analyses with the regard to author’s personality are older than
questionnaires and even task based psychological tests. According to W. Stern, Alfred
Binet before the development of the first IQ test was engaged in the study of the ways
of how the writer’s personality is expressed in the texts of his origin. Andrey Beliy
was adopting content analyses tools for investigation of Pushkin’s, Baratisnkiy’s and
Tutschev’s projective attitudes in their descriptions of nature as early as 1916. In
1930-50s first after the works of Murray &amp; Morgan [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and then after these of
McClelland, et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] content analyses of projective stories was the leading paradigm
in psychology and even now it holds strong positions in the domain of psychotherapy.
      </p>
      <p>
        Content analyses of projective texts, interviews and discussions have number of
advantages in comparison to psychological questionnaires. First of this advantages is
the freedom of choice for the topics and words for their expression – the subjects tell
us what is relevant for them and not choose between the variants of what is interesting
for us but may have no relation to the needs of a subject. The second advantage is that
projective stories possess permanent value, i.e. they may be used indefinite number of
times with ever new and corrected content analyses scales, while answers in
questionnaires are bond to respected questions. The third advantage is the amount of personal
information that may be acquired in addition to one that is gathered through the
generalized scales of content analyses. Generalized scales provide means for comparisons
of groups [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] while additional information may serve for deeper personality
analyses of a subject. Finally, concurrent use of questionnaires and content analyses of
projective stories provide complementary results those reveal both actual needs of the
subjects and specific features of ego-concept.
      </p>
      <p>
        Nevertheless, content analyses possess one major disadvantage in comparison to
the questionnaires based tests: it is the amount of labor that is required for its’
fulfillment. In order to work properly human coder has to deal with one – two scales at
once [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] this mean that each text has to be read several times depending on the
amount of the scales! Even in the case that all of the scales are coded simultaneously
(that inevitably will cause the growth of the number of mistakes) human processing of
big corpora of texts takes much effort and time.
      </p>
      <p>
        The need to overcome this main shortage of content analyses resulted in numerous
attempts of computerization of its routines those may be traced to 1970s [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and even
to 1960s, see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Contemporary state of computer based content analyses</title>
      <p>
        Surprisingly, but since the first attempts and till now, see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] the main paradigm
of content-analyses computerization has not changed significantly. It is still based on
the counting of single words in a text those are believed to relate to certain domains of
meaning (categories). As the target words are aggregated in dictionaries this paradigm
is named “dictionary based” [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Contemporary tools [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] permit finding of
collocations of words as well as grammatical parsing, in the sense that target words may be
used in the different grammatical forms; nevertheless it doesn’t change the essence of
the paradigm.
      </p>
      <p>
        Despite some positive feedback about the usage of dictionary based software tools
in psychological studies [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], generally psychologists are not confident in such tools.
Sometimes they can work but miss too much as well as produce false alarms too
often. Humans can express one and the same meaning using different words, and
depending on which word is connected to a target word the meaning of the whole
changes. For example, such words as “try”, “persistent” should refer to the
Achievement motive. But what if they are joined with “try to avoid”, “or persistently tried not
to meet with somebody”? These expressions correspond to avoidant behavior while
basing on the target words will result in counting them for the Achievement motive.
Psychological meaning of one and the same expression may change even dependently
on agens or patiens role of the teller. For example “I have beat him” is aggression, but
“I was beaten by him” is suffering from aggression, they should be not confused.
      </p>
      <p>Such considerations may raise skepticism towards the whole endeavor of computer
assessed content-analyses in psychological purposes. What kind of knowledge base it
must have in its foundation in order to be relevant and adequate? How this knowledge
base should be extracted from the experts? What is the appropriate form of storage of
this knowledge? And the most important one – when such software can acquire
practical usefulness?</p>
      <p>In order to transform this skepticism into the valuable source of making software
tools more adequate and useful reconsideration of the basics of the meaning
phenomenon is required and for this we have to return to the issue, the origin of logic, rhetoric
and text analyses.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The decisive role of judgment</title>
      <p>“No one of these terms [categories, single words], in and by itself, involves an
affirmation; it is by the combination of such terms that positive or negative statements
arise. For every assertion must, as is admitted, be either true or false, whereas
expressions which are not in any way composite such as 'man', 'white', 'runs', 'wins', cannot
be either true or false. (Aristotle “Categories”, Part 4. Translated by E. M. Edghill)</p>
      <p>Thus, in the very beginning of his “Organon” Aristotle teaches us (to put it
straight) that isolated words possess no actual meaning. They acquire it only within
the judgment i.e. proper junction of words, determined by the rules of grammar. Not
“collocations” but judgments should be the subject of content analyses. Moreover,
target judgments may be not only of the simple “S is P” type but rather complicated
predicative structures possibly of the several levels with some obligate and some
optional members. Practically it means that the collection of words in a text first should
be transformed into the collection of judgments and then within the latter collection
the search for target expressions has to be performed. Taking into consideration that
automatic parsers for different natural languages exist already for decades this crucial
step is not something that lays beyond the scope of contemporary technologies.
4</p>
    </sec>
    <sec id="sec-4">
      <title>The problem of context and the necessity of constant knowledge acquisition for content analyses software</title>
      <p>Grammatical parsers may solve the first stage of the meaning problem – find proper
grammatical junctions, but due to the polysemy they cannot solve the second part of
the meaning problem – that of the denotation i.e. relation of a judgment to the objects
of some consistent reality. For example, “naked conductor” is about a human or about
a wire? Humans in solution of such problems base on the very complex ontological
knowledge whereas one and the same words and even more or less similar
expressions can suite for different domains. In content analyses for the purposes of
psychology it is impossible to build such complicated ontology in advance. It may be formed
only gradually on the basis of computer learning of the software in constant
interaction with the human experts. The need for such constant interaction is also
determined by the following obstacles: 1) not all of the expressions can be
monosemantically resolved basing solely on the grammar rules, 2) not all of the
expressions are input according to the grammar rules by humans or speech recognition
software. The third and the most decisive reason for constant learning of such a
system is that humans produce ever new ways of speaking about the same matters, thus
categories of content should always be updated.</p>
      <p>It also has to be stated, that such systems must never acquire full autonomy.
Particularly, the possible use of content analyses in forensic procedures requires special
responsibility. Basing on the previous experience software supporting tools should
find and propose the most probable variants of certain expressions categorizations but
the final decision has to be left for the human expert, or at least humans must have
full access to the any step of analyses and categorization.</p>
      <p>Therefore, the software under discussion cannot be realized simply on the basis of
“neural” networks those somehow “learn” by changes of the weights of their elements
through the positive or negative feedback and those are characterized with the very
problematic extraction of the probabilities from their respected elements. The
probabilities of judgments relations to categories must be clear for human experts, and it
must be clear for them basing on what kind of the previous judgments such relations
were calculated. It means that the system must be able to form concordance for each
expression or may be even for each member of judgments, be it in the role of a
subject, or a predicate, or a predicate to a predicate, etc..
5</p>
    </sec>
    <sec id="sec-5">
      <title>The two main tasks of psychological content analyses and further specification of requirements to software for both of them</title>
      <p>
        Historically first task of content analyses for the psychological purposes was study of
motivation, psychological states and affects via projective stories [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
State of affairs in this field is although being characterized by the vigorous endeavors
of dictionary based approach in the past, see [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] contemporary due to the matters
described above is almost totally within the usage of the human coding.
      </p>
      <p>Successful transition from the single words (dictionary based) paradigm to the
judgment based paradigm in the field of projective texts analyses requires
consideration of the following matters:
1. Scales for software based content-analyses should be psychometrically proved,
their validity and sensitivity must be established, their coding instruction and
coding practice must be elaborated. It is hardly possible to develop simultaneously the
psychological part and the automatization part of the project, both are under the
risk of failure in this case.
2. Content analyses scales must form the system, i.e. balance each other, be not
isolated; otherwise there always will be the threat of expansion of the isolated scales
to the neighboring domain.</p>
      <p>Generally, psychological part must be developed previously to the software part
development starts.</p>
      <p>The other direction of content analyses application is the study of discussions or
“discourse” analyses. French word “discourse” means nothing else but “reasoning”.
On the one hand, it means that there is hardly much new in this field since the eight
books of Aristotle’s “Topics”, but on the other hand, development of Internet caused
exteriorization of numerous discussions on various important social issues. Processing
of them entirely by the means of human experts is evidently impossible. At the first
glance transition to the judgment based content analyses may be considered as the
radical step to the adequate understanding of the matter, indeed, reasoning consists of
rhetorical syllogisms (“enthymemes”) and the latter consist of the judgments
(propositions). Nevertheless, this field possesses many features those are specific for it both
in comparison to the projective stories content analyses and in comparison to the other
fields of reasoning, e.g. scientific.</p>
      <p>Aristotle comparing rhetorical syllogism to the “apodictic” one says that the major
premise of enthymeme is not something necessary and universal but something in
which the majority of concrete audience or the most prominent of them (judges,
rulers) will believe. This “majority” implies the study of amount of those who can share
respected views and in the case of the Internet discussions – activity of the users in
propagating them. Therefore, unlike previous task, internet discussions content
analyses demand registration of multiuser activity. The following variables should at least
be registered: 1) unique names (nicknames) of the participants, 2) total number of the
participants, 3) number of posts by each of the participants, 4) amount of the text from
each of the participant in 4a) absolute terms, and 4b) relative amount of text by each
participant, 5) nature of the message i.e. 5a) topical content – “reasoning”, 5b) link to
some content, 5c) pictorial (including the moving pictures of various forms) content.</p>
      <p>Unlike the case of projective stories the expert’s work in discourse analyses does
not presuppose fixed content categories. Experts are supposed to concentrate on the
polysemy tasks resolution. Do participants of discussions speak about one and the
same objects or the objects are different? Are the characteristics of the objects are
generally the same or differ? The experts work appears to be even more complicated
if taken into consideration that major premises in enthymemes are often truncated
(and maybe not even reflected by the participants of discussions). Instead of the
content categories the judgments within Internet discussions may be accumulated to some
major propositions (more or less equal statements, possibly with some gradual
differences) and the popularity of such should be studied alongside with the activity of
different users in propagation of them. It also has to be noted that grammatical
subjects and predicates should be differentiated from the ontological ones. For example,
in old and widespread expression “capitalist pig” grammatical S is “pig”, while
“capitalist” is the P. Ontologically of course relation is reversed, it is the affectively loaded
characteristic of the certain class of objects in the social world.</p>
      <p>Due to the high risks of subjectivity and biases in expert’s evaluation the whole
process of concrete judgments aggregation into the major propositions should be clear
and transparent for the human experts.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>Despite the differences between the two main tasks of the psychological content
analyses they have principal common feature – necessity of transition from the dictionary
based paradigm to the judgment based one. Technically this transition whereas not
simple is possible.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Financing</title>
      <p>The study was supported by the Russian Foundation for Basic Research, project №
18-00-00605 00606 (18-00-00605).
8</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Almayev</surname>
            <given-names>N.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murasheva</surname>
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Bessonova</given-names>
            <surname>Yu</surname>
          </string-name>
          .V.,
          <string-name>
            <surname>Kiselyova</surname>
            <given-names>N.I.</given-names>
          </string-name>
          <article-title>Content-analyses scales of the social motivation test. Results of correlation and factor analyses. Part 2</article-title>
          .
          <string-name>
            <surname>Eksperimental</surname>
          </string-name>
          <article-title>'naâ psihologiâ [Experimental Psychology</article-title>
          (Russia)],
          <year>2018</year>
          . Vol.
          <volume>11</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>108</fpage>
          -
          <lpage>119</lpage>
          . doi:
          <volume>10</volume>
          .17759/exppsy.2018110308. (In Russ., аbstr.
          <source>in Engl.)</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Almayev</surname>
            <given-names>N.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murasheva</surname>
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Bessonova</given-names>
            <surname>Yu</surname>
          </string-name>
          .V.,
          <string-name>
            <surname>Kiselyova</surname>
            <given-names>N.I.</given-names>
          </string-name>
          <article-title>Generalized scales of content analysis of projective narratives in test of social motivation (TSM)</article-title>
          .
          <source>Their validity and specifics. Part 1</source>
          .
          <string-name>
            <surname>Eksperimental</surname>
          </string-name>
          <article-title>'naâ psihologiâ [Experimental Psychology</article-title>
          (Russia)],
          <year>2016</year>
          . Vol.
          <volume>9</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>90</fpage>
          -
          <lpage>104</lpage>
          . doi:
          <volume>10</volume>
          .17759/exppsy.2016090409. (In Russ., аbstr.
          <source>in Engl.)</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gottschalk</surname>
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gleser G.C.</surname>
          </string-name>
          <article-title>The Measurement of Psychological States through the content analysis of Verbal Behavior</article-title>
          . Berkley: University of California Press,
          <year>1969</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hogenraad</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>McKenzie D.P.</given-names>
            ,
            <surname>Peladeau</surname>
          </string-name>
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Force and influence in Content Analysis: The Production of new social Knowledge // Quality and</article-title>
          <string-name>
            <surname>Quantity.</surname>
          </string-name>
          <year>2003</year>
          .
          <volume>37</volume>
          (
          <issue>1</issue>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Litvinova</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Litvinova</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Panicheva</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biryukova</surname>
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2018</year>
          )
          <article-title>Using Corpus Linguistics Tools to Analyze a Russian-Language Islamic Extremist Forum</article-title>
          . In Bodrunova S. (eds)
          <article-title>Internet Science</article-title>
          .
          <source>INSCI 2018. Lecture Notes in Computer Science</source>
          , vol
          <volume>11193</volume>
          . Springer, Cham
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>McClelland D.C.</surname>
          </string-name>
          <article-title>Power the inner experience</article-title>
          . Irvington Publishers. NY,
          <year>1975</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>McClelland</surname>
            <given-names>D.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atkinson</surname>
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lowell</surname>
            <given-names>E.L.</given-names>
          </string-name>
          <article-title>The achievement Motive</article-title>
          . Princeton, NJ: Van Nostrand,
          <year>1953</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Murray</surname>
            <given-names>H.A.</given-names>
          </string-name>
          (Ed.) Explorations in Personality. NY: Oxford University Press,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Scott</surname>
          </string-name>
          , M.
          <source>Oxford WordSmith Tools Version</source>
          <volume>4</volume>
          .0. Available at: http://www.lexically.net/downloads/version4/wordsmith.pdfl (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Vanheule</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Desmet</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meganck</surname>
            <given-names>R</given-names>
          </string-name>
          .
          <article-title>What the heart thinks the tongue speaks: a study on depression</article-title>
          and lexical choice // Psychological Reports.
          <year>2009</year>
          .
          <volume>104</volume>
          .
          <fpage>P473</fpage>
          -
          <lpage>481</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Winter</surname>
            <given-names>D.</given-names>
          </string-name>
          <article-title>The Power motive</article-title>
          . NY: Free Press,
          <year>1973</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>