<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Combining Predation Heuristics and Chat-Like Features in Sexual Predator Identification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>José María Gómez Hidalgo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrés Alfonso Caurcel Díaz</string-name>
          <email>andres.caurcel@optenet.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Optenet, Research and Development Department</institution>
          ,
          <addr-line>Las Rozas</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad Politécnica de Madrid, Dept. of Applied Intelligent Systems</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <abstract>
        <p>In this paper we present a system for sexual predator detection which combines two different approaches: a knowledge-based system that makes use of pattern matching according to hand-coded patterns that represent typical predator behaviors, and a learning-based system which employs surface linguistic features like capitalization and chat-like expressions. These approaches are combined in a chained fashion, being the learner applied to the suspicious predators as reported by the knowledge-based system. While the results of the system evaluation on the training collection are nice, the test run results for the official Sexual Predator Identification sub-task have shown much room for improvement.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        This paper describes the system designed and used by Optenet in the PAN Author
Identification task, Sexual Predator Identification sub-task. The system is intended to
validate our previous (unreported) research on predator detection in Spanish language, plus
our current research on age detection in Social Networks [
        <xref ref-type="bibr" rid="ref1 ref9">1,9</xref>
        ]3.
      </p>
      <p>In order to achieve this goal, we have adapted a previously existing
knowledgebased system that makes use of hand-coded predation patterns in Spanish language.
Additionally, we have augmented this system with learning capabilities based on
linguistic features, trying to separate adults from children, and from adults posing as
children. The overall results of our system when evaluated on the training collection are
nice, but the system has demonstrated very poor performance on the test runs. We
believe this is mainly due to three reasons: first, the adaptation of the knowledge-based
system has been done by the automatic translation of patterns from Spanish to English;
second, we do not make use of learning on any kind of word-level features (words, word
n-grams, character n-grams, etc.); finally, the previously existing knowledge-based
system makes use of some additional language-dependent mechanisms that we have not
adapted to English due to time restrictions.</p>
      <p>In the next sections, we present the general architecture and processing approach of
the system. We also describe the modules included in it, along with some examples and
conclusions.
3 These references are available in Spanish language only.</p>
    </sec>
    <sec id="sec-2">
      <title>General Architecture</title>
      <p>We have built a system that is composed of two modules:
1. A knowledge-based conversation filter (KBF), which processes single-speaker
conversations, and retains predator utterances while discarding most neutral and victim
conversations. This module analyzes each sentence and applies pattern matching to
detect suspicious ones, based on a set of hand-crafted patterns which correspond
to typical predator behavior. This sub-system largely reuses a previously existing
knowledge-based system for predator detection instant messaging conversations
written in Spanish.
2. A learning-based detection sub-system for Chat Language (CHL), which makes
use of chat-like and linguistic features, and it is trained on the retained
conversations. This sub-system highlights the candidate predators according to their
language/writing style. This approach is being tested on age detection in Social
Networks.</p>
      <p>We adopt a chaining approach. In the training phase, we have applied the following
process to the training data:
– All conversations are processed by the KBF, which reports a subset of users as
potential predators.
– Those conversations in which a user has been reported as a potential predator4, are
used as a training collection for the CHL. This module learns a classifier which is
stored for the classification phase.</p>
      <p>We roughly follow the same approach in the classification phase. The only
difference is that, instead of using the CHL sub-system to learn a classifier, we use it to
classify those users marked as suspicious by the KBF module.</p>
      <p>The Sexual Predator Identification sub-task requires not only to spot candidate
predators, but to highlight suspicious sentences in their conversations. During the
classification phase, our system reports the matched sentences by the KBF module in those
speakers reported by the whole system as predators.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The Knowledge-Based Conversation Filter</title>
      <p>
        The KBF builds on previously unreported research by Optenet on sexual predation
detection in chats, in the Spanish language. This module is based on the Spanish NGO
Protegeles [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] predator behavior characterization, which is quite similar to that reported
in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and latter used by Kontontathis et al. to build Chatcoder [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>It must be noted that chats used in the PAN task do not correspond to real grooming
cases, as the victim is a volunteer acting as a hook. Moreover, many of the cases are
quite fast5, while according to our experience in the real world, sex predators spend
4 Here we mean full conversations, including the victim or any neutral user in a false positive
example.
5 The harassment happens even in the first conversation between the predator and the volunteer.
several months when approaching and seducing a child before they meet on the real
world. Our previously existing system was designed to specifically target those slow
but more difficult cases.</p>
      <p>The KBF module features 51 hand-coded patterns in Spanish, which have been
automatically translated to English using the Google Translator in order to evaluate the
language portability of the system. The obtained patterns have not been corrected in the
case of translation mistakes. These patterns represent twelve typical predation behaviors
(BH) that correspond to four main predation phases:
1. Hooking: trying to locate the child (BH11), avoiding direct questions (BH12).
2. Fidelization: questions about family settings (BH21).
3. Seduction: personal questions (BH31), reducing sexual inhibition (BH32), sending
pictures (BH33), flattery (BH34), generating debt perception (BH35).
4. Harassment: nude and sexual pictures (BH41), virtual sex (BH42), trying to meet
(BH43), concept manipulation (e.g. child sex was accepted in ancient Greece –
BH44).</p>
      <p>The KBF includes some techniques for detecting other predator behaviors, but they
are not easy to port from Spanish to English. Due to the time limit of the competition,
we have deactivated those techniques in the English version of the system.</p>
      <p>In the table 1, we show several examples of patterns used to identify some of the
previous behaviors. The pattern-matching algorithm requires 2 or more unsorted words
in common with a sentence to fire, except for BH12 and BH44, in which a full ordered
match is required. We consider only the blank space as a word separator both in speaker
sentences and in patterns. A speaker is reported as suspicious if at least 10% of his/her
sentences match any behavior, or 3 or more in case of conversations with less than
ten sentences. These thresholds have been retained from the original KBF module in
Spanish.</p>
      <p>Behavior Pattern Behavior Pattern
BH11 “how old age” BH21 “where are your brothers mother father”
BH32 “consolation console touch” BH34 “handsome love you like you better”
BH41 “naked picture of you” BH44 “everyone does”</p>
      <p>We present an example of a speaker conversation tagged as suspicious by the KBF
sub-system in the table 2, along with the codes of the behaviors identified by the
patternmatching algorithm. We only report the lines corresponding to the suspicious speaker.
It must be noted in this example that there are some false positive lines (for instance,
several lines are recognized as BH31 while not being – personal questions), while other
lines that should be matched, do not (for instance, line 22 corresponds to BH34 –
flattery). We interpret this as a consequence of directly translating Spanish language
patterns to English without reviewing nor enriching them.</p>
      <p>Among the original training conversations provided in the task, which correspond to
139,573 speakers, this module reports 14,236 as potential predators, retaining 100% of
reported predators (recall = 1), and filtering out 89,89% of neutral or victim speakers.
4</p>
    </sec>
    <sec id="sec-4">
      <title>The Learning-Based Detection sub-system</title>
      <p>
        Over the years, some authors have been relatively successful in detecting the age of
speakers in online chats [
        <xref ref-type="bibr" rid="ref10 ref7">7,10</xref>
        ]. Based on these works, we have formulated the
hypothesis that linguistic and chat-like language properties can help to cluster speakers in three
groups: real kids speaking as digital natives, adults simulating kid language in chats
(as potential predators), and actual adults. Thus, we have selected six surface linguistic
properties as features for a learning-based module to be trained on the output of the
KBF sub-system:
– Number of uppercased letters.
– Number of words (guessing that chat-language tries to minimize communication
symbols).
– Number of SMS/Chat words (according to the dictionary available at [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]).
– Number of emoticons (according to the dictionary available at the same site).
– Number of typos (word that do not occur in the Freeling [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] English dictionary nor
in our SMS/Chat dictionary).
– Number of punctuation symbols (.,;:).
      </p>
      <p>The values for these features are taken as absolute numbers and as relative numbers
by dividing them among the number of lines in the conversation, and rounding them.
A vector of values is computed for each conversation speaker. In consequence, each
vector has twelve values. For instance, the values for the previous speaker conversation
example are the following ones:
70aca6a54d7d6b260273282143a685e0 ) [279; 255; 3; 24; 28; 266; 10; 9; 0; 0; 1; 10]</p>
      <p>
        We have used two classes (predator, neutral), and trained a WEKA [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] Naïve Bayes
Multinomial classifier with default parameters on the retained speakers from the KBF
sub-system. The evaluation of this classifier on the output of the KBF module taken
as training collection, and using 4-fold cross validation, leads to a recall of 0.986, a
precision of 0.639, and a F1 value of 0.775.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We have designed and evaluated an English-language sexual predator detection system
that largely reuses a previously existing one for the Spanish language. Our system
combines a knowledge-based predation detection system that makes use of hand-crafted
patterns, and a learning-based classifier that employs simple linguistic features.</p>
      <p>
        While the results of the evaluation on the training collection are nice, the
performance on the test run of the PAN Author Identification task, Sexual Predator
Identification sub-task, has been very poor. We believe this is mainly due to two reasons:
– First, the adaptation of the knowledge-based system has been done by the
automatic translation of patterns from Spanish to English. We believe that the obtained
patterns are not correct, so they must be reviewed by a native speaker with
experience on Internet chats. Moreover, they should be augmented with specific patterns
for the English language, and the numbers of matched patterns may be optimized
by a learning based classifier. It must be noted that the mistakes made by this
module may impact on the linguistic learning-based sub-system as well, making the
effectiveness of the learned classifier much worse than expected.
– Second, and according to the ground truth provided by the PAN organizers, it is
clear for us that word-level features (specifically, word n-grams and character
ngrams) should be used as well. Given the architecture of our system, we believe the
most straightforward way to do this, is training an additional text-based classifier
on the suspicious users reported by the KBF, or, most likely, on all the users, and
combine it with the linguistic classifier using stacking [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
– Third, we have not used some of the techniques currently implemented in the
knowledge-based system because they are highly language-dependent and due to
the time limit for submitting the test run results.
      </p>
      <p>We plan to do these improvements in future versions of our system, and to port them
to the Spanish language sexual predator detector as well.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The work described here has been partly funded by the Spanish Ministry of Economy
and by the Spanish Centre for Industrial Technological Development (CDTI), under
the research project “WENDY: WEb-access coNfidence for chilDren and Young”
(TSI020100-2010-452).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Gómez</given-names>
            <surname>Hidalgo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Caurcel</surname>
          </string-name>
          <string-name>
            <surname>Díaz</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Avances tecnológicos en la protección del menor en redes sociales</article-title>
          . Novática, Revista de la
          <source>ATI (218)</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>The WEKA data mining software: an update</article-title>
          .
          <source>SIGKDD Explorations Newsletter</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kontostathis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bayzick</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGhee</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leatherman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Comparison of rule-based to human analysis of chat logs</article-title>
          .
          <source>In: Proceedings of the First International Workshop on Mining Social Media (MSM09)</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Lingo2Word Home Page: http://www.lingo2word.
          <source>com (June</source>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Olson</surname>
            ,
            <given-names>L.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daggs</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ellevold</surname>
            ,
            <given-names>B.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>T.K.K.</given-names>
          </string-name>
          :
          <article-title>Entrapping the innocent: Toward a theory of child sexual predators' luring communication</article-title>
          .
          <source>Communication Theory</source>
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <fpage>231</fpage>
          -
          <lpage>251</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Padró</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Collado</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reese</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lloberes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castellón</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Freeling 2.1: Five years of open-source language processing tools</article-title>
          .
          <source>In: Proceedings of the International Conference on Language Resources and Evaluation</source>
          ,
          <string-name>
            <surname>LREC</surname>
          </string-name>
          <year>2010</year>
          ,
          <volume>17</volume>
          -
          <fpage>23</fpage>
          May
          <year>2010</year>
          , Valletta, Malta.
          <source>European Language Resources Association</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pendar</surname>
          </string-name>
          , N.:
          <article-title>Toward spotting the pedophile telling victim from predator in text chats</article-title>
          .
          <source>In: Proceedings of the International Conference on Semantic Computing</source>
          . pp.
          <fpage>235</fpage>
          -
          <lpage>241</lpage>
          . IEEE Computer Society, Washington, DC, USA (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Protegeles Home Page: http://www.protegeles.
          <source>com (June</source>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Santos</given-names>
            <surname>Sierra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Sánchez Ávila</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Carmonet</surname>
          </string-name>
          <string-name>
            <surname>Bravo</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Guerra</given-names>
            <surname>Casanova</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Santos</given-names>
            <surname>Sierra</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Control de edad en redes sociales mediante biometría facial</article-title>
          .
          <source>In: XII Reunión Española sobre Criptología y Seguridad de la Información (RECSI</source>
          <year>2012</year>
          )
          <article-title>(</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tam</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martell</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          :
          <article-title>Age detection in chat</article-title>
          .
          <source>In: Proceedings of the International Conference on Semantic Computing</source>
          . pp.
          <fpage>33</fpage>
          -
          <lpage>39</lpage>
          . IEEE Computer Society, Washington, DC, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wolpert</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          :
          <article-title>Stacked generalization</article-title>
          .
          <source>Neural Networks</source>
          <volume>5</volume>
          ,
          <fpage>241</fpage>
          -
          <lpage>259</lpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>