<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Two-step Approach for Effective Detection of ⋆ Misbehaving Users in Chats</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Esaú Villatoro-Tello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Juárez-González</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hugo Jair Escalante</string-name>
          <email>hugojair@ccc.inaoep.mx</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Montes-y-Gómez</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luis Villaseñor-Pineda</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Technologies Department, Universidad Autónoma Metropolitana (UAM), Unidad Cuajimapla</institution>
          ,
          <addr-line>Mexico City</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Laboratorio de Tecnologías del Lenguaje, Coordinación de Ciencias Computacionales Instituto Nacional de Astrofísica</institution>
          ,
          <addr-line>Óptica y Electrónica (INAOE)</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <abstract>
        <p>This paper describes the system jointly developed by the Language Technologies Lab from INAOE and the Language and Reasoning Group from UAM for the Sexual Predators Identification task at the PAN 2012. The presented system focuses on the problem of identifying sexual predators in a set of suspicious chatting. It is mainly based on the following hypotheses: (i) terms used in the process of child exploitation are categorically and psychologically different than terms used in general chatting; and (ii) predators usually apply the same course of conduct pattern when they are approaching a child. Based on these hypotheses, our participation at the PAN 2012 aimed to demonstrate that it is possible to train a classifier to learn those particular terms that turn a chat conversation into a case of online child exploitation; and, that it is also possible to learn the behavioral patters of predators during a chat conversation allowing us to accurately distinguish victims from predators.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        It is well known that the World Wide Web (WWW) has vastly penetrated into social
living and allows connecting people from different geographic regions through new forms
of communication. Examples of such communication forms are instant messaging
services, chat rooms, social networks (e.g., Facebook, Twitter, etc.) and blogs. These
services have become very popular tools for personal as well as for group communication,
as they are cheap, easy to use, virtual and private in nature [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Such online services allow users to hide their personal information behind the
monitor; which, on the one hand, makes this type of communication a source of fun, but
on the other hand, it also represents a threat. The privacy and virtual nature of these
services augment the possibilities of some heinous acts which one may not commit in
the real world. Examples of such acts are the online paedophiles who “groom” children,
⋆ This work was done under partial support of CONACYT (project grants 134186 and 106013).</p>
      <p>
        We also thank SNI-Mexico, INAOE and UAM for their assistance.
that is, who meet underage victims online, engage in sexually explicit text or video chat
with them, and eventually convince the children to meet them in person. According
to the National Center for Missing &amp; Exploited Children [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and the Office of
Juvenile Justice and Delinquency Prevention3, one out of every seven children receives an
unwanted sexual solicitation online.
      </p>
      <p>
        Traditionally, a term that is used to describe such malicious actions with a potential
aim of sexual exploitation or emotional connection with a child is referred as “Child
Grooming” or “Grooming Attack” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which has been defined by [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] as: a
communication process by which a perpetrator applies affinity seeking strategies, while
simultaneously acquiring information about and sexually desensitizing targeted victims in order
to develop relationships that result in need fulfilment (e.g. “physical sexual
molestation.”).
      </p>
      <p>Nowadays, the usual way to catch these sexual predators is for trained law
enforcement officers or volunteers to pose as children in online chat rooms, thus predators fall
into the trap and are identified. However, online sexual predators always outnumber the
law enforcement officers and volunteers. An organization that employs this
methodology to catch sexual predators is the Perverted Justice group, located in the United
States. This organization has been able to convict more than 500 predators since 20044.
Nevertheless, there is a great need for software applications that can flag suspicious
online chats automatically, either as a tool to aid law enforcement officials or as parental
control features offered by chat service providers.</p>
      <p>In this paper we propose a novel methodology that faces the problem of sexual
predators identification as a text classification task by means of a supervised approach.
The identification process is divided in two main stages: the Suspicious Conversations
Identification (SCI) stage and the Victim From Predator disclosure (VFP) stage.
Performed experiments showed that it is possible to train a classifier to learn those
particular terms that turn a chat conversation into a case of online child exploitation; and, that
it is also possible to learn the behavioral patters of predators during a chat conversation
allowing us to accurately distinguishing victims from predators.</p>
      <p>The rest of the paper is organized as follows. Section 2 presents recent work on the
task of sexual predators identification. Section 3 describes the proposed methodology
for identifying both suspicious chat conversations and the sexual predator within a chat
conversation. Section 4 presents the experimental settings as well as the results achieved
in the context of the PAN 2012 competition. Finally, Section 6 depicts our conclusions
and formulates directions for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Traditionally, the problem of identifying grooming attacks has been tackled through
text classification strategies. In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] the problem is reduced to the task of
distinguishing predators from victims within chat conversations. Such conversations are known
to be cases of grooming attacks, particularly they used a set of 701 conversations
obtained from the perverted-justice web page. In order to solve the problem, Pendar et al.
3 http://www.ojjdp.gov/
4 http://www.perverted-justice.com/
separated those interventions that belong to the set of victims from those that belong
to the set of predators, i.e., a two-class problem. Next, authors remove the stopwords,
and then computed word unigrams, bigrams and trigrams. Subsequently, they processed
them with the classification algorithms Support Vector Machine (SVM) and k-nearest
neighbors (k-NN). Authors performed several experiments varying the number of
features from 5000 to 10000, and concluded that the k-NN algorithm with k equal to 30 and
employing a feature vector of 10000 elements provides the most effective classification
resulting in a value of f -measure that equals to 0.943.
      </p>
      <p>
        Similarly, in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Michalopoulos et al. proposed a decision making method to be
used for recognizing potential grooming threats by extracting information from
captured dynamic dialogues. Their proposed method makes use of the following three
classification classes: i) Gaining Access: indicate predators intention to gain access to the
victim; ii) Deceptive Relationship: indicate the deceptive relationship that the predator
tries to establish with the minor, and are preliminary to a sexual exploitation attack; and
iii) Sexual Affair: clearly indicate predator’s intention for a sexual affair with the victim.
The classification process computes the probability that a captured dialogue belongs to
each one of the above classes. At the end, their system decides if there is a threat in the
conversation by means of a linear combination of the computed probabilities. For their
experiments, authors employed a set of 219 chat conversations (73 for each class), they
removed all the stopwords and applied a spelling correction strategy. Michalopoulos et
al. concluded that Naïve Bayes is the most appropriate technique, not only for reaching
the highest average classification score of 96%, but also for being fastest than all the
other algorithms that were evaluated.
      </p>
      <p>
        Finally, the work proposed by RahmanMiah et al. faces the problem of grooming
attack identification from a more general point of view [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Contrary to the work
proposed in [
        <xref ref-type="bibr" rid="ref3 ref5">5,3</xref>
        ], the authors of [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] define a method for identifying when a conversation
is a case of child exploitation instead of detecting which user is the predator directly.
The system proposed by RahmanMiah et al. defines three classes of conversations: i)
Child Exploitation: cases of grooming attacks; ii) Sexual Fantasies: conversations
between adults with a high degree of sexual content; and iii) general chatting: general
conversations with no sexual content. The proposed system applies traditional text
categorization techniques in combination with psychometric and categorical information
provided by LIWC (Linguistic Inquiry and Word Count [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]). For their experiments,
authors employed a set of 392 conversations, they do not apply any pre-processing to
the texts neither a spelling correction process. Authors conclude that psychometric and
categorical information can be used by classifiers as a feature set to predict the
suspected child exploitation in chats. These psychometric features significantly improve
the performance of Naïve Bayes classifiers to predict child-exploitation type chats.
      </p>
      <p>Our proposal differs from previous works in that it attacks both problems at once,
i.e., we are able to identify when a chat conversation is a case of child exploitation and
subsequently we are able to tell which user is the sexual predator. Thus our proposal can
be used for both purposes (detecting suspect conversations and identifying predators);
besides, we show that the two-step approach outperforms a single-stage method.</p>
    </sec>
    <sec id="sec-3">
      <title>Proposed Method</title>
      <p>The proposed method for detection of misbehaving users in chats is based on two main
hypotheses: (i) terms used in the process of child exploitation are categorically and
psychologically different than terms used in general chatting; and (ii) predators usually
apply the same course of conduct pattern when they are approaching a child. Accordingly,
we propose a new methodology for solving the problem of sexual predators
identification, which is divided in two main stages: the Suspicious Conversations Identification
(SCI) stage and the Victim From Predator disclosure (VFP) stage. Figure 1 shows the
general architecture of the proposed system.</p>
      <p>
        Notice that the goal of the first stage is to act as a filter, i.e., it helps distinguishing
between general chatting and possible cases of online child exploitation; in this way,
the set of conversations to be analyzed by the VFP module will be reduced. Hence we
can focus only on conversations that potentially include sexual predators for a more
fine grained analysis. Consequently, the goal of the second stage is to identify (to point
at) the potential predator from a possible case of child exploitation. The associated
classification problem is less complex than trying to discriminate between predators
and typical users directly, see Section 4.4.
Our proposed system does not include any module for preprocessing the texts, i.e.,
we did not remove any punctuation mark, stopwords and, neither apply a stemming
process. The main reason for not applying any preprocessing was because the text in
chat conversations had unique characteristics that distinguish them from any other type
of text [
        <xref ref-type="bibr" rid="ref2 ref6 ref7">2,6,7</xref>
        ], for example, chat conversations do not follow any grammar rules (i.e.,
are grammatically informal and unstructured), plus the frequent orthographical errors
and the common use of abbreviations and emoticons.
      </p>
      <p>
        We believe that using emoticons and intentional misspelled words may contain
valuable contextual information in a chat text. For example, in the grooming phase the
perpetrator may amend the relationship by an emphasized “soryyyyyyyyy” when the child
felt threatening by any obtrusive language. Another example may be the emoticon for
“hug (&gt;:d&lt;)” and “kiss (:-*)” for a soft introduction of sexual stage. Preserving such
information makes traditional language processing tools, such as stemmers and POS
taggers, unsuitable for processing the chat texts [
        <xref ref-type="bibr" rid="ref2 ref6">2,6</xref>
        ].
3.2
      </p>
      <sec id="sec-3-1">
        <title>Filtering stage</title>
        <p>Although we did not apply any pre-processing stage, it is important to mention that
we did apply a pre-filtering stage to all the conversations that were given for the PAN
2012 sexual predator competition. The goals of the pre-filtering stage were: i) to help
us focusing only in the most important cases and, ii) to reduce the computational cost
for automatically processing all the information.</p>
        <p>This pre-filtering stage consisted in removing all the conversations that accomplish
at least one of the following conditions: 1) Conversations that had only one
participant, 2) Conversations that had less than 6 interventions per-user and 3) Conversations
that had long sequences of unrecognised characters (apparently images). Table 1 shows
information about the training data before and after applying this pre-filtering stage.</p>
        <p>As it can be seen, by means of the pre-filtering stage we are able to reach a
substantial reduction ratio (90% approximately) of conversations/users. It is important to
notice that by doing this pre-filtering, we also removed a few sexual predators. Thus,
even if our proposed system works perfectly we will not be able to identify the 100%
of the sexual predators. Nevertheless, we think the information from interventions of
removed predators was not enough to effectively recognize them as predators anyways.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Suspicious Conversations Identification</title>
        <p>As we have mentioned before, our system faces the problem of sexual predators
identification as a text classification task (TC). Accordingly, for training the SCI classifier
(Figure 1) we employed traditional TC techniques to construct a model that distinguish
between general chatting and cases of child exploitation.</p>
        <p>In order to properly train our SCI classifier, we labeled as suspicious conversations
all the chat conversations were at least one predator appears, resulting in 5790
nonsuspicious conversations and 798 suspicious ones. During the experimentation phase,
the SCI classifier represents the conversations by means of the bag of words
representation (BOW) employing either a boolean or a TF-IDF weighting scheme. Since we
did not apply any preprocessing stage to the chat texts, we obtained features vectors of
117015 elements.
3.4</p>
      </sec>
      <sec id="sec-3-3">
        <title>Victim From Predator disclosure</title>
        <p>Similarly to the SCI classifier, our VFP stage was designed using traditional TC
techniques. The goal of the VFP classifier was to recognize sexual predators in suspicious
chat conversations, as detected by the SCI method. For training the VFP classifier we
divided all the text conversations, where a predator was involved, into interventions.
This means that if a text chat involved two different users, we had two sets of
interventions. Therefore, we used as examples of victims the interventions of the users that had
a conversation with a predator, resulting in 194 examples of victims; and as examples
of predators we used the interventions of the 136 users already labeled as predators.
For the experiments performed, the VFP classifier employs a BOW representation
using either a boolean or a TF-IDF weighting scheme. For the VFP classifier we obtained
features vectors of 16709 elements.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Setup</title>
      <sec id="sec-4-1">
        <title>Data set</title>
        <p>For our experiments we used the data set provided in the context of the PAN 2012 Lab:
Sexual Identification competition. As we mentioned in Section 3.2 we were given for
training a total of 66928 different chat conversations, where 97690 different users are
involved and only 148 are tagged as sexual predators (See Table 1). Additionally, a test
data set for evaluation was provided. Such corpus contained 155129 chat texts, where
218702 different users are involved and only 250 are tagged as sexual predators. For a
more detailed explanation on how the training and test corpora were constructed please
visit http://pan.webis.de/.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Classification methods</title>
        <p>
          Two classifiers from the CLOP toolbox [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] are used in the text classification
experiments; these are Neural Networks (NN) and Support Vector Machines (SVM)
classifiers. The NN classifier was set as a two layer neural network with a single hidden layer
of 10 units. For the SVM we tried linear and polynomial kernels.
        </p>
        <p>During the development phase we adopted two-fold cross validation to estimate the
performance of our methods using training data only. This validation technique was
used for all of our experiments. For the final evaluation of our system we used the test
data provided by organizers of PAN 2012, see Section 4.1. An analysis of the results
using both training and test data sets is given in the following section. The evaluation
of training-set results was carried out mainly by means of the classification accuracy,
which indicates the overall percentage of text chats correctly classified. Additionally,
due to the class imbalance, we also report results in terms of F1 measure. Regarding the
final evaluation of the system on test data, we used the measures proposed by organizers,
namely: F-0.5 measure, precision and recall.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Baseline definition</title>
        <p>As baseline we used the traditional paradigm for solving the problem of sexual
predators identification. Figure 2 provides a general view of the baseline definition.</p>
        <p>As can be noticed, the problem of identifying sexual predators is performed in one
single step. For the baseline experiment we employed a BOW representation using
either a boolean or a TF-IDF weighting scheme. By following the same procedure
established for the SCI and the VFP classifiers, under this configuration we obtained features
vectors of 117015 elements.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Experimental results</title>
        <p>In this section we report experimental results obtained by the components of our
twostep approach, as well as the results obtained with the baseline method, using training
data. Next, in Section 4.5, we report the performance of the system in the test data set.
Baseline results. Table 2 shows results obtained by the baseline configuration. In order
to evaluate if a dimensionality reduction strategy could be helpful for the classifier, we
performed several experiments varying the size of the features vectors that were used to
train the classifier. For this purpose we employed the well known information gain (IG)
method to rank and preserve the most distinctive features.</p>
        <p>Recall that the size of the features vector is 117015 elements, hence using only 10%
of the features means that the classifiers employed only 11,702 features to represent the
11,038 chat conversations. The baseline configuration faces the problem of learning a
model that helps to classify between normal users and sexual predators from a highly
unbalanced corpora, i.e., 10,902 normal users and only 136 sexual predators.</p>
        <p>Table 2 indicates that using a binary weighting scheme allows a better performance
for the classifier. It is also possible to observe that reducing the dimensionality of the
feature’s vector does not allow a significant improvement. On the contrary, it decreases
the ability of the classifier to accurately distinguish between normal users and sexual
predators. Because of these results we decided to not use any dimensionality reduction
strategy for both SCI and VFP classifiers.</p>
        <p>SCI results. Table 3 shows results obtained from the SCI classifier. As mentioned in
previous sections, the aim of this classifier is to distinguish between general chatting
and possible cases of online child exploitation. It is worth mentioning that the
training data employed by the SCI classifier represented an unbalanced corpus, since there
are 5790 conversation labeled as general chatting and only 798 text chats labeled as
cases of online child exploitation, however, obtained results showed that it is possible
to accurately distinguish suspicious conversations.</p>
        <p>Experimental results showed that both classifiers (i.e., SVM and NN) are suitable
for solving the problem of classifying suspicious conversations. Contrary to the results
obtained with the baseline configuration, the SCI classifier using SVM as classification
method obtained better results when chat conversations are represented by means of a
BOW considering a tf-idf weighting scheme.</p>
        <p>We concluded from these experiments that, terms used in the process of child
exploitation are categorically and psychologically different than terms used in general
chatting, which allows to train a classifier to learn those particular terms and accurately
detect cases of online child exploitation.</p>
        <p>VFP results. Table 4 shows results obtained from the VFP classifier. As we mentioned
in Section 3.4, the aim of this particular module is, once a conversation has been tagged
as a suspicious, to point at the sexual predator, i.e., to tell which user is the victim and
which one is the predator.</p>
        <p>Obtained results showed that the proposed methods are adequate for solving the
problem of classifying victims and predators. Similarly to the results obtained in the
SCI classifier, using SVM as classification method obtained better results when a tf-idf
weighting scheme was employed. Nevertheless, for the case of the VFP classifier, the
best results were obtained when using NN algorithm and a binary weighing scheme.</p>
        <p>From the experiments performed in this section we were able to show evidence that
suggest predators apply the same course of conduct pattern when they are approaching a
child, and that our proposed method it is able learn these behavioral patters of predators
during a chat conversation allowing us to accurately distinguish victims from predators.
4.5</p>
      </sec>
      <sec id="sec-4-5">
        <title>PAN 2012 competition results</title>
        <p>Previous sections described obtained results employing the training data set, and all
experiments were performed in a controlled scenario. However, the main goal of the PAN
2012 Lab was to evaluate in real scenarios the proposed method for sexual predators
identification. This section reports official results obtained by our system on test data
from the PAN 2012 competition.</p>
        <p>In order to apply our proposed methodology, the first step consisted on applying the
filtering stage (Section 3.2) to test data. Table 5 shows some statistics of the test data
before and after applying the filtering stage. From this table we can observe similar
reduction ratios as in training data. A total of 28 predators were removed by our filtering
approach. We manually analyzed the interventions of removed predators and found
that most of them contained a few characters only; we consider that with such little
information it is not possible to effectively identify to removed predators.</p>
        <p>After filtering the test data, we were able to apply our proposed method (Figure 1).
The next step was to represent all the remaining conversations (i.e., 15,330) into the
generated model for the SCI classifier. Following, the chat conversations that the SCI
classified as suspicious were divided into interventions and represented accordingly to
the model generated for the VFP classifier.</p>
        <p>Our team submitted three different runs: i) Baseline: it corresponds to the
configuration showed in Figure 2 employing as classification method a Neural Network and
using a binary weighting scheme with no dimensionality reduction; ii) SCI(NN-B)&amp;
VFP(NN-TF-IDF): it means that the SCI module was configured for using a NN and
a binary weighting scheme, whereas the VFP module used a NN with a tf-idf
weighting scheme; and iii)SCI(NN-B)&amp; VFP(NN-B): it means that both the SCI and the VFP
modules were configured for using a NN and a binary weighting scheme.</p>
        <p>Official results of submitted runs are showed in Table 6. The leading evaluation
measure was F-0.5 measure, which emphasizes the importance of the precision of the
system, which is particularly important for sexual predator detection as mentioned by
the organizers of the PAN 2012 competition.</p>
        <p>As it is possible to observe, the best configuration was the third one, i.e., using
in both modules a binary weighting scheme. Thus showing that the hypotheses of our
work were right. Indeed the configuration SCI(N N − B)&amp;V F P (N N − B) obtained
the highest performance among the 16 teams that participated in the sexual predator
identification track of PAN 2012. The closest entry (snider12-run-2012-06-16-0032)
achieved an F-0.5 measure of 0.9168, whereas the average was of 0.5105. The results
obtained by our group are promising and motivate us to pursuing several future work
directions.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5 Identifying predators’ bad behavior</title>
      <p>An additional task to that of sexual predator detection, proposed by organizers of PAN
2012, was to search for those lines (interventions) that reveal predator’s bad
behavior. Traditionally, such lines are manually identified and used as evidence to convict
paedophiles. PAN 2012 participants were encouraged to propose new ideas to
automatically solve this issue.</p>
      <p>
        We approached the line-detection task with a language-models based approach. Our
main idea was based in the following statement; it is well known that every predator
follows three main stages when approaching a child [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]: i) gain access to the victim,
ii) involve the victim in a deceptive relationship and, iii) launch and prolong a sexually
abuse relationship. Based on these facts, we believe that if we can generate language
models from each one of the stages mentioned above, we will be able to find those lines
that represent a bad behavior. We were particularly interested in the 2nd and the 3rd
stages, since from our point of view these should be the most critical sections within a
child exploitation chat conversation.
      </p>
      <p>Our proposed solution works as follows: first we automatically divide all the
conversations where a predator appears in three sections, such division is made without
considering any type of contextual frontiers, i.e., we did not identify where a child
approaching stage begins or ends. Next we generated the language model (lm) of the 2nd
and the 3rd parts5. Finally, for a user that is tagged as predator, we compute the
perplexity against the lm of each one of its interventions, and we delivered as the most
distinctive lines of bad behaving those with the minor perplexity value.</p>
      <p>For the competition, we proceed as follows: from the set of users labeled as sexual
predators by our system (SCI(NN-B)&amp; VFP(NN-B)), we select the 50 most distinctive
lines and delivered as examples of predators bad behaving. Table 7 shows examples of
the lines that were delivered. Official evaluation results indicated that by following this
procedure we were able to identify just 1 revealing line. We believe that our proposed
idea could be very effective, although more work is necessary in order to improve its
performance. Although this result seems absolutely negative we have to mention that
the winner participant of this sub-task submitted 63, 290 lines
(grozea12-run-2012-0614-1706b).</p>
      <p>From the 2nd parts From the 3rd parts
&lt;text&gt;what do you want me to be?&lt;/text&gt; &lt;text&gt;what do u want me to say&lt;/text&gt;
&lt;text&gt;what do u want me to say&lt;/text&gt; &lt;text&gt;what do u want me to wear&lt;/text&gt;
&lt;text&gt;what do u want me to wear&lt;/text&gt; &lt;text&gt;what do you want me to be?&lt;/text&gt;
&lt;text&gt;do u want to talk to me too&lt;/text&gt; &lt;text&gt;do u want me to do it&lt;/text&gt;</p>
      <p>Table 7. Examples of lines with the minor perplexity values.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>We have proposed a new methodology for detecting sexual predators in text chats. Our
proposal differs from traditional approaches in that it divides the problem in two stages:
5 We used the SML toolkit http://svr-www.eng.cam.ac.uk/ prc14/toolkit.html to this end.
the Suspicious Conversations Identification (SCI) stage and the Victim From Predator
disclosure (VFP) stage. The goal of the first stage is to work as a filter, i.e., it helps
distinguishing between general chatting and possible cases of online child exploitation;
in this way, the set of conversations to be analyzed will be reduced, hence we can
focus only on conversations that potentially include sexual predators for a more fine
grained analysis. Consequently, the goal of the second stage is to identify (to point at)
the potential predator from a possible case of child exploitation.</p>
      <p>Performed experiments showed that it is possible to train a classifier to learn those
particular terms that turn a chat conversation into a case of online child exploitation;
and, that it is also possible to learn the behavioral patters of predators during a chat
conversation allowing us to accurately distinguish victims from predators. Our
participation in the PAN 2012 forum showed that the proposed methodology is able to produce
very good results in a realistic scenario, obtaining an F-measure (β = 0.5) of 0.8936,
which was the best ranked result among all of the participants.</p>
      <p>
        As future work we plan to include some linguistic features, such as proposed by [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
We believe that the inclusion of such type of features can be helpful for increasing the
recall levels of our proposed system. In addition, we also believe that in the process of
identifying the interventions that depict the predator’s bad behavior this type of
information could be very helpful.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Harms</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Grooming</surname>
          </string-name>
          :
          <article-title>An operational definition and coding scheme</article-title>
          .
          <source>In Sex Offender Law Report</source>
          , Vol.
          <volume>8</volume>
          ,
          <issue>Num</issue>
          . 1. pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kucukyilmaz</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cambazoglu</surname>
            <given-names>B. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aykanat</surname>
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Can F</surname>
          </string-name>
          .
          <article-title>Chat mining: predicting user and message attributes in computer-mediated communication</article-title>
          .
          <source>In Information Processing and</source>
          Management Vol.
          <volume>44</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>1448</fpage>
          -
          <lpage>1466</lpage>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Michalopoulos</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Mavridis.</surname>
          </string-name>
          <article-title>Utilizing document classification for grooming attack recognition</article-title>
          .
          <source>In IEEE Symposium on Computers and Communications (ISCC</source>
          <year>2011</year>
          ), pp.
          <fpage>864</fpage>
          -
          <lpage>869</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>O</given-names>
            <surname>'Connell R. A</surname>
          </string-name>
          <article-title>Typology of Child Cyber- sexploitation and Online Grooming Practices</article-title>
          . In Cyberspace Research Unit, University of Central Lancashire.
          <year>2003</year>
          . Retrieved from http://image.guardian.co.uk/sys-files/Society/documents/2003/07/17/Groomingreport.pdf (accessed
          <year>August 2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Pendar</surname>
            <given-names>N.</given-names>
          </string-name>
          <article-title>Toward Spotting the Pedophile Telling victim from predator in text chats</article-title>
          .
          <source>In IEEE International Conference on Semantic Computing. Irvine California USA</source>
          , pp.
          <fpage>235</fpage>
          -
          <lpage>241</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>RahmanMiah</surname>
            <given-names>M. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yearwood</surname>
            <given-names>J.</given-names>
          </string-name>
          , and Kulkarni S.
          <article-title>Detection of child exploiting chats from a mixed chat dataset as text classification task</article-title>
          .
          <source>In Proceedings of the Australian Language Technology Association Workshop</source>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>165</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Rosa K. D.</surname>
          </string-name>
          , and Ellen J.
          <article-title>Text classification methodologies applied to micro-text in military chat</article-title>
          .
          <source>In Proceedings of the eight IEEE International Conference on Machine Learning and Applications (ICMLA '09)</source>
          , pp.
          <fpage>710</fpage>
          -
          <lpage>714</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Saffari</surname>
          </string-name>
          and
          <string-name>
            <given-names>I</given-names>
            <surname>Guyon</surname>
          </string-name>
          .
          <article-title>Quick Start Guide for CLOP</article-title>
          .
          <source>Technical report, Graz-UT and CLOPINET</source>
          , May,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Wolak</surname>
            <given-names>J</given-names>
          </string-name>
          ., Mitchell K., and Finkelhor D.
          <article-title>Online victimization of youth: Five years later</article-title>
          .
          <source>In National Center for Missing &amp; Exploited Children Builletin 07-06-025</source>
          , National Center for Missing &amp; Exploited
          <string-name>
            <surname>Children</surname>
            , Alexandia,
            <given-names>VA</given-names>
          </string-name>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>