<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mixed Style Feature Representation and B0-maximal Clustering for Style Change Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Castro-Castro</string-name>
          <email>danielbaldauf@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlos Alberto Rodríguez-Losada</string-name>
          <email>carlosarl1999@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafael Muñoz</string-name>
          <email>rafael@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Software and Computing systems, Alicante University</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Oriente University</institution>
          ,
          <addr-line>Santiago de Cuba</addr-line>
          ,
          <country country="CU">Cuba</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>The goal of Style Change Detection task in a document is to determine if it was written by more than one author and in such case, to delimit which paragraph (or more generally a portion of text) corresponds to each one of them. The objective of our proposal is to build a paragraph representation based on general Style Feature computed considering characters, lexical and syntactic features, without the use of semantic words. The paragraphs were grouped employing a non overlapped variant of the B0-maximal clustering algorithm, where the overlapping was eliminated considering the order of paragraphs in the document.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Authorship detection is important for determining which author or group of authors
should get credit for writing a given document. In particular, in our digital modern
society, it is a complex task when the objective is to determine who wrote a piece of
digital text or if a document could be written by more than one author.</p>
      <p>
        Thanks to the research community and in particular to the organizers of PAN 3
evaluation forum [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], in recent years there is a growing interest in sharing methods and
algorithms to solve many of the tasks involved in Authorship Attribution (AA). One of
these tasks is the Style Change Detection in a document, with the purpose of detection
if a document was written by only one author or more than one, and in the last scenario,
what piece of text corresponds to each one of the authors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Overviews of the past style change detection task [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ][
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] resumed the description
of the task, approaches presented by participants and the results obtained. It is important
to highlight, that a priori, there is no information about authors or the numbers of them
involved in a problem, that’s why, the tasks are mainly solved considering text clustering
solutions.
      </p>
      <p>One of the key aspects tackled to solve the task, corresponds to, the representation
of textual contents, and in the majority of proposals, it was used the Bag of Word model
considering lexical and syntactical linguistic features. For clustering algorithms have
been used hierarchical and non-hierarchical traditional methods. Also, due to the nature
of the task, the clusters of documents may not be overlapped, because a document or a
paragraph (depending of the practical problem proposed) belongs to an unique author,
so then, a document or paragraph must be part of just one cluster (or group).</p>
      <p>In section 2 our proposal is described with emphasis in paragraph representation
based on the construction of a Mixed Style Set of Features and the clustering algorithm
employed. Section 3 presents the results obtained and the conclusions of the work, and
in Section 4 a brief discussion of the main problems.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Proposal for Style Change Detection - PAN 2020</title>
      <p>Our main goal is to determine clusters of paragraphs in which is considered that all
paragraphs in a cluster are written by the same author. If the algorithm obtains more than
one cluster, then the document was written by more than one author and the number of
distinct authors corresponds to the number of clusters.</p>
      <p>
        In the next two sections are described the representation of paragraphs and the
clustering algorithm. The computational representation is based on the formulations
proposed in Logical Combinatorial Pattern Recognition 4 and the clustering algorithm is
explained in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
2.1
      </p>
      <sec id="sec-2-1">
        <title>Paragraph Mixed Style Feature Representation and Similarity</title>
        <p>The objective of our proposal was to build a paragraph representation based on general
Style Feature computed considering characters, lexical and syntactic features, without
the use of content words. In Figure 1, are illustrated the paragraph representation and
similarity functions implemented to compare two of them.</p>
        <p>The paragraph representation is build considering a finite mixed set of 185 features
from three types of data values, Boolean, Float or n-gram Vector. At the left corner
in Figure 1 there is a section "Features examples" with three examples of the type of
features analyzed.</p>
        <p>Features were structured in six subsets considering different textual layers on the
text. These layers are boolean, character, sentence, paragraph, syntactic and the text.</p>
        <p>Examples for each of the layers subset of features:
1- Boolean layer: Uses the same word to finish a sentence and to begin the next sentence.
2- Character layer: Average length of words.
3- Sentence layer: Average number of words. Average number of distinct prepositions.
4- Paragraph layer: Average number of sentences. Average number of words.
5- Syntactic layer: Proportion of nouns over adjective.
6- Text layer: Average length of sentence. Bag of Words of conjunctions.
In order to compare two representations, one comparison criteria (CC) for each type of</p>
        <sec id="sec-2-1-1">
          <title>4 https://www.uci.cu/reconocimiento-logico-combinatorio-de-patrones</title>
          <p>Figure 1. Description of paragraph representation and similarity function
data feature value is introduced. The three CC formulas are exposed at the right section
in Figure 1. For features of Bool type, two features are similar, if they have the same
value (true or false), see CCbool. For features of Float type, two features are similar,
if the difference between values are less than a predefined threshold, see CCfloat. For
features of Vector type, two features are similar if the similarity between them is greater
than a predefined threshold, see CCvector. We used M inM ax5 similarity to compare
two vector. Finally, the two paragraphs are similar, if the number of features, in which
they are similar, are greater than a percentage defined, see F (Pi; Pj ).
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>B0-maximal Clustering Method</title>
        <p>The clustering proposal generates all the subsets of paragraphs in order to achieve that
the similarity between each of the paragraphs in a cluster should be larger than a
predefined B0 parameter. The B0-maximal clustering algorithm obtains compact groups of
paragraphs and some overlapped groups.</p>
        <p>For the task, this overlapping needs to be eliminated and we used an approach based
on the order of paragraphs in the document. If a paragraph could be part of two or more
clusters, it will be considered only in the cluster where a paragraph with the lower index
in the order of appearance in the text exists. This decision is based on the assumption
that the style in a document are characterized by the style reflected in the firsts
paragraphs and in general the main author tend to write the majority of the paragraphs and
the firsts one. To accomplish that, the overlapped paragraphs were sorted by their index
of appearance in the document. When the cluster assignment was defined, then all edges
from these paragraph to other clusters were eliminated.</p>
        <p>In Figure 2 and Figure 3 are presented an example based on a graph construction,
where the vertices are the paragraphs and the number of the vertices, the order of the
paragraphs in the document. The edges that connect two vertices represent that the
similarity of vertices is greater or equal than a B0 parameter and in our proposal we use
a percentage of similar features.</p>
        <p>The Figure 2 corresponds to the output of the clustering algorithm, and it can be seen
that paragraph 4 and 5 could be part of two clusters. Considering the heuristic explained
to eliminate the overlapping, the final clusters will correspond to the two illustrated at
Figure 3.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <p>
        The data-set distributed contains documents for two problems of Style Change
Detection, a narrow data-set and a wide data-set [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The description of the data and evaluation
measures are discussed on the overview published for the task.
      </p>
      <p>In Table 1, are resumed the average results for task1 and task2 considering results
in both data-set. For task1 the objective was to answer if a document was written by
one author or more than one. In task2 had to be answered, in which paragraph (could
be more than one) of the text there was a style change. As an additional data, the
organizers informed, that a maximum of three authors could be involved in a document,</p>
      <sec id="sec-3-1">
        <title>5 https://rdrr.io/cran/stylo/man/dist.minmax.html</title>
        <p>but our proposal is not restricted by a predefined number of clusters. Task1 was
evaluated by F1 measure and task2 using micro F measure. Using train and validation
data-set distributed for the task, it was selected the values for parameters B0, and ,
considering all combinations of these three.</p>
        <p>As a baseline we considered for task1 that the answer was always multi-authored,
and as the data-set are balanced in the number of problems for multi-authored and
single-authored documents, the result is 0.5. Similar result is obtained if the answer
were single-authored for all documents. For task 2 we could not compare the results
with a baseline.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>Using training and validation data-set, we got better results processing the wide
dataset than the narrow one, and this is interesting, considering that we did not use content
(topic related) words, concluding that the syntactic and structural style features are used
differently when the topics change. At the contrary, we got no significant difference of
style between authors when they wrote about the same topic.</p>
      <p>Several of the features get duplicated values, because they capture the same values,
considering that two or more structural layers are fused, for example, when the unit to
be analyzed as a document is a paragraph, then the paragraph layer and text layer are
considered distinct but they are the same.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>It was presented a proposal based on a paragraph representation, considering general
Style Features at character, lexical and syntactic layers of analysis, without the use
of topic or content words. The paragraphs were grouped employing a non overlapped
variant of the B0-maximal clustering algorithm, where the overlapping was eliminated
considering the order of the paragraph in the document.</p>
      <p>As future work could be interesting to combine semantic and topic vector
representation as features of the mixed model in order to distinguish between paragraph of
different topics. Also, the heuristic employed to eliminate the overlapping scenarios can
be improved, if some characteristics of the groups are considered, for example: the size,
strength of similarity or the adjacency of paragraphs.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Eva</given-names>
            <surname>Zangerle</surname>
          </string-name>
          , Maximilian Mayerl,
          <string-name>
            <surname>G.S.M.P.B.S.:</surname>
          </string-name>
          <article-title>Overview of the Style Change Detection Task at PAN 2020</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Eickhoff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Névéol</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . (eds.)
          <article-title>CLEF 2020 Labs and Workshops, Notebook Papers</article-title>
          .
          <source>CEUR-WS.org (Sep</source>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Gil-García</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Badía-Contelles</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pons-Porrata</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A parallel algorithm for incremental compact clustering</article-title>
          .
          <source>In: Euro-Par. Lecture Notes in Computer Science</source>
          , vol.
          <volume>2790</volume>
          , pp.
          <fpage>310</fpage>
          -
          <lpage>317</lpage>
          . Springer (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kestemont</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Overview of the author identification task at PAN-2018: cross-domain authorship attribution and style change detection</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Soulier</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . (eds.) Working Notes of CLEF 2018 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Avignon, France,
          <source>September 10-14</source>
          ,
          <year>2018</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>2125</volume>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2018</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2125</volume>
          /invited_paper_2.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiegmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>TIRA Integrated Research Architecture</article-title>
          . In: Ferro,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <surname>C</surname>
          </string-name>
          . (eds.)
          <article-title>Information Retrieval Evaluation in a Changing World</article-title>
          . Springer (Sep
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Overview of the author identification task at PAN-2017: style breach detection and author clustering</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          , T. (eds.) Working Notes of CLEF 2017 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Dublin, Ireland,
          <source>September 11-14</source>
          ,
          <year>2017</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <year>1866</year>
          .
          <article-title>CEUR-WS.org (</article-title>
          <year>2017</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-1866/invited_paper_3.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Zangerle</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Overview of the style change detection task at PAN 2019</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.E.</given-names>
            ,
            <surname>Müller</surname>
          </string-name>
          , H. (eds.) Working Notes of CLEF 2019 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Lugano, Switzerland, September 9-
          <issue>12</issue>
          ,
          <year>2019</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>2380</volume>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2019</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2380</volume>
          /paper_243.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>