<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Determining the proximity of groups in social networks based on text analysis using big data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A S Mukhin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>I A Rytsarev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>R A Paringer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A V Kupriyanov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D V Kirsh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Image Processing Systems Institute of RAS - Branch of the FSRC "Crystallography and Photonics" RAS</institution>
          ,
          <addr-line>Molodogvardejskaya street 151, Samara, Russia, 443001</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Samara National Research University</institution>
          ,
          <addr-line>Moskovskoe Shosse 34А, Samara, Russia, 443086</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>521</fpage>
      <lpage>526</lpage>
      <abstract>
        <p>The article is devoted to the definition of such groups in social networks. The object of the study was selected data social network Vk. Text data was collected, processed and analyzed. To solve the problem of obtaining the necessary information, research was conducted in the field of optimization of data collection of the social network Vk. A software tool that provides the collection and subsequent processing of the necessary data from the specified resources has been developed. The existing algorithms of text analysis, mainly of large volume, were investigated and applied.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Currently, social networks are booming: every day their users send billions of messages and leave
millions of comments under the relevant posts. The analysis of such content is of great importance
for many areas of business. For example, it is impossible to overestimate the impact of Internet
marketing on the promotion of goods and services. However, clear understanding of user requests is
essential to use these mechanisms effectively. The source of such information can be the materials
published by users of social networks, as well as the shares and reposts by users and the entire
communities. Thus, the issue of determining the proximity of groups in the social network
Vkontakte using the BigData technology, considered in this paper, is certainly a relevant objective
and a task of great scientific importance in the field of data analysis.</p>
      <p>
        Data processing from social networks is very popular now. For example, in the article [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
proposes a text normalization with deep convolutional character level embedding (Conv-char-Emb)
neural network model for SA of unstructured data. This model can tackle the problems: (1)
processing the noisy sentence for sentiment detection (2) handling small memory space in word level
embedded learning (3) accurate sentiment analysis of the unstructured data. In the article [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], authors
introduce SS3, a novel supervised learning model for text classification that naturally supports these
aspects. SS3 was designed to be used as a general framework to deal with ERD problems. In the
article [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], authors propose a nonparametric model (NPMM) which exploits auxiliary word
embeddings to infer the topic number and employs a “spike and slab” function to alleviate the
sparsity problem of topic-word distributions in online short text analyses. NPMM can automatically
decide whether a given document belongs to existing topics, measured by the squared Mahalanobis
distance. In the article [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], examine the long-term relationship between signals derived from nine
years of unstructured social media microblog text data and financial market developments in five
major economic regions. Employing statistical language modeling techniques we construct
directional sentiment metrics. In the article [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the authors propose a background clustering
technology for discussion. Compared with the traditional methods, background future clustering
keeps the constrains caused by data sparseness and spatio-temporal dependence off, and can be used
for unpredictable activities discovery
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Social network data collection</title>
      <p>
        The social network Vkontakte was selected as a data source for this study [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The reasons for this
choice are as follows:
• the network provides open access to its data (no restrictions on accessing the server data);
• Vkontakte is the most popular social network in Russia and the fifth most popular social
network in the world;
• Vkontakte is a full-fledged social network (unlike Twitter and Instagram, which are
microblogs) allowing to create thematic communities, which are particularly interesting for
this study.
      </p>
      <p>As part of this study, a Python software package was developed, containing an authorization
module, a data collection module, and a filtration module. This software package allows to collect
data and filter them to take the relevant information only.</p>
      <p>Within this study, the developed software package was used to collect more than 8,000 posts and
over 280,000 comments on them from the two most popular communities of the city of Samara
(“Podslushano Samara” and “Uslyshano Samara”) and from the community of the Samara
University students (“Podslushano Samarsky Universitet”).</p>
      <p>Streaming data obtained from social networks contains a lot of service information. Only the
relevant data is important for further analysis; therefore, it is necessary to separate the service
information from the relevant data. The software package pre-processing module structures the
collected data and filters the relevant and the service fields.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Determination of the proximity of groups using BigData technology</title>
      <p>
        To determine the proximity of groups, several metrics for the comparison of word indexes were
considered: Euclidean distance, city-block distance and Mahalanobis distance [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. The Euclidean
distance was chosen, since it is most suitable for this experiment according to the following criteria:
1) It is the most widely used and universal metric;
2) The Euclidean distance is calculated based on the original, not the standardized data.
      </p>
      <p>
        To calculate this metric, attribute vectors were formed between the groups by combining two
word indexes or more into a common one [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Weight was assigned to each word in the word index,
thus each group took the form of a vector of attributes (words) with own weights. In this paper, it
was decided to use the word frequency count as the weight [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ].
      </p>
      <p>
        Such an approach for calculating the weights of words in word indexes using traditional methods and
technologies requires huge computational resources and takes a long time when the volume and the
number of analyzed word indexes increases, so it was decided to use BigData technology and
computational clusters for this purpose [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. At this stage, an algorithm involving MapReduce
technology was developed, that rejected non-informative parts of the word index (words consisting
of less than three or more than fifteen characters) and also counted the frequency of words in the
text. As a result, three word indexes were developed, the elements of which had their own weights,
one of them is presented in Fig. 1.
      </p>
      <p>At the next step, it was decided to use two word indexes (of the groups “Podslushano Samara”
and “Uslyshano Samara”) to get a common word index, and to use the other word index (of the
Data Science
group “Podslushano Samarsky Universitet”) for test counting. The common word index consisted of
overlapping words with the weights recalculated according to formula 2.
in the first word index;  2 is the weight of the word in the second word index.
where:  ( 1,  2) is the weight of the word in the common word index;  1 is the weight of the word</p>
      <p>The last step was to calculate first the distances between the resulting word index and the groups
it was based on, and then the distance between the resulting word index and the test group. The
Euclidean formula of distance between the two groups was used to measure this value.
 ( 1,  2) =
 1 +  2</p>
      <p>( ,  ) =</p>
      <p>(  −   )2

 =1
(1)
(2)
The results are provided in the Table 1.</p>
      <p>In order to show the applicability of the proposed method for calculating distances between
groups using a common dictionary, we calculated the distances between groups without using a
common dictionary and with its use. Table 3 presents the results of calculations without a common
dictionary. Table 4 presents the results of calculating the Euclidean distances between all pairs of
groups and the templates of common dictionaries built on their basis.</p>
      <p>Comparing the obtained results, we can notice that the distances for calculations using a common
dictionary are less than calculations without it. From this we can assume that the use of the method
of finding a common dictionary for calculating Euclidean distances is justified and applicable for
solving the problem posed.</p>
    </sec>
    <sec id="sec-4">
      <title>5. The study of the dependence of the volume of the general dictionary used to calculate the</title>
    </sec>
    <sec id="sec-5">
      <title>Euclidean distance between groups</title>
      <p>We investigate at what volume of a general dictionary the results of determining the degree of
similarity of groups among themselves give the most informative readings. To do this, we carry out
an experimental calculation of the Euclidean distances for groups numbered 1 and 2 between them
and their common vocabulary by changing the dimensions of the common vocabulary. For the study,
we will choose the size of the dictionary equal to the greatest number of unique words for the second
group (18.948 words), the small size of the general dictionary (300 words) and several intermediate
values. Table 5 shows the results of calculations of this experiment.</p>
      <p>After analyzing the results obtained, it can be noted that for the anomalously large and, on the
contrary, anomalously small size of the general dictionary, the results turned out to be as
noninformative as possible. Most likely this is explained by the fact that with a small dictionary, for the
most part, only the most common words that do not carry more information and are approximately
equally found in the texts of both groups, for the maximum size of a common dictionary, the
situation is fundamentally opposite, tk. Many rare words come into account that are found only in
one of the groups, and therefore the results show such an abnormally large scatter. When analyzing
the results produced for the intermediate sizes of the general dictionary, it is seen that the values of
the distances cease to have strong leaps relative to each other when using a common dictionary of
about 3/5 of the amount of unique words for the group with the highest number. Such a result is due
to the fact that with such a volume the most non-informative words and words are cut off.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>Within the framework of this study, a set of software modules was developed allowing to determine
the distance between the communities of the social network Vkontakte. As a result of the work, a
common word index was compiled, on the basis of which the degrees of proximity between 3
communities were determined. In the future, the results of the work can be used to develop
algorithms for determining the proximity of larger groups and communities using the BigData
technology.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was financially supported by the Russian Foundation for Basic Research under grant #
1929-01135, # 18-37-00418, # 17-01-00972 and by the Ministry of Science and Higher Education
within the State assignment to the FSRC “Crystallography and Photonics” RAS No.
007GZ/Ch3363/26 (theoretical results).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Arora</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kansal</surname>
            <given-names>V 2019</given-names>
          </string-name>
          <article-title>Character level embedding with deep convolutional neural network for text normalization of unstructured data for Twitter sentiment analysis Social Network Analysis and Mining 9(1) 12</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Burdisso</surname>
            <given-names>S G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Errecalde</surname>
            <given-names>M</given-names>
          </string-name>
          and
          <string-name>
            <surname>Montes-</surname>
          </string-name>
          y
          <string-name>
            <surname>-Gómez M 2019</surname>
          </string-name>
          <article-title>A text classification framework for simple and effective early depression detection over social media streams</article-title>
          <source>Expert Systems with Applications</source>
          <volume>133</volume>
          <fpage>182</fpage>
          -
          <lpage>197</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Chen</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gong</surname>
            <given-names>Z</given-names>
          </string-name>
          and
          <string-name>
            <surname>Liu W 2019 A Nonparametric</surname>
          </string-name>
          <article-title>Model for Online Topic Discovery with Word Embeddings Information Sciences</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Groß-Klußmann</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>König</surname>
            <given-names>S</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ebner</surname>
            <given-names>M 2019</given-names>
          </string-name>
          <article-title>Buzzwords build Momentum: Global Financial Twitter Sentiment and the Aggregate Stock Market Expert Systems with Applications</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Zhu</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <article-title>Du J 2018 Background feature clustering and its application to social text</article-title>
          <source>Information Processing Letters</source>
          <volume>136</volume>
          <fpage>44</fpage>
          -
          <lpage>48</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Xu</surname>
            <given-names>X 2007</given-names>
          </string-name>
          <string-name>
            <surname>Scan:</surname>
          </string-name>
          <article-title>a structural clustering algorithm for networks Proceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining 824-833</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Rytsarev</surname>
            <given-names>I A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozlov</surname>
            <given-names>D D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kravtsova</surname>
            <given-names>N S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kupriyanov</surname>
            <given-names>A V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liseckiy</surname>
            <given-names>K S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liseckiy</surname>
            <given-names>S K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paringer</surname>
            <given-names>R A</given-names>
          </string-name>
          and
          <string-name>
            <surname>Samykina N Yu 2018</surname>
          </string-name>
          <article-title>Application of the principal component analysis to detect semantic differences during the content analysis of social networks</article-title>
          <source>CEUR Workshop Proceedings</source>
          <volume>2212</volume>
          <fpage>262</fpage>
          -
          <lpage>269</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Rytsarev</surname>
            <given-names>I A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kupriyanov</surname>
            <given-names>A V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirsh</surname>
            <given-names>D V</given-names>
          </string-name>
          and
          <string-name>
            <surname>Liseckiy K S 2018</surname>
          </string-name>
          <article-title>Clustering of social media content with the use of BigData technology</article-title>
          <source>Journal of Physics: Conference Series</source>
          <volume>1096</volume>
          (
          <issue>1</issue>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Rytsarev</surname>
            <given-names>I A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirsh</surname>
            <given-names>D V</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kupriyanov</surname>
            <given-names>A V</given-names>
          </string-name>
          <year>2018</year>
          <article-title>Clustering of media content from social networks using BigData technology</article-title>
          <source>Computer Optics</source>
          <volume>42</volume>
          (
          <issue>5</issue>
          )
          <fpage>921</fpage>
          -
          <lpage>927</lpage>
          DOI: 10.18287/
          <fpage>2412</fpage>
          - 6179
          <string-name>
            <surname>-</surname>
          </string-name>
          .
          <source>-2018-42-5-921-927</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Mikhaylov</surname>
            <given-names>D V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozlov A P and Emelyanov G M 2016Extraction</surname>
          </string-name>
          <article-title>of knowledge and relevant linguistic means with efficiency estimation for the formation of subject-oriented text sets</article-title>
          <source>Computer Optics</source>
          <volume>40</volume>
          (
          <issue>4</issue>
          )
          <fpage>572</fpage>
          -
          <lpage>582</lpage>
          DOI: 10.18287/
          <fpage>2412</fpage>
          -6179-2016-40-4-
          <fpage>572</fpage>
          -582
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Rytsarev</surname>
            <given-names>I A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kupriyanov</surname>
            <given-names>A V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirsh</surname>
            <given-names>D V</given-names>
          </string-name>
          and
          <string-name>
            <surname>Liseckiy K S 2018</surname>
          </string-name>
          <article-title>Clustering of social media content with the use of BigData technology</article-title>
          <source>Journal of Physics: Conference Series</source>
          <volume>1096</volume>
          (
          <issue>1</issue>
          ) DOI:
          <fpage>10</fpage>
          .1088/
          <fpage>1742</fpage>
          -6596/1096/1/01208
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Kropotov</surname>
            <given-names>Y A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Proskuryakov</surname>
            <given-names>A Y</given-names>
          </string-name>
          and
          <string-name>
            <surname>Belov</surname>
            <given-names>A A</given-names>
          </string-name>
          <year>2018</year>
          <article-title>Method for forecasting changes in time series parameters in digital information management systems</article-title>
          <source>Computer Optics</source>
          <volume>42</volume>
          (
          <issue>6</issue>
          )
          <fpage>1093</fpage>
          -
          <lpage>1100</lpage>
          DOI: 10.18287/
          <fpage>2412</fpage>
          -6179-2018-42-6-
          <fpage>1093</fpage>
          -1100
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>