<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Utilization of Information Interpolation using Geotagged Tweets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Masaki Endo</string-name>
          <email>e-mail endou@uitec.ac.jp</email>
          <email>endou@uitec.ac.jp</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Masaharu Hirota</string-name>
          <email>e-mail hirota@mis.ous.ac.jp</email>
          <email>hirota@mis.ous.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hiroshi Ishikawa</string-name>
          <email>hiroshi@tmu.ac.jp</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Okayama University of Science</institution>
          ,
          <addr-line>Okayama</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Polytechnic University</institution>
          ,
          <addr-line>Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Tokyo Metropolitan University</institution>
          ,
          <addr-line>Tokyo, Japan, e-mail ishikawa</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Along with the spread of social media, it has become possible to extract events occurring in the real world in real time. A benefit of analysis using data with position information is that it can accurately extract an event from a target area to be analyzed. However, because social media data include few data with location information, the amount to analyze is insufficient for almost all areas: we cannot fully extract most events. Therefore, efficient analytical methods must be devised for the accurate extraction of events with position information, even in areas with few data. For this study, we use geotagged tweets along with interpolation to estimate the best time to observe biological seasons when doing sightseeing such as cherry-blossom and autumn leaf viewing in areas and sightseeing spots. Herein, we explain the analysis results obtained using information interpolation and analysis of cherry blossoms in Japan during 2017.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        INTRODUCTION
Because of the wide dissemination and rapid performance
improvement of various devices such as smart phones and
tablets, diverse and vast data are generated on the web.
Particularly, social networking services (SNSs) have
become popular because users can post data and various
messages easily. Twitter [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], an SNS that provides a
microblogging service, is used as a real-time communication tool.
Numerous tweets have been posted daily by vast numbers
of users. Twitter is therefore a useful medium to obtain,
from a large amount of information posted by many users,
real-time information corresponding to the real world.
We specifically consider tourist information provision using
real-time information from Twitter. According to a survey
reported in the Inbound Landing-type Tourism Guide [2] by
the Ministry of Economy, Trade and Industry (METI),
tourists want real-time information and local unique
seasonal information posted on websites. Current websites
© 2018. Copyright for the individual papers remains with the authors.
Copying permitted for private and academic purposes.
      </p>
      <p>UISTDA '18, March 11, Tokyo, Japan.
provide similar information in the form of guidebooks.
Nevertheless, the information update frequency of that
medium is low. Because each local government, tourism
association, and travel company independently provides
information about travel destination locales, it is difficult
for tourists to collect information for “now” tourist spots.
Therefore, providing current, useful, real-world information
for travelers by capturing changes of information in
accordance with the season and relevant time period of the
tourism region is important for the travel industry.
Tourist information for best times requires a peak period,
which means that the best time is neither a period after or
before falling flowers, but a precisely defined period to
view blooming flowers. Furthermore, the best times differ
among regions and locations. Therefore, for each region
and location, it is necessary to estimate the best time for
phenological observations. Estimating best-time viewing
periods requires the collection of large amounts of
information having real-time properties. For this study, we
use Twitter data obtained from many users throughout
Japan. We use Twitter, a typical microblogging service, and
also use geotagged tweets that include position information
sent in Japan to ascertain the best time (peak period) for
biological season observation by region. The geotagged
tweets are useful as social indicators reflecting real-world
circumstances. They are a useful resource supporting a
realtime regional tourist information system in the tourism field.
Therefore, our proposed method might be an effective
means of estimating the best time to view events other than
biological seasonal observations.</p>
      <p>To analyze information of each region from Twitter data, it
is necessary to specify a location from tweet information.
Geotagged tweets can identify places. Therefore, they are
effective for analysis. However, because geotagged tweets
account for only a very small proportion of the total
information content of tweets, it is not possible to analyze
all regions. We propose a method to provide tourists with
information about sightseeing spots after processing a small
amount of data in real time by interpolation using
geotagged tweets.</p>
      <p>
        RELATED WORKS
Along with rising SNS popularity, real-time information
has increased. Analysis using real time data has become
possible. Many studies have examined efficient methods for
analyzing large amounts of digital data. Some studies have
been conducted to predict real world phenomena using
large amounts of social big data. Phithakkitnukoon et al. [
        <xref ref-type="bibr" rid="ref4">3</xref>
        ]
analyze details of traveler behavior using data from mobile
phone GPS location records such as embarkation places,
destinations, and traveling mode on a personal level.
Mislove et al. [4] develop a system that infers a Twitter
user’s feelings from tweet text and which visualizes
changes of emotion in space–time. Based on research to
detect events such as earthquakes and typhoons, Sakaki et
al. [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ] propose a method to estimate real-time events from
Twitter tweets. Cheng et al. [
        <xref ref-type="bibr" rid="ref7">6</xref>
        ] estimate Twitter users’
geographical positions at the time of their contributions,
without the use of geotags, by devoting attention to the
geographical locality of words from text information in
articles posted on Twitter. Yamada [7] analyzes Japanese
blog data and proposes a method to identify seasonal words
using simple autocorrelation analysis. Krumm et al. [
        <xref ref-type="bibr" rid="ref9">8</xref>
        ]
propose a method to detect events using time series analysis
of geotagged tweet volumes from localized areas. Various
studies have analyzed spatiotemporal data, but research to
estimate viewing periods using interlinkage is a new field.
OUR PROPOSED METHOD
We describe the best-time estimation method of organisms
by analysis using geotagged tweets that include organism
names. Best-time estimation, as defined for this paper, is
estimation of the period during which creatures at tourist
spots are useful for sightseeing. Such information can be
useful reference information when visiting tourist spots. It
supports estimation of the period during which a tourist can
enjoy the four seasons by viewing cherry blossoms and
autumn leaves. However, geotagged tweets are far fewer
than tweets without geotags. For that reason, although it is
possible to estimate the best time in a prefecture unit or
municipality, finely honed analyses have been impossible.
Nevertheless, the best time to visit sightseeing spots can be
estimated with finer granularity using the method with
interpolation proposed in this paper.
      </p>
      <p>
        In the following subsections, we describe the collection of
geotagged tweets to be analyzed, character preprocessing
for conducting analysis, and interpolation using Kriging.
Data collection
This section presents data collection. Geotagged tweets sent
from Twitter are a collection target. The range of geotagged
tweets includes the Japanese archipelago (120.0°E –
154.0°E, and 20.0°N – 47.0°N) as the collection target. The
collection of these data was done using a streaming API [
        <xref ref-type="bibr" rid="ref10">9</xref>
        ]
provided by Twitter Inc.
      </p>
      <p>
        Next, we describe the number of collected data. According
to a report presented by Hashimoto et al. [
        <xref ref-type="bibr" rid="ref11">10</xref>
        ], among all
tweets originating in Japan, about 0.18% are geotagged
tweets: this is a rare characteristic for text data. However,
the geotagged tweets we collected are an average of 500
thousand tweets per day. We use about 250 million
geotagged tweets from 2015/2/17 through 2017/5/13. In
addition, although the author has a Twitter account, which
is necessary to use the API, the author never tweets.
Moreover, even for personal account holders, tweets do not
include location information in some areas. Therefore we
do not consider excluding the tweets of the author’s own
account. Using these data, we calculated the best time for
flower viewing, as estimated using the processing described
in the following sections.
      </p>
      <p>Preprocessing
This section presents preprocessing. Preprocessing includes
reverse geocoding and morphological analysis, as well as
database storage for data collected through the processing
described in the previous subsection.</p>
      <p>
        From latitude and longitude information in the individually
collected tweets, reverse geocoding is useful to identify
prefectures and municipalities by town name. We use a
simple reverse geocoding service [
        <xref ref-type="bibr" rid="ref12">11</xref>
        ] available from the
National Agriculture and Food Research Organization in
this process.
      </p>
      <p>
        Morphological analysis divides the collected geotagged
tweet morphemes. We use the “Mecab” morphological
analyzer [
        <xref ref-type="bibr" rid="ref13">12</xref>
        ].
      </p>
      <p>Preprocessing accomplishes necessary data storage for
besttime viewing, as estimated based on results of the
processing of the data collection, reverse geocoding, and
morphological analysis. Data used for this study were the
tweet ID, tweet post time, tweet text, morphological
analysis result, latitude, and longitude.</p>
      <p>
        Interpolation using Kriging
This section presents the method of interpolation, for which
we used Kriging [
        <xref ref-type="bibr" rid="ref14">13</xref>
        ], an estimation method used for
estimating values for points where information was not
acquired. It is impossible to estimate from the number of
geotagged tweets of each sightseeing spot when conducting
detailed analysis at each sightseeing spot. Therefore,
geotagged tweets that have seven significant digits and the
same latitude and longitude information are judged to have
originated from the same spot. As an example, the tweet's
position information of (latitude, longitude) =
(34.93162536621094, 135.72979736328125) is truncated to
(latitude, longitude) = (34.93162, 135.72979). Then, the
tweets from the same point were counted for each date.
Furthermore, by dividing the total for each point by the
total number of tweets on each day, we calculated the
weight of each spot.
      </p>
      <p>We attempted estimation by interpolation using data
aggregated for each spot. The estimated value of the target
data at a point S0 is shown in formula (1) as a weighted
average of the measured values Z(Si) (i = 1, 2..., N) at N
points Si around point S0. Then we assigned value Z to
tweets including the target word and Z. Here, N represents
the 30 nearby targeted tweets. λ denotes a spherical model
with decreased influence as distance increases. As
described in this paper, weighting is done only by the
number of tweets existing in the same spot. However,
consideration of the weight of the tweet itself, such as using
the number of retweets to tweets, is also necessary for
future studies.

  :Unknown weighting of measured value at  -th position
 0 :Predicted position
 :Number of measurements
EXPERIMENTS
In this section, we explain the experiment for information
interpolation for cherry blossoms in 2017, using the method
described in the previous section.</p>
      <p>We are conducting
estimation experiments for cherry blossoms, autumn leaves,
and other phenomena from 2015. As described herein, we
used the period of cherry blossoms in 2017 while studying
the interpolation method to improve estimation accuracy.
The following subsections describe experimental datasets.
Datasets
Datasets used for this experiment were collected using
streaming API, as described for data collection. The data,
which include about 250</p>
      <p>million items, are geotagged
tweets from Japan during 2015/2/17 – 2017/5/13. The
estimation experiment conducted to ascertain the best-time
viewing of cherry blossoms uses the target word “cherry
blossom,” which is “桜”, “さくら”, and “サクラ” in
Japanese. We analyzed tweet texts that include the target
word. About 100,000 tweets during the experiment period
included the subject word.</p>
      <p>The subject of the experiment was set as tourist spots in</p>
    </sec>
    <sec id="sec-2">
      <title>Tokyo. In this report, we describe “Takao</title>
    </sec>
    <sec id="sec-3">
      <title>Mountain,” “Showa</title>
    </sec>
    <sec id="sec-4">
      <title>Memorial</title>
      <p>Park,”
“Shinjuku</p>
      <p>Gyoen,”
and
“Rikugien.” Figure 1 presents the target area locations. A, B,
C, and</p>
      <p>D
in the figure respectively
denote
“Takao</p>
    </sec>
    <sec id="sec-5">
      <title>Mountain,” “Showa</title>
      <p>Memorial Park,” “Rikugien,” and
“Shinjuku Gyoen.” In this experiment, about 30,000 tweets
including the target word in Tokyo were found. In this
experiment, all tweets made by the same user are also used
as targets for analysis if they are tweets including the target
word.</p>
      <p>We conducted experiments of the following two kinds
using these datasets. The first is an experiment using the
number of tweets including the target
word
and the
sightseeing
spot
name
without
interpolation.</p>
    </sec>
    <sec id="sec-6">
      <title>This</title>
      <p>experiment was compared as Baseline to confirm the
usefulness of interpolation proposed in this paper. The
second is an experiment using interpolation. In this
experiment, we used Kriging in the earlier section.
We present the example of Mt. Takao in Hachioji city.
city from 2017/1/1 to 2017/5/13. The point denoted by A is
Mt. Takao. Few geotagged tweets are related to cherry
blossoms near Mt. Takao. For that reason, one cannot
estimate the best time merely using tweets from A.
Therefore, for this study, we interpolated the amount of
information by Kriging using information of Hachioji city
with Mt. Takao. Interpolation is done on a daily basis. The
experiment results are presented in the next subsection.</p>
      <p>A</p>
      <p>B
16 km
32 km</p>
      <p>C
D 6 km</p>
      <p>The light gray part of Figure 3 portrays an experimentally
obtained result from interpolation results including the
tourist spots. Apparently, A was able to produce an estimate
using the proposed method by increasing the number of
tweets using interpolation with surrounding tweets. For B
and C, the useful information was included in the tweet not
co-occurring with the tourist spot name. Therefore, we
confirmed interpolation related to other kinds of cherry
blossoms in early March of B and C and late April of C. In
addition, for D, there are days when it can be determined
more accurately by interpolating the number of tweets.
Therefore, these results confirmed the possibility of
estimating the peak period, even for an area without tweets,
using data interpolation and overall tweet number
interpolation.</p>
      <p>CONCLUSION
As described in this paper, we proposed an interpolation
method to improve the accuracy of tourism information
related to phenological observations. The results of the
cherry blossom experiment conducted at the tourist spots in
Tokyo in 2017 confirmed the trend of improved estimation
accuracy using information interpolation. We confirmed the
possibility of applying this proposed method to the
estimation of viewpoints and sightseeing spots with few
tweets. However, in regions with no geo-tagged tweets,
another method must be considered. Research can be
conducted in the future to verify whether similar results are
obtained for other biological seasonal observations.</p>
      <p>ACKNOWLEDGMENTS
This work was supported by JSPS KAKENHI Grant Nos.
16K00157 and 16K16158, and by a Tokyo Metropolitan
University Grant-in-Aid for Research on Priority Areas
“Research on Social Big Data.”</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Twitter</surname>
          </string-name>
          .
          <article-title>It's what's happening</article-title>
          .
          <source>2017. Retrieved September 2</source>
          ,
          <year>2017</year>
          from https://Twitter.com/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>Ministry of Economy. Trade and Industry</article-title>
          . Inbound
          <string-name>
            <surname>Landing-Type Tourism Guide</surname>
          </string-name>
          .
          <year>2017</year>
          . Retrieved November 10,
          <year>2017</year>
          from
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>http://www.mlit.go.jp/common/001091713.pdf (in Japanese).</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          3.
          <string-name>
            <given-names>S.</given-names>
            <surname>Phithakkitnukoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Teerayut</given-names>
            <surname>Horanont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Witayangkurn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Siri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sekimoto</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Shibasaki</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Understanding tourist behavior using large-scale mobile sensing approach: A case study of mobile phone users in Japan</article-title>
          .
          <source>Pervasive and Mobile Computing</source>
          Volume
          <volume>18</volume>
          :
          <fpage>18</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Niels</given-names>
            <surname>Rosenquist</surname>
          </string-name>
          .
          <source>Understanding the Demographics of Twitter Users</source>
          .
          <year>2011</year>
          .
          <source>In Proceedings of the Fifth International AAAI Conference on Weblogs and Social</source>
          Media (icwsm
          <year>2011</year>
          ),
          <fpage>554</fpage>
          -
          <lpage>557</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          5.
          <string-name>
            <given-names>T.</given-names>
            <surname>Sakaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Okazaki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matsuo</surname>
          </string-name>
          .
          <article-title>Earthquake shakes Twitter users: real-time event detection by social sensors</article-title>
          .
          <source>2010. WWW</source>
          <year>2010</year>
          ,
          <volume>851</volume>
          -
          <fpage>860</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          6.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kaneko</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Yanai</surname>
          </string-name>
          .
          <article-title>Visual Event Mining from the Twitter Stream</article-title>
          .
          <year>2010</year>
          . WWW '16 Companion,
          <source>In Proceedings of the 25th International Conference Companion on World Wide Web</source>
          ,
          <fpage>51</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Yamada</surname>
          </string-name>
          .
          <article-title>Detecting two types of seasonal words using simple autocorrelation analysis</article-title>
          .
          <source>2017. IEEE Big Data 2017 Workshops. In Proceedings of the Second International Workshop on Application of Big Data for Computational Social Science.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          8.
          <string-name>
            <given-names>J.</given-names>
            <surname>Krumm</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Horvitz</surname>
          </string-name>
          .
          <article-title>Eyewitness: identifying local events via space-time signals in Twitter feeds</article-title>
          .
          <year>2015</year>
          .
          <source>In Proceedings of the 23rd SIGSPATIAL International Conference on Advances in Geographic Information Systems (SIGSPATIAL '15)</source>
          . ACM, New York, NY, USA,, Article
          <volume>20</volume>
          , 10 pages. DOI: https://doi.org/10.1145/2820783.2820801
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Twitter</given-names>
            <surname>Developers</surname>
          </string-name>
          .
          <source>Twitter Developer official site</source>
          .
          <source>2017. Retrieved April 2</source>
          ,
          <year>2017</year>
          from https://dev.twitter.com/
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          10.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hashimoto</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Oka</surname>
          </string-name>
          .
          <article-title>Statistics of Geo-Tagged Tweets in Urban Areas (&lt;Special Issue&gt;Synthesis and Analysis of Massive Data Flow)</article-title>
          .
          <year>2012</year>
          . JSAI vol.
          <volume>27</volume>
          ,
          <issue>4</issue>
          :
          <fpage>424</fpage>
          -
          <lpage>431</lpage>
          (in Japanese).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          11.
          <string-name>
            <given-names>National</given-names>
            <surname>Agriculture</surname>
          </string-name>
          and Food Research Organization.
          <article-title>Simple reverse geocoding service</article-title>
          .
          <source>2017. Retrieved November 18</source>
          ,
          <year>2017</year>
          from https://www.finds.jp/rgeocode/index.html.ja
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          12.
          <string-name>
            <surname>MeCab. Yet Another</surname>
            Part-of-Speech and
            <given-names>Morphological</given-names>
          </string-name>
          <string-name>
            <surname>Analyzer</surname>
          </string-name>
          .
          <year>2017</year>
          . Retrieved November 10,
          <year>2017</year>
          from http://taku910.github.io/mecab/
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          13.
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Oliver</surname>
          </string-name>
          .
          <article-title>Kriging: A Method of Interpolation for Geographical Information Systems</article-title>
          .
          <year>1990</year>
          .
          <source>International Journal of Geographic Information Systems</source>
          vol.
          <volume>4</volume>
          :
          <fpage>313</fpage>
          -
          <lpage>332</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>