<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Accessibility by public transport predicts residential real estate prices: a case study in Helsinki region</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Indre˙ Zˇ liobaite˙</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aalto University, Dept. of Computer Science</institution>
          ,
          <addr-line>Espoo</addr-line>
          ,
          <country country="FI">FINLAND</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Helsinki Institute for Information Technology (HIIT)</institution>
          ,
          <addr-line>Espoo</addr-line>
          ,
          <country country="FI">FINLAND</country>
        </aff>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>This pilot study investigates how considering
accessibility could help to model prices of
residential real estate more accurately.</p>
      <p>We introduce
two novelties from the price modeling point of
view (1) defining accessibility as travel time by
public transport, in addition to geographic
distance, and (2) considering dynamic points of
interest from check-ins into social networks, in
addition to fixed location community centers. Our
case study focuses on the Helsinki region.
We
model price per square meter as a linear function
of apartment characteristics, and characteristics
of the neighborhood, including accessibility by
public transport and social activities. The
resulting models show good predictive performance,
as compared to baselines not taking accessibility
into account. We discover that apartment price
relates to the geographical distance from the city
center, but accessibility by public transport to
local centers of interest is more informative than
just the geographical distance to those centers.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Modeling real estate prices has long been of interest to
researchers and practitioners, and it is employed for various
purposes related to investment, lending or taxation.
Arguably all city residents, even non-specialists, intuitively
understand that the price of a residential apartment
positively relates to the size of the apartment, and negatively
reProceedings of the 2 nd International Workshop on Mining Urban
Data, Lille, France, 2015. Copyright c 2015 for this paper by its
authors. Copying permitted for private and academic purposes.
lates to the distance to the city center. Professional real
estate price models include many more features of apartments
and environment, such as age, construction type, floor, or
population characteristics in the neighborhood.</p>
      <p>
        Residential real estate prices are typically modeled using
so called hedonic models
        <xref ref-type="bibr" rid="ref2 ref5">(Case &amp; Quigley, 1991; Sirmans
et al., 2005)</xref>
        , where the price of a house is assumed to be
affected by the structural characteristics of the house
itself, characteristics of the neighborhood, and
environmental characteristics.
      </p>
      <sec id="sec-2-1">
        <title>While in real estate domain research</title>
        <p>mainly focuses on identifying factors that impact pricing,
in machine learning and data mining research real estate
price modeling mainly focuses on developing sophisticated
predictive models beyond linear regression (Chopra et al.,
2007; Fu et al., 2014).</p>
        <p>A
literature review
on
hedonic
pricing
models
(Bartholomew
&amp;</p>
        <p>Ewing, 2011)
finds the structural
characteristics typically include the age and the size of
the house, the number of bedrooms, and the presence of
different amenities such as a garage.</p>
      </sec>
      <sec id="sec-2-2">
        <title>The effect of the</title>
        <p>location of the house on housing prices is often captured by
physical proximity to a central business district (CBD) or a
regional center. The literature review finds evidence of an
inverse relationship between pricing and distance to CBD
in studies on various cities around the world.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Another</title>
        <p>access-related characteristic often used in hedonic models
is the proximity of the house to a transit station, measured
in air distance or walking distance. This attribute is used to
capture the effect transit has on relative accessibility of a
CBD or a regional center. Here the results are more mixed,
with the majority of studies suggesting pricing premiums
for housing located near to a transit station, and a higher
premium for transit stations that provide a higher degree of
relative proximity to a CBD.
The era of big data provides access to new data sources,
such as public transport, traffic and social mobility data,
that potentially relate to real estate prices (at least
intuitively we know that people consider mobility, and social
factors when buying an apartment). Integrating such data
could help to model residential real estate prices more
precisely, and, as a result, better understand urban mobility
patterns and activities. Such models can contribute to
managing, coordinating and long term planning of mobility, and
overall development of modern smart cities.</p>
        <p>Our pilot study investigates to what extent accessibility of a
neighborhood relates to residential real estate prices. This
case study focuses on the Helsinki region. We model price
per square meter as a linear function of apartment
characteristics, and characteristics of the neighborhood, including
accessibility by public transport and social activities. Our
main hypothesis is that prices are more related to travel
times than travel distances, and local centers of activities
than the city center. The resulting models show good
predictive performance, as compared to baselines not taking
accessibility into account. We discover that an apartment
price relates to the geographical distance from the city
center, but accessibility by public transport to local centers of
interest is more informative than just the geographical
distance to those centers.</p>
        <p>Our study introduces two conceptual novelties in modeling
prices of residential real estate: (1) to measure
accessibility, we consider travel times in addition to distances, and
(2) we consider dynamic local points of interest, defined by
4square1 check-ins (people posting their location and
activity on a social network), in addition to community centers
at fixed locations.</p>
        <p>The remainder of the paper is organized as follows. Section
2 describes data acquisition and feature engineering.
Section 3 presents the results of the experimental case study,
and Section 4 concludes the study.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Data acquisition and feature engineering</title>
      <p>Our dataset consists of three parts: real estate data
describing characteristics of the apartments, location data
describing points of interest and community centers, and
accessibility data describing point-to-point distances and travel
times. We make our dataset publicly available2 for
research.</p>
      <p>1http://foursquare.com
2http://www.zliobaite.com/datahel.zip
●
● ●● ●●●●●●●●●
●● ●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●● ●
●●●●●●●●●●●●●●●●●●●●● ● ●
● ●● ●</p>
      <sec id="sec-3-1">
        <title>2.1. Real estate data</title>
        <p>The sales price data comes from a Finnish web portal
Oikotie3, which is the most popular marketplace for
residential real estates sales and rental. Our dataset consists
of apartments in the capital region (Helsinki, Espoo,
Vantaa and Kauniainen municipalities) advertised for sales on
October 24, 2014. The pricing data is based on sales ads,
as sales transaction prices are not available for the public.
We exclude apartments that do not provide a street address
(hence no coordinates), and for which size is not available.
Moreover, we filter out very large apartments (size more
than 300 m2), very old apartments (built earlier than 1850),
far away apartments (distance to metro more than 20 km),
extremely cheap (price pr square meter less than 1200 eur)
and extremely expensive apartments (price per square
meter more than 12000 eur), because we aim at focusing on
modeling prices of mainstream apartments and avoiding
extreme outliers. After filtering our dataset includes 8337
apartments. Figure 1 plots all the apartment locations.</p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Location data</title>
        <p>We consider two types of location data: fixed location,
and dynamic points of interest. Fixed location data
includes the city center, for which the Stockmann
department store is used as a proxy (coordinates found by hand
via Google maps), and local community centers,
approximated by H&amp;M shop (a chain of clothing shops)
locations in Helsinki region (also found by hand from Google
Stockman</p>
        <p>HM
●
center (Stockmann). The black rectangular indicates the area from
6:00, violet 6:00-10:00, red 10:00-14:00, brown 14:00-18:00,
orwhich point of interest data is collected.
ange 18:00-22:00.
maps). Stockmann is a well-known location in the centre of
Helsinki. H&amp;M shops are typically present in larger
shopping malls. Shopping malls are local centers of attraction.
We hope that H&amp;M serves as a proxy for local centers in
the neighborhoods. Figure 2 plots the community centers
and the city center location.</p>
        <p>Dynamic points of interest are obtained from an existing
dataset of 4square check-ins (Le Falher et al., 2015). Each
check-in in the dataset corresponds to one user’s visit to
one venue (restaurant, cafeteria, store, etc) with known
geographic location, at a particular time. The data cover user
activity between March and July 2014 in the inner Helsinki
city. To extract points of interest, we perform k-means
clustering on the geographic locations of check-ins, using
k = 20. Each of the k centroids identified defines one point
of interest. Note that we extract points of interest both on
top of all check-ins contained in the dataset, regardless of
the the time of the day they occur, as well as separately for
check-ins that occur at separate time intervals in the day
(five 4-hour intervals from 2am to 10pm ). Figure 3 plots
the points of interest for each time interval.</p>
      </sec>
      <sec id="sec-3-3">
        <title>2.3. Accessibility features</title>
        <p>Accessibility data connects apartments with point of
interest.</p>
        <p>We consider two types of accessibility features: air
distance from an apartment to the location of a point of
interest, and travel time by public transport from an
apartment to the point of interest (including walking time).
Air distance is measured in kilometers from coordinate of
the apartment to coordinate of the point of interest, as
D
s1
s2
=
=
=</p>
        <p>Re · arccos(s1 + s2), where
cos(lat 1) ∗ cos(lat 2) ∗ cos(lon2 − lon1),
sin(lat 1) ∗ sin(lat 2),
where Re is the radius of Earth (set to Re = 6371km),
(lat 1, lon1) are the coordinates of the apartment, and
(lat 2, lon2) are the coordinates of the point of interest.
Travel time by public transport between two coordinates is
measured using a freely available tool Reititin4, developed
by BusFaster Ltd and researchers at University of Helsinki.</p>
        <sec id="sec-3-3-1">
          <title>We use the default settings.</title>
          <p>In addition to accessibility between apartments and points
of interest we also include the distance from an apartment
to the nearest metro station. The address of Metro stations
is listed on Helsinki Metro’s website5 and their geographic
coordinates are collected via manual queries to the Google
Maps API6. Note that in the Helsinki region metro runs
only to the eastern part of the city, therefore, we do not
necessarily expect a regular behavior from this feature. A
regular behavior would be a higher price if there is a metro
stop nearby.</p>
          <p>4http://blogs.helsinki.fi/saavutettavuus/
tyokaluja/metropaccess-reititin/
5http://www.hel.fi/hki/hkl/en/HKL+Metro
6https://developers.google.com/maps/
2
m
_
r
e
p
_
ce
ir
We see that the location features are strongly correlated
with each other, therefore, many may be redundant.
Nevertheless, some of those could potentially be expected to be
more informative than others, therefore, we consider them
all. We can also see than most of the location features are
negatively correlated with community centers and dynamic
points of interest. We already have seen similar tendencies
in the scatterplots. This behavior is along with a common
intuition that apartments near points of interest should be
more expensive.</p>
          <p>The correlation and scatter plots analyzed features
one-byone. In the next section we will consider predictive models
that use sets of features for modeling apartment prices.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Case study</title>
      <p>The goal of this pilot case study is to investigate whether
accessibility information helps to model real estate prices,
as compared to using only geographical location
information. In addition, we investigate informativeness of
dynamic points of interest (derived from social networks) as
opposed to stationary fixed points of interest.</p>
      <sec id="sec-4-1">
        <title>3.1. Experimental protocol</title>
        <p>We model price per square meter. An alternative would
be to model the total price. We choose the former as the
target variable, because price per square meter is easier to
interpret and compare across neighborhoods.</p>
        <p>We limit our analysis to linear regression, which is easily
Ne
interpretable. Note, however, that some of the features are
expressed as non-linear functions of simpler features (e.g.
fyear is a non-linear function of a building’s age, as
explained above). The ordinary least squares procedure (the
standard implementation in R) is used for estimating the
model parameters.</p>
        <p>For assessing the performance we use two common
accuracy measures: coefficient of determination (R2) and mean
absolute error (MAE). Coefficient of determination is a
relative accuracy measure, where 1 means the best possible
performance, and 0 means the performance is equivalent to
random. Mean absolute error indicates error in the units of
the target variable, 0 is an ideal performance, the higher the</p>
        <sec id="sec-4-1-1">
          <title>MAE, the worse the performance.</title>
          <p>We report R2 and MAE measured on the whole dataset
used for model fitting (fit) and via 10 fold cross-validation
(cv), which iteratively fits a model on 90% of the data, and
tests on the remaining part. Cross-validation scores provide
an indication of how models would generalize to unseen
data.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Performance of base models</title>
        <p>Base models do not use any accessibility information, and
use only very basic location information. The first base
model (Size-year) does not use location at all, and is based
only on size of the apartment and its construction year
(fyear ). The second model in addition uses basic
location information, encoded as raw geographical coordinates,
centered in the old town of the city.</p>
        <p>The resulting models for price per square meter are:
price
=</p>
        <p>3722 − 4.97 × size + 0.91 × fyear ,
and
=
+
price
5643 − 5.14 × size + 0.78 × fyear +
38.9 × east − 147.7 × north.</p>
        <p>The models are consistent with common intuition: the
larger the apartment, the cheaper the price per square
meter; older or newer apartments with respect to 1970
construction year are more expensive; the further to the north
from the sea and the city center, the cheaper. Easterness
has a positive effect, which is somewhat inconsistent with
a common intuition that cheaper neighborhoods are in the
east. However, this can be explained by the range of data
(see Figure 4). Data extends further to the west than to the
east, therefore, western apartments are on average further
from the center, and thus cheaper.</p>
        <p>Table 2 reports predictive accuracies of the base models.
We can make two observations. First, Size-year-location
model already performs quite well with the cross-validation
of determination (the higher, the better), MAE - mean absolute
error (the smaller the better).
location model with one additional accessibility feature at a time.
formance differs only a little, which suggests that there is
no notable overfitting, and the model could use more
informative input features.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Predictive power of accessibility</title>
        <p>Next we test whether adding accessibility information
helps to predict more accurately. We test accessibility to
the city center (Stockmann), local centers (HM), and
dynamic centers of interest (4square check-ins) overall and at
morning peak times (from 6:00 to 10:00). We compare
informativeness of using air distance as a feature to using the
total travel time by public transport.</p>
        <p>We use the base model Size-year-location as a starting
point, add one feature at a time to it, and measure the
accuracy. Table 3 reports the results.</p>
        <p>From the resulting accuracies we can see that
accessibility has some predictive power, as in all cases the predictive
performance improves as compared to the base model. The
results indicate that the distance to the city center
(Stockmann) is more informative than the travel time by
public transport. However, accessibility to the local centers
(fixed centers H&amp;M and dynamic centers 4square) by
public transport is more informative than just the air distance
to those centers. In other words, it seems that an apartment
price relates to the overall geographical location, but
accessibility to local centers of interest is more important than
just the geographical distance to those centers. This is an
interesting finding for exploring in detail in future studies.
We report selected models. Metro distance is intuitive - the
=
−
=
−
=
+
=
+
ment:
price
apartment:
price
closer to metro, the more expensive is the apartment:
price
5300 − 3.94 × size + 0.77 × fyear −
62.2 × east − 50.8 × north − 158.4 × metro.
Adding metro distance shrinks other coefficients, which
suggests that earlier this feature was indirectly captured.
More importantly, adding metro distance changes the
direction of the easternness coefficient from positive to
negative. Now it is more intuitive keeping in mind
peculiarities of Helsinki residential neighborhoods, where overall
the east is considered cheaper than the west.</p>
        <p>Stockmann distance is as well intuitive - the closer to the
center, the more expensive is the apartment:
price
5698 − 3.78 × size + 0.72 × fyear +
3.3 × east − 59.8 × north − 117.8 × dstock .</p>
        <p>Shortest H&amp;M travel time is intuitive - the shorter the travel
time to the local center, the more expensive is the
apart5681 − 5.00 × size + 0.78 × fyear +
37.0 × east − 139.6 × north − 39.3 × thm .</p>
        <p>Shortest 4square travel time is intuitive - the shorter the
travel time to a center of interest, the more expensive is the
5659 − 3.55 × size + 0.77 × fyear +
13.5 × east − 103.3 × north − 31.4 × t4sq .
price
5729 − 4.06 × size + 0.71 × fyear +
31.6 × east − 94.5 × north +
73.0 × metro − 139.0 × dstock +
21.6 × thm − 12.5 × t4sq2
=
+
+
+
We can see some interesting relations, reflecting
peculiarities of the Helsinki region. First, the longer the metro
distance, the higher the price, while one could expect the
opposite. Our interpretation is that the metro distance
captures what was not very successfully captured by the
easternness.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion</title>
      <p>We have experimentally explored several models for real
estate prices in Helsinki region, focusing our analysis on
accessibility by public transport and dynamic points of
interest, obtained via check-ins into social networks.
We
have found that even a basic account for accessibility
features helps to improve the accuracy of price estimates. We
have discovered that an apartment price relates to the
geographical distance from the city center, but accessibility by
public transport to local centers of interest is more
informative than just the geographical distance to those centers.
Integrating such data could help to model residential real
estate prices more precisely, and, as a result, better
understand urban mobility patterns and activities. Such
models can contribute to managing, coordinating and long term
planning of mobility, and overall development of modern
smart cities.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The authors thank Antti Ukkonen for insightful
discussions. Research leading to these results was partially
sup2014.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>Journal of Planning Literature</source>
          ,
          <volume>26</volume>
          (
          <issue>1</issue>
          ):
          <fpage>18</fpage>
          -
          <lpage>34</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Case</surname>
          </string-name>
          ,
          <source>Bradford and Quigley</source>
          ,
          <string-name>
            <surname>John M.</surname>
          </string-name>
          <article-title>The dynamics of real estate prices</article-title>
          .
          <source>The Review of Economics and Statistics</source>
          ,
          <volume>73</volume>
          (
          <issue>1</issue>
          ):
          <fpage>50</fpage>
          -
          <lpage>58</lpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>Discovering the hidden structure of house prices with a non-parametric latent</article-title>
          <source>In Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '07</source>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Fu</surname>
          </string-name>
          , Yanjie, Xiong, Hui, Ge, Yong, Yao, Zijun, Zheng, Yu, and
          <string-name>
            <surname>Zhou</surname>
          </string-name>
          , Zhi-Hua.
          <article-title>Exploiting geographic dependencies for real estate appraisal: A mutual perspective of ranking and clustering</article-title>
          .
          <source>In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '14</source>
          , pp.
          <fpage>1047</fpage>
          -
          <lpage>1056</lpage>
          ,
          <string-name>
            <surname>Le</surname>
            <given-names>Falher</given-names>
          </string-name>
          , Geraud, Gionis, Aris, and
          <string-name>
            <surname>Mathioudakis</surname>
          </string-name>
          , Michael.
          <article-title>Where is the Soho of Rome? : Measures and algorithms for finding similar neighborhoods in cities</article-title>
          .
          <source>In The 9th International AAAI Conference on Web and Social Media</source>
          ,
          <string-name>
            <surname>ICWSM</surname>
          </string-name>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Sirmans</surname>
            , Stacy, Macpherson, David, and Zietz,
            <given-names>Emily.</given-names>
          </string-name>
          <article-title>The composition of hedonic pricing models</article-title>
          .
          <source>Journal of Real Estate Literature</source>
          ,
          <volume>13</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>44</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>