<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring chance in NCAA basketball</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>INSA Lyon</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <abstract>
        <p>There seems to be an upper limit to predicting the outcome of matches in (semi-)professional sports. Recent work has proposed that this is due to chance and attempts have been made to simulate the distribution of win percentages to identify the most likely proportion of matches decided by chance. We argue that the approach that has been chosen so far makes some simplifying assumptions that cause its result to be of limited practical value. Instead, we propose to use clustering of statistical team pro les and observed scheduling information to derive limits on the predictive accuracy for particular seasons, which can be used to assess the performance of predictive models on those seasons. We show that the resulting simulated distributions are much closer to the observed distributions and give higher assessments of chance and tighter limits on predictive accuracy.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In our last work on the topic of NCAA basketball [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], we speculated about
the existence of a \glass ceiling" in (semi-)professional sports match outcome
prediction, noting that season-long accuracies in the mid-seventies seemed to
be the best that could be achieved for college basketball, with similar results
for other sports. One possible explanation for this phenomenon is that we are
lacking the attributes to properly describe sports teams, having difficulties to
capture player experience or synergies, for instance. While we still intend to
explore this direction in future work,1 we consider a different question in this
paper: the in uence of chance on match outcomes.
      </p>
      <p>Even if we were able to accurately describe sports teams in terms of their
performance statistics, the fact remains that athletes are humans, who might
make mistakes and/or have a particularly good/bad day, that matches are
refereed by humans, see before, that injuries might happen during the match, that
the interaction of balls with obstacles off which they ricochet quickly becomes
too complex to even model etc. Each of these can affect the match outcome
to varying degrees and especially if we have only static information from
before the match available, it will be impossible to take them into account during
prediction.
1 Others in the sports analytics community are hard at work doing just that, especially
for \under-described sports such as European soccer or NFL football.</p>
      <p>While this may be annoying from the perspective of a researcher in sports
analytics, from the perspective of sports leagues and betting operators, this is a
feature, not a bug. Matches of which the outcome is effectively known beforehand
do not create a lot of excitement among fans, nor will they motivate bettors to
take risks.</p>
      <p>Intuitively, we would expect that chance has a stronger effect on the outcome
of a match if the two opponents are roughly of the same quality, and if scoring
is relatively rare: since a single goal can decide a soccer match, one (un)lucky
bounce is all it needs for a weaker team to beat a stronger one. In a fast-paced
basketball game, in which the total number of points can number in the two
hundreds, a single basket might be the deciding event between two evenly matched
teams but probably not if the skill difference is large.</p>
      <p>
        For match outcome predictions, a potential question is then: \How strong is
the impact of chance for a particular league? ", in particular since quantifying
the impact of chance also allows to identify the \glass ceiling" for predictions.
The topic has been explored for the NFL in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which reports
      </p>
      <p>The actual observed distribution of win-loss records in the NFL is
indistinguishable from a league in which 52.5% of the games are decided at
random and not by the comparative strength of each opponent.
Using the same methodology, Weissbock et al. [6] derive that 76% of matches
in the NHL are decided by chance. As we will argue in the following section,
however, the approach used in those works is not applicable to NCAA basketball.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Identifying the impact of chance by Monte Carlo simulations</title>
      <p>The general idea used by Burke and Weissbock2 is the following:
1. A chance value c 2 [0; 1] is chosen.
2. Each out of a set of virtual teams is randomly assigned a strength rating.
3. For each match-up, a value v 2 [0; 1] is randomly drawn from a uniform
distribution.</p>
      <p>{ If v c, the stronger team wins.</p>
      <p>{ Otherwise, the winner is decided by throwing an unweighted coin.
4. The simulation is re-iterated a large number of times (e.g. 10; 000) to smooth
results.</p>
      <p>Figure 1 shows the distribution of win percentages for 340 teams, 40 matches
per team (roughly the settings of an NCAA basketball season including playoffs),
and 10; 000 iterations for c = 0:0 (pure skill), c = 1:0 (pure chance), and c = 0:5.</p>
      <p>
        By using a goodness of t test { 2 in the case of Burke's work, F-Test in the
case of Weissbock's { the c-value is identi ed for which the simulated distribution
ts the empirically observed one best, leading to the values reproduced in the
2 For details for Weissbock's work, we direct the reader to [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ].
      </p>
      <p>0.14
0.12
0.1
0.04
0.02
itrno 0.08
o
p
roP 0.06</p>
      <p>Pure skill
Pure chance</p>
      <p>Chance=0.5
0 0
0.2
0.4</p>
      <p>0.6
introduction. The identi ed c-value can then be used to calculate the upper limit
on predictive accuracy in the sport: since in 1 c cases the stronger team wins,
and a predictor that predicts the stronger team to win can be expected to be
correct in half the remaining cases in the long run, the upper limit lies at:
(1</p>
      <p>c) + c=2,
leading in the case of
{ the NFL to: 0:475 + 0:2625 = 0:7375, and
{ the NHL to: 0:24 + 0:36 = 0:62
Any predictive accuracy that lies above those limits is due to the statistical
quirks of the observed season: theoretically it is possible that chance always
favors the stronger team, in which case predictive accuracy would actually be
1:0. As we will argue in the following section, however, NCAA seasons (and not
only they) are likely to be quirky indeed.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Limitations of the MC simulation for NCAA basketball</title>
      <p>A remarkable feature of Figure 1 is the symmetry and smoothness of the
resulting curves. This is an artifact of the distribution assumed to model the
theoretical distribution of win percentages { the Binomial distribution { together with
the large number of iterations. This can be best illustrated in the \pure skill"
setting: even if the stronger team were always guaranteed to win a match,
realworld sports schedules do not guarantee that any team actually plays against
representative mix of teams both weaker and stronger than itself. A reasonably
strong team could still lose every single match, and a weak one could win at a
reasonable clip. One league where this is almost unavoidable is the NFL, which
consists of 32 teams, each of which plays 16 regular season matches (plus at
most 4 post-season matches), and ranking \easiest" and \hardest" schedules in
the NFL is an every-season exercise. Burke himself worked with an empirical
distribution that showed two peaks, one before 0:5 win percentage, one after. He
argued that this is due to the small sample size ( ve seasons).</p>
      <p>0.045
0.04
0.035
0.03
itrno 0.025
o
p
roP 0.02
0.015
0.01
0.005
0 0
2008, observed
2009, observed
2010, observed
2011, observed
2012, observed
2013, observed
0.4</p>
      <p>0.5</p>
      <p>Win percentage
0.1
0.2
0.3
0.6
0.7
0.8
0.9
1
The situation is even more pronounced in NCAA basketball, where 340+
Division I teams play at most 40 matches each. Figure 2 shows the empirical
distribution for win percentages in NCAA basketball for six season (2008{2013).3
While there is a pronounced peak for a win percentage of 0:5 for 2008 and 2012,
the situation is different for 2009, 2010, 2011, and 2013. Even for the former
two seasons, the rest of the distribution does not have the shape of a Binomial
distribution. Instead it seems to be that of a mix of distributions { e.g. \pure
skill" for match-ups with large strength disparities overlaid over \pure chance"
for approximately evenly matched teams.</p>
      <p>NCAA scheduling is subject to conference memberships and teams will try to
pad out their schedules with relatively easy wins, violating the implicit
assumptions made for the sake of MC simulations. This also means that the \statistical
quirks" mentioned above are often the norm for any given season, not the
exception. Thought to its logical conclusion, the results that can be derived from
the Monte Carlo simulation described above are purely theoretical: if one could
observe an effectively unlimited number of seasons, during which schedules
are not systematically imbalanced, the overall attainable predictive
accuracy were bound by the limit than can be derived by the simulation. For a given
3 The choice of seasons is purely due to availability of data at the time of writing and
we intend to extend our analysis in the future.
season, however, and the question how well a learned model performed w.r.t. the
speci cities of that season, this limit might be too high (or too low).
2008, observed, KS=0.0
2008, simulated, cluster bias, KS=0.0526
MC simulation, chance=0.41, KS=0.0748</p>
      <p>MC simulation, chance=0.525, KS=0.0694
0.06
0.05
0.04
n
o
itrop 0.03
o
r
P
0.02
0.01
0 0
0.2
0.4</p>
      <p>0.6
Win Percentage
0.8
1</p>
      <p>As an illustration, consider Figure 3.4 The MC simulation that matches the
observed proportion of teams having a win percentage of 0:5 is derived by
setting c = 0:42, implying that a predictive accuracy of 0:79 should be possible.
The MC simulation that ts the observed distribution best, according to the
Kolmogorov-Smirnov (KS) test (overestimates the proportion of teams having a
win percentage of 0:5 along the way), is derived from c = 0:525 (same as Burke's
NFL analysis), setting the predictive limit to 0:7375. Both curves have visually
nothing in common with the observed distribution, yet the null hypothesis { that
both samples derive from the same distribution { is not rejected at the 0.001
level by the KS test for sample comparison. This hints at the weakness of using
such tests to establish similarity: CDFs and standard deviations might simply
not provide enough information to decide whether a distribution is appropriate.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Deriving limits for speci c seasons</title>
      <p>The ideal case derived from the MC simulation does not help us very much in
assessing how close a predictive model comes to the best possible prediction.
Instead of trying to answer the theoretical question: What is the expected limit
to predictive accuracy for a given league?,
we therefore want to answer the practical question: Given a speci c season, what
was the highest possible predictive accuracy?.
4 Other seasons show similar behavior, so we treat 2008 as a representative example.</p>
      <p>To this end, we still need to nd a way of estimating the impact of chance
on match outcomes, while taking the speci cities of scheduling into account. The
problem with estimating the impact of chance stays the same, however: for any
given match, we need to know the relative strength of the two teams but if we
knew that, we would have no need to learn a predictive model in the rst place.
If one team has a lower adjusted offense efficiency than the other (i.e. scoring
less), for example, but also a lower adjusted defensive efficiency (i.e. giving up
fewer points), should it be considered weaker, stronger, or of the same strength?</p>
      <p>Learning a model for relative strength and using it to assess chance would
therefore feed the models potential errors back into that estimate. What we can
attempt to identify, however, is which teams are similar.
4.1</p>
      <p>Clustering team pro les and deriving match-up settings</p>
      <p>
        We describe each team in terms of their adjusted efficiencies, and their Four
Factors, adopting Ken Pomeroy's representation [
        <xref ref-type="bibr" rid="ref5">4</xref>
        ]. Each statistic is present
both in its offensive form { how well the team performed, and in its defensive
form { how well it allowed its opponents to perform (Table 1). We use the
averaged end-of-season statistics, leaving us with approximately 340 data points
per season. Clustering daily team pro les, to identify ner-grained relationships,
and teams' development over the course of the season, is left as future work.
As a clustering algorithm, we used the WEKA [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] implementation of the EM
algorithm with default parameters. This involves EM selecting the appropriate
number of clusters by internal cross validation, with the second row of Table 2
showing how many clusters have been found per season.
      </p>
      <p>Season 2008 2009 2010 2011 2012 2013</p>
      <p>Number of Clusters 5 4 6 7 4 3</p>
      <p>Cluster IDs in Tournament 1,5 4 2,6 1,2,5 3,4 2
Table 2. Number of clusters per season and clusters represented in the NCAA
tournament</p>
      <p>As can be seen, depending on the season, 340 teams do not separate into
many different statistical pro les. Additionally, as the third row shows, only
certain clusters, representing relatively strong teams, make it into the NCAA
tournament, with the chance to eventually play for the national championship
(and one cluster dominates, like Cluster 5 in 2008). These are strong indications
that the clustering algorithm does indeed discover similarities among teams that
allow us to abstract \relative strength". Using the clustering results, we can
reencode a season's matches in terms of the clusters to which the playing teams
belong, capturing the speci cities of the season's schedule.</p>
      <p>Table 3 summarizes the re-encoded schedule for 2008. The re-encoding
allows us to esh out the intuition mentioned in the introduction some more:
teams from the same cluster can be expected to have approximately the same
strength, increasing the impact of chance on the outcome. Since we want to take
all non-chance effects into account, we encode pairings in terms of which teams
has home-court. The left margin indicates which team has home court in the
pairing: this means, for instance, that while teams from Cluster 1 beat teams
from Cluster 2 almost 80% of the time when they have home court advantage,
teams from Cluster 2 prevail in almost 57% of the time if home court advantage
is theirs. The effect of home court advantage is particularly pronounced on the
diagonal, where unconditional winning percentages by de nition should be at
approximately 50%. Instead, home court advantage pushes them always above
60%. One can also see that the majority of cases teams were matched up with a
team stronger than (or as strong as) themselves. Table 3 is the empirical
instantiation of our remark in Section 3: instead of a single distribution, 2008 seems
to have been a weighted mixture of 25 distributions.5 None of these speci cities
can be captured by the unbiased MC simulation.
4.2</p>
      <p>Estimating chance
The re-encoded schedule includes all the information we need to assess the effects
of chance. The win percentage for a particular cluster pairing indicates which of
the two clusters should be considered the stronger one in those circumstances,
and from those matches that are lost by the stronger team, we can calculate the
chance involved.</p>
      <p>Consider, for instance, the pairing Cluster 5 { Cluster 2. When playing at
home, teams from Cluster 5 win this match-up in 82.85% of the cases! This is
5 Although some might be similar enough to be merged.
the practical limit to predictive accuracy in this setting for a model that always
predicts the stronger team to win, and in the same way we used c to calculate
that limit above, we can now inverse the process: c = 2 (1 0:8285) = 0:343.
When teams from Cluster 5 welcomed teams from Cluster 2 on their home court
in 2008, the overall outcome is indistinguishable from 34.3% of matches having
been decided by chance.</p>
      <p>The impact of chance for each cluster pairing, and the number of matches
that have been played in particular settings, nally, allows us to calculate the
effect of chance on the entire season, and using this result, the upper limit for
predictive accuracy that could have been reached for a particular season.
1. Teams were described in terms of adjusted efficiencies and Four Factors {
adding or removing statistics could lead to different numbers of clusters and
different cluster memberships.
2. Predictive models that use additional information, e.g. experience of players,
or networks models for drawing comparisons between teams that did not play
each other, can exceed the limits reported in Table 4.</p>
      <p>The table also indicates that it might be less than ideal to learn from preceding
seasons to predict the current one (the approach we have chosen in our previous
work): having a larger element of chance (e.g. 2008) could bias the learner against
relatively stronger teams and lead it to underestimate a team's chances in a more
regular season (e.g. 2010).
5</p>
    </sec>
    <sec id="sec-5">
      <title>Simulating seasons</title>
      <p>With the scheduling information and the impact of chance for different pairings,
we can simulate seasons in a similar manner to the Monte Carlo simulations we
have discussed above, but with results that are much closer to the distribution
of observed seasons. Figure 3 shows that while the simulated distribution is not
equivalent to the observed one, it shows very similar trends. In addition, while
the KS test does not reject any of the three simulated distributions, the distance
of the one resulting from our approach to the observed one is lower than for the
two Monte Carlo simulated ones.</p>
      <p>The gure shows the result of simulating the season 10; 000 times, leading to
the stabilization of the distribution. For fewer iterations, e.g. 100 or less,
distributions that diverge more from the observed season can be created. In particular,
this allows the exploration of counterfactuals: if certain outcomes were due to
chance, how would the model change if they came out differently? Finally, the
information encoded in the different clusters { means of statistics and co-variance
matrices { allows the generation of synthetic team instances that t the cluster
(similar to value imputation), which in combination with scheduling information
could be used to generate wholly synthetic seasons to augment the training data
used for learning predictive models. We plan to explore this direction in future
work.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Summary and conclusions</title>
      <p>In this paper, we have considered the question of the impact of chance on the
outcome of (semi-)professional sports matches in more detail. In particular, we
have shown that the unbiased MC simulations used to assess chance in the
NFL and NHL are not applicable to the college basketball setting. We have
argued that the resulting limits on predictive accuracy rest on simplifying and
idealized assumptions and therefore do not help in assessing the performance of
a predictive model on a particular season.</p>
      <p>As an alternative, we propose clustering teams' statistical pro les and
reencoding a season's schedule in terms of which clusters play against each other.
Using this approach, we have shown that college basketball seasons violate the
assumptions of the unbiased MC simulation, given higher estimates for chance,
as well as tighter limits for predictive accuracy.</p>
      <p>There are several directions that we intend to pursue in the future. First,
as we have argued above, NCAA basketball is not the only setting in which
imbalanced schedules occur. We would expect similar effects in the NFL, and
even in the NBA, where conference membership has an effect. What is needed
to explore this question is a good statistical representation of teams, something
that is easier to achieve for basketball than football/soccer teams.</p>
      <p>
        In addition, as we have mentioned in the preceding section, the exploration of
counterfactuals and generation of synthetic data should help in analyzing sports
better. We nd a recent paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] particularly inspirational, in that the authors
used a detailed simulation of substitution and activity patterns to explore
alternative outcomes for an NBA playoff series.
6. Weissbock, J., Inkpen, D.: Combining textual pre-game reports and statistical data
for predicting success in the national hockey league. In: 27th Canadian AI 2014,
Montreal, QC, Canada, May 6-9, 2014. pp. 251{262 (2014)
using machine learning techniques: some results and lessons learned (originally
in "MLSA13", workshop at ECML/PKDD 2013). arXiv preprint arXiv:1310.3607
(2013)
A
      </p>
    </sec>
    <sec id="sec-7">
      <title>Clustered schedules for different seasons</title>
      <p>Cluster 1 Cluster 2 Cluster 3 Cluster 4
Cluster 1 133/197
Cluster 2 210/227
Cluster 3 261/308
Cluster 4 210/211
46/182
231/352
192/357
341/374
105/272
262/374
409/663
424/448
1/45
76/247
56/261
515/818</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Burke</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Luck and n outcomes 1</article-title>
          . http://archive.advancedfootballanalytics.com/
          <year>2007</year>
          /08/luckand-n -outcomes.html,
          <source>Accessed</source>
          <volume>22</volume>
          /06/2015
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>Data Mining: Practical Machine Learning Tools and Techniques with Java Implementations</article-title>
          . Morgan Kaufmann (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>h</year>
          .,
          <string-name>
            <surname>Keshri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iyengar</surname>
          </string-name>
          , G.:
          <article-title>Graphical model for baskeball match simulation</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          In: MIT Sloan Conference (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          4.
          <string-name>
            <surname>Pomeroy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Advanced analysis of college basketball</article-title>
          . http://kenpom.com
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          5.
          <string-name>
            <surname>Weissbock</surname>
          </string-name>
          , J.:
          <article-title>Theoretical predictions in machine learning for the nhl: Part ii</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Zimmermann</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moorthy</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <source>Predicting college basketball match outcomes Cluster 1 Cluster 2 Cluster 3 Cluster 4 Cluster 1 108/201 110/320 20/119 19/121 Cluster 2 362/416 610/960 105/354 175/394 Cluster 3 197/197 458/500 264/418 191/251 Cluster 4 179/191 373/454 111/245 163/258</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>Table 8. Wins and total matches for different cluster pairings</article-title>
          ,
          <source>2012 Cluster 1 Cluster 2 Cluster 3 Cluster 1 507/807 89/374 272/567 Cluster 2 569/607 622/967 518/578 Cluster 3 435/611 119/381 358/572</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>Table 9. Wins and total matches for different cluster pairings, 2013</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>