Showing posts with label regression toward the mean. Show all posts
Showing posts with label regression toward the mean. Show all posts

June 1, 2011

Two months in, a Bayesian look at the standings

At the end of April, I posted "Early season standings and Bayes" that took two different approaches to regressing the early season standings to come up with a prediction for the eventual result for the full season.

So here we are at the end of May, and there's been a lot of movement in the standings, so here's an update to the spreadsheet. Although the Phillies and the Indians remain at the top of the standings, they are starting to regress downwards. At the end of April, both teams were "on pace" to win 112 games in the season, but the regression showed a more modest result of 93 wins. A month later both teams are "on pace" to win 100, but the Bayesian approach suggests that they are more likely to win 91 games.

If you are a Twins or an Astros fan, there is no solace in the fact that both teams have not regressed toward the mean over the past month, but have instead continued to play at roughly the same level they exhibited in April. The regression model now predicts the Twins will end up at 65 wins, which would be the lowest in MLB. Of course, this prediction is based only on the team's performance to date -- it doesn't consider the number of injuries the Twins are currently dealing with.

-30-

May 1, 2011

Early season standings and Bayes

Early season performance has been a hot topic this year (not that it isn't a topic of discussion every year).  I wrote about it, using a simple approach of assuming that every team is .500, and a more recent addition in the blogosphere is Rob Neyer's take.

Last week Kincaid over at 3-D baseball had great post that used Boston's 2-10 start to go down a detailed and more sophisticated Bayesian path to estimating the team's true talent.  Tango posted a link to Kincaid's blog, and added a few details that incorporate actual observations.  A key element in this is that the observed spread of talent is wider than the theoretical .500 level of all teams.  (If all teams were .500, the random component would result in a standard deviation of 0.039. In reality, the standard deviation is wider, at 0.071 -- the implication of this is that there are real talent differences between teams, with some teams having a true talent level above .500 and others below.)

My modest contribution to this thread is here:  a Google doc spreadsheet that show all of the MLB team's current record (as of 2011-04-30), and then takes two different Bayesian-based methods to predict each team's final season outcome.

The first set are the yellow columns, which replicate Kincaid's "shortcut" approach, with the implied regression of 69 games noted by Tango. The blue columns take a different approach that uses the standard deviation of both the observed performance to date and the long-term observations (every MLB team season outcome from 1961-2010) as the prior.

The difference in the result generated between these two approaches is relatively modest.  (It's worth noting that the relative position on the standings does not change.) What is apparent is that with roughly 25 games played this season, there are solid differences appearing in the team performances.  This method forecasts that Cleveland and Philadelphia will regress downward from .692 (a 112 win season) to .573 and end up with 93 wins. At the bottom of the table, it suggests that the Twins will improve from .346 (58 wins on the season) to .444 and a much more respectable 74 wins over the course of the season.

-30-



April 11, 2011

Social mobility toward the mean

The April 8, 2011 edition of the BBC radio program More or Less* includes a discussion of regression toward the mean in the context of social mobility stats in the U.K.  Most of the analysis has focussed on the impact social class has on long-term education outcomes. In particular, much has been made of the fact that the analysis suggests that the low ability children from high social class catch up and pass the high ability children from low social class. 


But in the broadcast Daniel Read, professor at Warwick Business School, has offered a critique (link to written version) that points out that the analysis has not accounted for "one of the oldest statistical problems of all" (the BBC's description): regression toward the mean. The source of the problem is correctly identified as the bias introduced by including only the highest and lowest performers in the groups shown in the chart. The children closest to the mean for that social class have been excluded.

Because only the extreme ends of the education outcomes tests of the two social class groups have been selected, the poorest performers naturally show improvements while the higher performers show declines. From the broadcast:
It's not that it's [the differences in outcomes between social class] all fluke. But if there's any element of luck at all -- which there surely is, because we're talking about ability tests for toddlers -- then we have to allow for what we'd expect to happen when that luck fails to last.  And what we'd expect to happen is pretty much what the graph in the government's social mobility strategy shows, which is that the next time you test the children all the high performers have dropped off. But especially the poorer kids who, remember, Nick Clegg says were disadvantaged from birth. And all the lower performers have caught up, but especially the richer kids. And then as you continue to test, the richer kids gain on the poorer kids at a very much less dramatic pace.
The easiest way to spot the regression towards the mean? The enormous change from the first to the second measurement, as much of the selection bias at the first measurement point disappears. The high and low performers were selected not on the basis of their long-term outcomes, but on the results of the first test. In subsequent tests the children in these extreme cases will move toward the mean, and closer to their "true talent".

Accounting for regression toward the mean does not mean that social class doesn't have a relationship with education outcomes. But accounting for the regression toward the mean would moderate the magnitude of the difference between the two social classes.

*The linked page has a text summary of the program, a copy of the chart in question, streaming audio of the program, and links to the podcast and supporting documents. The item begins at roughly 17' 25" of the podcast.

-30-

December 22, 2010

The ERA distribution curve

NOTE: Tango and MGL at The Book, through the "Lemonade" thread, have critiqued the analysis below and pointed out errors in my assumptions. These errors mean that my closing conclusions are wrong -- good math, but bad statistics on my part. Through this post, you will find italicized text describing my errors.
Revised January 5, 2011




J.C. Bradbury's recent blog postings (here and here) have included histograms showing the distribution of ERA across major league pitchers for the 2009 season.  For his analysis, Bradbury omitted those pitchers with fewer than 100 batters faced -- in both his blog and book Hot Stove Economics, he justifies this due to the wide variance in ERA scores, much of which will be due to the small number of "samples" for each pitcher.  (As we saw in my earlier post about Bo Hart, it's possible for an average player to do very well over the short term; the inverse applies too.)

But a few comments on Bradbury's blog from readers ask about the impact that those "missing cases", who account for nearly a third (28%) of all individuals who pitched in MLB in 2009, would have on the curve.

Here's the answer: 

Figure 1: MLB Pitching, 2009 -- Number of Pitchers by ERA, by Number of Batters Faced





Incorporating the <100 BFP pitchers (the black chunks of each bar) adds pitchers across the whole range, although they are skewed to the right (i.e. higher ERAs).  While there is a stack on the left with very low ERAs, there's a bigger group of players with an ERA greater than 10. (The highest ERA of this group was 135.00.)

NOTE 1: ERA is a poor measure to use for this type of evaluation -- for pitchers with a low number of batters faced or innings pitched, it's easy for huge numbers to appear. That 135.00 ERA is the equivalent of 15 earned runs with only a single recorded out.  These exaggerated values then lead to an upward distortion of the mean for the group.  A better measure would be wOBA, or other measure that resembles a probability between 0 and 1.

The table below shows the average ERA of this group and three other groups based on the number of batters faced.  What we see is that the <100 BFP pitchers have a higher ERA than those who pitched more frequently.  (This difference is statistically significant.)  In spite of the variation in their ERAs, this group on average are less skilled than the other three groupings of pitchers.

NOTE 2: This is where I went wrong. The math is correct, but there is bias in the sample that I ignored. We can be fairly confident that pitchers who get off to a poor start won't get many opportunities to pitch -- and therefore won't get the opportunity to regress to the mean. Pitchers who do better at the start of their season will continue to pitch, and regress to the mean.  This process may take them some time, which may push them over the arbitrary line of 100 batters faced.  Thus the statistical significance is an artifact of the bias.

Figure 2: MLB Pitching, 2009 -- Average ERA, by Number of Batters Faced



In a thread on The Book blog that covered this same topic, I made a similar statement (reply #8): "What I’m trying to say is that our best estimate of the “true talent” of this group is an ERA of 8.11 [in the current case, 8.72], and that estimate is quite accurate". That statement got a response from Tango (reply #9) of "That is not accurate. If you look at how those pitchers who faced fewer than 100 batters did in the season preceding or the season following, THAT will give you a much better indicator of the true talent level."

So let me clarify.  The average level of skill of the pitchers who faced fewer than 100 batters in 2009, is an average ERA of 8.72. Although Tango is correct in his assertion that the poorest performers would regress upwards, by the same token the best pitchers (some of whom managed a 0.00 ERA in their short stint) would get worse. But if we were to let all 188 of them continue to pitch, we can be 95% certain that the "true" ERA of the group would end up somewhere between 6.92 and 10.52.

Even the lower bound (i.e. the lowest score we would expect with our more rigorous testing) is higher than the highest range from the other groups.

NOTE 3:  My statement above would be correct, if it were not for the bias in the sample.  My belief had been that this group would regress not to the MLB average, but to the average of the <100 BFP pitchers.  But because of the selection bias, this does not hold true.
Here's a simple example to demonstrate how this works. Think of the probability professor's favourite tool, the coin toss. If we have a penny and toss it repeatedly -- say, 10,000 times -- and recorded the result each time, the proportion of heads would very accurately reflect the true probability of the individual penny. And we'd need plenty of tosses to get an accurate measure of the single penny.

But what if instead of one penny we had 188 pennies, and we varied the number of tosses each penny got? Although the average number of tosses would be 50, some pennies might get only one toss, while others would get as many as 100 tosses. Some of those short sequences might come up all heads, while others would heavily favour the tails. On average, though, across the 188 pennies, we would find that the group average was a close reflection of "true average" of the group.

NOTE 4: The error in the initial assumption causes my coin flipping analogy to fall apart.  If “success” is a head, then the coin that comes up heads >0.5 will keep being flipped, possibly with enough flips to no longer be part of the “low flip” group (over that arbitrary threshold).  Meanwhile, a coin that runs tails more often will get pulled from the trials quickly, and end up <0.5 and with few flips.  Thus, as a group, the coins with a smaller number of flips will end up looking worse than those that keep getting flipped.  Selection bias causes an apparent difference, where none really exists.


And so it is with the pitchers in question. If they were like the other pitchers in MLB, we would expect that some of the <100 batters faced pitchers would have ERAs above the league average, while others would fall below. What we see, however, is that while there is a wide variation, the average is substantially higher than the other groups of pitchers.

NOTE 5:  ...because of selection bias!  The lesson:  selection bias can crop up anywhere, even if you are not the one doing the selecting.

-30-

December 10, 2010

Slugging regression II

Building on my previous post, this time around we'll look at a bigger group of hitters, those with at least 75 at-bats in both 2007 and 2008. This is a total of 360 players.
Theoretically with fewer at-bats, we would see a greater number of very high SLG values and also a larger number of below-average SLG values. But we've already seen hints that player talent gets evaluated early on (in the previous post, I identified the fact that the worst SLG in the 400+ group wasn't as awful to the same degree as the best hitters are good).

How to read the charts below: in both cases, there are 25 players plotted. Those that fall between the 100% and zero lines are regressing to the league mean. And the closer they are to the line, the bigger the regression. As shown in Figure 1, 22 of the top sluggers regressed toward the mean in 2008, 3 improved (led by Albert Pujols) and none fell below the league average.

For these players, 66% of their 2008 SLG score was accounted for by their 2007 SLG (and therefore the league average accounted for 44%).

An interesting observation is that these players are by and large the same as the 400+ AB group I dealt with in the previous post. Of the 25, 19 had 400+ ABs in both years. And of the remaining 6, 4 of the players had below 400 in 2007 and then over 400 in 2008. This group includes familiar names -- Josh Hamilton, David Murphy, and Cody Ross. All of them are young sluggers who did well in a short stint in 2007, and were given the opportunity to continue to play in 2008.

Figure 1: Top 25 SLG (2007), minimum 75 at-bats

For hitters at the bottom of the slugging table, we see a similar pattern of regression. Figure 2 shows SLG "improvement" in the opposite direction: the closer the bar gets to the bottom of the chart, the bigger the improvement. Thus of the 25 players, 17 regressed toward the mean without achieving it, and 3 others exceeded the league average (the ones who fell "below zero"). The remaining 5, on the other hand, started out below average in 2007 and fared worse in 2008.
For this group, the previous year's SLG accounted for only 55% of their 2008 SLG.
And for this group of 25, they are decidedly not the same players as the least sluggerly of the 400+ AB group. Only 1 -- Jason Kendall -- appears in both lists.

In short, the "survivor bias" that keeps good players active with opportunities to hit has an inverse impact on the players at the bottom of the stack. Unless they show a huge improvement that brings them much closer to the league average, these players seem to be destined to part-time roles.

Figure 2: Bottom 25 SLG (2007), minimum 75 at-bats

This analysis is best described as "proof of concept". Certainly any conclusions drawn should be tentatively stated, and a more robust analysis over a greater number of seasons is warranted.

-30-

Slugging regression

Tango issued a multi-part challenge, of which the first part is:
1. Take the top 10 in SLG in each of the last 10 years, and tell me what the overall average SLG of these 100 players was in the following year.


The point of the challenge is to demonstrate that top performing players will regress toward the mean in subsequent seasons, and that the year under consideration accounts for, as a rule of thumb, 70% of the next season's performance, and the league average (to which their performance regresses) the other 30%.


In algebraic terms, X is predicted to be 70% when
X = (SLG2 - LSLG)/(SLG1-LSLG)
Where:
SLG1 is season 1 slugging average,
SLG2 is season 2 slugging average,
LSLG is the average league slugging average (from season 1)

Leo quickly responded (comment #1 to Tango's post), with his calculation that for slugging, 73.3% was accounted for by the player's average in the first season. To my way of thinking, the challenge has been met -- job well done, Leo!

But as I started to think about it further, I began to wonder how far through the rankings this rule of thumb holds -- as we approach the league average, the player's SLG and the league SLG become one and the same number. And at the opposite end of the scale -- the non-sluggers -- do they regress upwards towards the mean?

So my first step was to simplify the challenge, and only look at two consecutive seasons, 2007 and 2008. Using only those players who had a minimum of 400 at bats each season, I pruned the list down to 129 players in both the NL and AL. Simplifying matters further is the fact that the 2007 SLG for the NL was the same as the AL -- .423. So for my "top sluggers" I then looked at the top 25 across both leagues.

The result: for these 25 players, on average, 66% of their 2008 SLG was accounted for through their 2007 score. A few percentage points from Tango's rule of thumb, but close enough.

Charting the results shows that all but two of the top 25 sluggers regressed downwards towards the mean. And of the two, only one improved dramatically: Albert Pujols (who inched up still further in 2009, before regressing ever-so-slightly in 2010). Were Pujols not in the mix, the 2007 SLG would account for only 62% of the 2008 scores.






Another interesting observation is that of these top performers, not one fell so far in 2008 to end up with a SLG below the league average. That's not to say that it wouldn't happen, but it suggests that at the extreme end of the performance curve, as determined over the course of a full season, top performers really are above average. (NOTE: further testing required!)

But what of the other end of the ranking? I looked at the lowest performing players that I had selected, and the rule of thumb does not work. From the bottom up, the percentage explained was 87%, 84%, -4.4%, -28%, ...

At this point, I started to wonder -- why minus values? A quick check of the numbers, and I saw that these players regressed up, and to a point above the league average.

So what's different about the bottom of the range? It's simple: survivorship bias. My "sample" of 139 players who had 400+ ABs in each of 2007 and 2008, while ensuring I found the top hitters, automatically excluded those weak-slugging players who don't get many plate appearances but who collectively drag down the league average. Thus the "worst" players of the 139 with lots of ABs were not (by and large) far from the league average. The bottom of the list was Jason Kendall, who slugged .309 in 2007 for the A's and the Cubs while catching. Perform much worse than that, and you'll end up playing Triple A. Or in Kendall's case, for the Royals.


On deck: regression toward the mean, SLG with 75+ ABs.


-30-

November 22, 2010

Bo knows probability



Cardinals' second baseman Bo Hart



Over on 3-D Baseball, Kincaid has a nice explanation of regression to the mean in a post titled "On Correlation, Regression, and Bo Hart". The blog entry starts with the story of Bo Hart, who got called up to the Cardinals in June 2003, and promptly hit .412 over his first 75 at-bats. Since Kincaid wrote a regression to the mean article, you can guess where Hart's season went -- he finished with 286 at-bats and a .277 average.

But Kincaid flirts with a few notions that I think are worth following in a bit more detail.

First up, what are the odds that a .277 hitter will break .400 across a string of 75 at-bats?

The answer is roughly 1 in 200.

This is calculated through the fact that the binomial distribution approximates the normal distribution -- in English, if you repeat a set of binomial trials, the histogram of the count of success rates for the trials will look like the normal curve. This leads us to the probability density function, which allows us to state the probability that a value (in this case, a batting average of .412) falls at a certain point given the mean value (.277).

Using Bo Hart's season batting average of .277 as his "true talent" (or "population mean") across 75 at-bats, we can calculate the standard deviation of the distribution (0.052). We then determine that .412 lies at 2.60 standard deviations from the mean (2.60=[.412-.277]/.052). As a probability, 2.60 standard deviations is 0.5% -- or 1 in 200.

What was unusual about Bo Hart is that his 1 in 200 string of successful at-bats occurred at the beginning of his Major League career. Calculating that probability is a task for another day.

In my next post I will explore Kincaid's statements about evaluating "true talent" based on a number of observations. Specifically, I'll delve into the following questions: "At what point can we be relatively certain about our inferences of true talent based on observed performance? 75 PAs is not enough, and one million is plenty, but what about 1000?"

-30-