Showing posts with label real life. Show all posts
Showing posts with label real life. Show all posts

December 14, 2011

OT: Unemployment Rates in the U.S.A.

A great example of a truly awful graph was posted on Flowing Data, starting a conversation on The Book.

I posted the following comment there:

Wexler/27 beat me to the BLS data sets... I will  note that the "discouraged" numbers can be found in the "characteristics of the unemployed" tables.  The different ways of parsing what constitutes "unemployment" are ways to try to get to the nuance in why people aren't working.  The narrow "actively seeking work" definition of defining who is in the labour market is a way to cut through demographic changes (e.g. in the post-WWII period, when most women were June Cleavering and not seeking work), increases in post-secondary enrollments, etc.
With that said, the increase in people who have thrown in the towel is, to me, one of the most disturbing parts of the current recession.
Wexler/7 and MGL/8 raise the question of attribution -- how much is the President (in the current circumstance, Obama) responsible for unemployment rates?
I dug up some historical U.S. unemployment rate data going back to 1948, and I have posted a chart of it to my own blog (because I have no idea how to do that directly).
A summary:  the current increase in unemployment started in the last year of G.W. Bush's 2nd term (rising from 5.0% in January 2008 to 6.8% by the November, the month of the election).  Going back, there was a peak in unemployment during the first G.W. presidency, and another that straddled G.H.W. Bush and Clinton.  And the worst unemployment rates since the Great Depression (higher rates and with a longer peak than the current phase) were during the first term of Ronald Reagan's presidency.  


Here's a chart of U.S. unemployment rates from January 1948 to November 2011:

And for those of you wanting to focus on the more recent period, the past 20 years (less a month):





I obtained the U.S. Department of Labor data set from the Economic Research pages of the Federal Reserve Bank of St. Louis here, and have converted that text file to an Excel file.

More Bureau of Labor Statistics (BLS) unemployment data can be found here, here and here.

-30-

May 6, 2011

When labour market research goes to the ballpark

In a recently issued paper called "Productivity, Wages, and Marriage: The Case of Major League Baseball", economists Francesca Cornaglia and Naomi E. Feldman examine the "marriage premium" -- the fact that controlling for other influencing factors, married men earn more than unmarried men. In most situations, a variety of confounding variables muddy the waters -- things like geographic location, differences across occupations, and poor productivity measures. Cornaglia and Feldman innovatively use information from MLB to control for those variables.

Derek Jeter, the exception that proves the rule.

The abstract:

Using a sample of professional baseball players from 1871 - 2007, this paper aims at analyzing a longstanding empirical observation that married men earn significantly more than their single counterparts holding all else equal. There are numerous conflicting explanations, some of which reflect subtle sample selection problems (that is, men who tend to be successful in the workplace or have high potential wage growth also tend to be successful in attracting a spouse) and some of which are causal (that is, marriage does indeed increase productivity for men). Baseball is a unique case study because it has a long history of statistics collection and numerous direct measurements of productivity. Our results show that the marriage premium also holds for baseball players, where married players earn up to 20% more than those who are not married, even after controlling for selection. The results are generally robust only for players in the top third of the ability distribution and post 1975 when changes in the rules that govern wage contracts allowed for players to be valued closer to their true market price. Nonetheless, there do not appear to be clear differences in productivity between married and nonmarried players. We discuss possible reasons why employers may discriminate in favor of married men.

You can hear Dr. Cornaglia discuss the research on the BBC programme More or Less (2011-04-29), starting at roughly 10'50".

My initial reaction regards neither the findings nor the methodology, but the fact that other than a mention of the Lahman database, the list of references does not include any of work of the sabermetric research community. At one point in the discussion of productivity measures the authors write "Most modern-day baseball enthusiasts and commentators consider the latter two statistics [OPS and EqA] to be the most accurate measures of a player’s productivity", but the authors neither refer to any authority to support that statement nor discuss fact that others have critiqued those measures.

This is not the first time that academics have utilized the contributions of the sabermetric community in supporting their research (in this case, it provides a vital element in the foundation of the productivy measure) but then failed to acknowledge that work. For a well-reasoned discussion of that topic, please read Phil Birnbaum's "Chopped liver II".

-30-

April 11, 2011

Social mobility toward the mean

The April 8, 2011 edition of the BBC radio program More or Less* includes a discussion of regression toward the mean in the context of social mobility stats in the U.K.  Most of the analysis has focussed on the impact social class has on long-term education outcomes. In particular, much has been made of the fact that the analysis suggests that the low ability children from high social class catch up and pass the high ability children from low social class. 


But in the broadcast Daniel Read, professor at Warwick Business School, has offered a critique (link to written version) that points out that the analysis has not accounted for "one of the oldest statistical problems of all" (the BBC's description): regression toward the mean. The source of the problem is correctly identified as the bias introduced by including only the highest and lowest performers in the groups shown in the chart. The children closest to the mean for that social class have been excluded.

Because only the extreme ends of the education outcomes tests of the two social class groups have been selected, the poorest performers naturally show improvements while the higher performers show declines. From the broadcast:
It's not that it's [the differences in outcomes between social class] all fluke. But if there's any element of luck at all -- which there surely is, because we're talking about ability tests for toddlers -- then we have to allow for what we'd expect to happen when that luck fails to last.  And what we'd expect to happen is pretty much what the graph in the government's social mobility strategy shows, which is that the next time you test the children all the high performers have dropped off. But especially the poorer kids who, remember, Nick Clegg says were disadvantaged from birth. And all the lower performers have caught up, but especially the richer kids. And then as you continue to test, the richer kids gain on the poorer kids at a very much less dramatic pace.
The easiest way to spot the regression towards the mean? The enormous change from the first to the second measurement, as much of the selection bias at the first measurement point disappears. The high and low performers were selected not on the basis of their long-term outcomes, but on the results of the first test. In subsequent tests the children in these extreme cases will move toward the mean, and closer to their "true talent".

Accounting for regression toward the mean does not mean that social class doesn't have a relationship with education outcomes. But accounting for the regression toward the mean would moderate the magnitude of the difference between the two social classes.

*The linked page has a text summary of the program, a copy of the chart in question, streaming audio of the program, and links to the podcast and supporting documents. The item begins at roughly 17' 25" of the podcast.

-30-

July 27, 2010

Skill, luck, and more than a little style

Pictured: "Mr. May", Dave Winfield, comes through in the clutch in the 1992 World Series with, in his words, "One stinkin' little hit." The 11th-inning double drove in two runs and sealed the World Series win for the Blue Jays.


The BBC has posted an article and calculator ("Can chance make you a killer?") that is used to demonstrate the challenges in differentiating luck from skill. In this case, a simple scenario with fixed parameters is linked to a calculator that generates the range of possibilities.

While I'm not sure how this could be used in a baseball setting, it is a very good tool for demonstrating that it can be difficult -- particularly if you just look at "the numbers" in a selective way -- to make definitive statements about a player's ability. Such as, say, clutch hitting.

(Acknowledgement: The Book.)

July 26, 2010

Baseball imitates real life

How understanding luck in baseball can help understanding real life, or at least your investment portfolio: "Untangling skill and luck" by Michael J. Mauboussin.

Mauboussin uses a variety of sabermetric analysis, including Jim Albert's 2004 paper “A Batting Average: Does It Represent Ability or Luck?” and Tango's True Talent Level analysis.