Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Wednesday, July 21, 2010

Time series graph

There are lots of time series graphs available, but sometimes there are more issues that you want to include.  This is sometimes accomplished with bubble charts, but here's a nice example using geography as the "x-axis":

Saturday, July 17, 2010

Graphs can say one thing, or another

First, I am not posting this to make any statement about Paul Krugman, or anything like that.  But when comparing time series data, choosing the starting point can be more imporant than it seems, as this blogger demonstrates.  You could imagine investing strategies being compared in a similar fashion.  One graph shows strategy A is better, but with the same data and a different start date, another graph would show that strategy B is better. 

Another issue is the choice of the other data sets - they can also be chosen to look "fair" but those time series contain special features that make the analysis less than fair.  For example, compare company X with specially chosen company Y, because of the special charges for Y in a particular quarter make some observation about X seem more true.

Simple comparisons are not always so simple.

Sunday, April 11, 2010

A post about something I know little about

One of the questions I like to pose in my classes is "how do you know?"  Doing analysis of data and possible decision leads to lots of wild speculation, and this question sometimes brings us back to recognizing that widely held theories are actually quite tenuous.  One observation can bring down a whole theory, a "black swan" event.  I like to use examples from physics, even though I am not a physicist, because I find that students often think that we "know" stuff there.  Science is settled, even if other areas are not.

Here is an interesting (to me) article about measuring light and comparing the measurements to accepted physical theories.  At the end, something to think about:
There’s also a possibility that the explanation could be even more far-reaching, such as that the universe is not expanding and that the big bang theory is wrong.
How certain are you of the things you know?

Saturday, March 6, 2010

Rankings

Here is the most admired companies list for this year.  There are quite a few ways to set up the mechanics for the 'scoring' and there is not much in the way of confidence intervals in the article.  Would the article be more or less effective if there were statistical ties in the rankings?  How many ties would there be?

Here are the undergraduate business school rankings for 2010.

Monday, January 25, 2010

Christian statistics

Not statistics that act in a Christian way, but rather trying to use statistics to understand the state of Christianity.  Is it growing?  Shrinking?  Healthy?  Dying?  Turns out that describing statistics collected and making inferences is tricky, and the link describes some nice examples of where things can go wrong.  Thanks to Brad for the link.

Sunday, January 10, 2010

When is an average or not an average?

First, I have not followed up on the reference in this Instapundit blog post, so I do not know whether this is really the way the British Met Office really computes average temperatures:
In fact, the Met still asserts we are in the midst of an unusually warm winter — as one of its staffers sniffily protested in an internet posting to a newspaper last week: “This will be the warmest winter in living memory, the data has already been recorded. For your information, we take the highest 15 readings between November and March and then produce an average. As November was a very seasonally warm month, then all the data will come from those readings.”
But suppose for a minute that an average for a whole season (year) really is fifteen days, and that the highest fifteen all (most) came from November. Would that really represent the average for the winter? Can you think about why they would not use every day's reading from November through March? Should they use just the high for the day, or the high and the low for each day? And where should the reading come from? How many locations around England would be enough for what would seem like a good average calculation? What if averages in the past were computed differently? Could you make comparisons?

Average seems like such an easy concept - but it often is quite tricky. To say nothing about trying to understand the variability of temperatures (do they compute standard deviations?)

For example, the fifteen day record could occur in a season where the other one hundred and thirty five days were pretty cold...

Monday, December 28, 2009

What do we know?

I am naturally somewhat of a skeptic, but it is hard to not read something and think about how smart the scientists are, and how much they know, and what will soon be possible. One of the things that is really interesting is brain research. Suppose we could know how we work?

All kinds of new tools allow researchers to "watch" your brain work. Or can they? How do they "know"?

Here some interesting reading about the difficulties in measuring, knowing and statistics.

Here is another that makes you think twice when we assume that science is a clean process that follows a straight line to the truth, getting it right along the way.

And here's a cartoon view of how this works (or doesn't work).

Friday, December 11, 2009

Cartoon with statistics

This cartoon site has quite a bit of "geek" humor, so I like it. It also has an interesting feature where if you move the mouse over the cartoon, a hidden message appears. This one has a comment about "the mother of all sampling biases." It may not seem funny, but it is a really great example. Some of his other cartoons are pretty good. This one makes me wonder about teaching statistics.

Friday, November 27, 2009

Simpson's paradox

Averages seem like simple things, but not always. Consider this simple baseball example:

Tony and Joe are competitive friends and so they compare batting averages. At the All-Star break, Tony is batting .300 and Joe is only batting .290. Joe mentions that batting in the second half of the season is more important, and so he and Tony agree to compare their batting averages for the second half of the season (and only the second half). When they finally meet, it turns out that Tony batted .390 in the second half of the season. Joe did better, too, but only batted .375. Tony wins both halves of the season.

Question: who's batting average was higher for the entire year? Turns out we don't know, and it could very easily be Joe! (I'll post an example later)

What the paradox states is that averages for subgroups can demonstrate relationships that are inconsistent with averages for different subgroups or the overall averages. So Tony could win both halves of the season, but have a lower batting average for the entire season.

I do not know the details about Climategate, but it is very interesting to me, with all the statistics. Are average temperatures going up or down or whatever. Here's an article that about midway through mentions this batting average paradox. I wonder if I have a new example of the paradox involving averages.

Wednesday, October 28, 2009

Telling a story with a graph

This website has some interesting graphs, and I especially like this one. You have to think for a second before it makes sense, but once it does, it provides a nice explanation for some of the mess people have gotten themselves into (they believed the fiction).

Friday, October 16, 2009

Bubble charts

New charts that are popular these days are called bubble charts. They show data over time, usually over a couple of axes by using bubbles of different sizes. Now I don't know about the source data for this page, but the graph is pretty cool. And a little scary.

Saturday, September 26, 2009

Means and medians

Here is an article from the Wall Street Journal about the housing market in Detroit. Here's how the story starts:

On a grassy lot on a quiet block on a graceful boulevard stands the answer to a perplexing question: Why does the typical house in Detroit sell for $7,100?

The brick-and-stucco home at 1626 W. Boston Blvd. has watched almost a century of Detroit's ups and downs, through industrial brilliance and racial discord, economic decline and financial collapse. Its owners have played a part in it all. There was the engineer whose innovation elevated auto makers into kings; the teacher who watched fellow whites flee to the suburbs; the black plumber who broke the color barrier; the cop driven out by crime.

The last individual owner was a subprime borrower, who lost the house when investors foreclosed.

Then the article sites some statistics:

And the median selling price for a home stood at a paltry $7,100 as of July, according to First American CoreLogic Inc., a real-estate research firm -- down from $73,000 three years earlier. A typical house in Cleveland sells for $65,000. One in St. Louis goes for $120,000.
Now I understand that median house prices are the standard way of talking about what is "typical." Using the mean creates problems when most houses are of one value, but there are some outlier, high-priced houses.

But in this article, is median misleading in a different way? Did the house on Boston Blvd sell for $7,100? How many houses were sold in Detroit? Voluntarily? Can the dynamics associated with foreclosures make the median a poor choice to report?

Often we try to summarize with a few statistics, and that can be useful. But here I think there is need for more information to really understand this situation. The reporter probably had access to that data, if they wanted to report it.

Monday, August 31, 2009

Statistics versus anecdotes

When you report on something and you have statistics and anecdotes, which should win? Which should be the basis for your conclusions? If you had data on a large part of the US population and then you talked to your brother-in-law, which would be more important to report?

Check out this NYTimes article about the Exodus from Facebook, and think about how statistics and andecdotes are used. Ouch.

Wednesday, August 26, 2009

Fun with probabilities

Here's a real life application of probability theory. It is similar to the sports betting company business model:

You get an mailing list (email, it is cheaper), and you pick a big football game. You send half the list an advisory that one of the teams will win, and you send the other half an advisory that the other will win, along with the opportunity to subscribe to your newsletter for a low, low price.

To the half of the people that got the correct prediction, you send another letter, picking another big game, and sending half the prediction that one team will win, half the prediction that the other team will win.

People who get two correct predictions in two weeks will be amazed and will be more likely to subscribe to your newsletter.

But you can split that group in half, and send out a third prediction, then split the people who got three weeks in a row right and send out a fourth letter, and so on. You might recycle some of the losers, too, so that some people see that your accuracy is 5 out of 6 weeks and so on.

Imagine getting a letter where the predictions were right five weeks in a row! What is the probability of that? This newsletter must be really good!

Or the newsletter writer knows some probability...

Friday, August 7, 2009

NYTimes article about statistics

I have to link to this article, For Today's Graduate, Just One Word: Statistics mostly because of this quote:
“I keep saying that the sexy job in the next 10 years will be statisticians,” said Hal Varian, chief economist at Google. “And I’m not kidding.”
Works for me. There's also a nice example at the end of the article about how just finding relationships in data is not always enough:
For example, in the late 1940s, before there was a polio vaccine, public health experts in America noted that polio cases increased in step with the consumption of ice cream and soft drinks, according to David Alan Grier, a historian and statistician at George Washington University. Eliminating such treats was even recommended as part of an anti-polio diet. It turned out that polio outbreaks were most common in the hot months of summer, when people naturally ate more ice cream, showing only an association, Mr. Grier said.

Tuesday, March 18, 2008

What do we know?

Science depends on statistics to allow us to "know" something. We experiment and show differences and try to show that those differences are not just random, but show some underlying effect.

Sometimes, even statistics are not enough, when enough people in a field are convinced of the current theories -- which are theories because they can never be "proved." An example is the struggle of two scientists who saw the data tell them that ulcers were not caused by what "everyone" knew they were caused by, but rather by a bacteria. Publishing their results was frustrated because the reviewers knew they must be wrong.

Lucky for us they persevered, and eventually on the Nobel Prize.

Tuesday, February 26, 2008

Pie Charts

Most pie charts are not very efficient at showing relationships (often you can just write the numbers instead). I finally found a pie chart that does more.

Wednesday, February 13, 2008

Housing prices

Here's an interesting set of graphs of the movement of housing prices over the years. Think about what the graphs "say" and the implications of the way they are constructed. What changes could you make to tell a the kind of story that the text argues for?