Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Hypothetical Statistics Question

Let \(\phi(x,\mu,\sigma)\) denote the normal distribution with mean \(\mu\) and standard deviation \(\sigma\) evaluated at \(x.\) Consider the model distribution $$p(x,\mu)=(1-t)\cdot \phi(x,\mu-3t,1)+t\cdot \phi(x,\mu+3\cdot(1-t),t^2),$$ for some known, fixed, but very small \(t>0.\) Since \(t\) is small, \(p\) looks very much like a normal distribution with variance \(1\) and mean \(\mu,\) except for a tall spike around \(\mu+3.\) See the picture below. The mean of \(p\) is \(\mu.\) Note in the last term, the variance is \(t^4.\) The total mass of the spike is \(t,\) thus small, and so the cumulative distribution for \(p\) will look very much like the standard cumulative distribution for \(\phi(x,\mu,1).\)

Suppose we have a sample \(x_0,\) assumed to be from a distribution \(p\) of the form above, but of unknown mean. To repeat, \(t\) is known. Consider the hypothesis: $$H_0: \mu=x_0-3\cdot(1-t)$$ Do we reject \(H_0\)?

If \(H_0\) is true, then \( p(x_0,\mu)\approx t \cdot \phi(x_0,x_0,t^2) = 1/(t \sqrt{2\pi}),\) and this is very large, about \(1/t\) times larger than \(p(\mu,\mu).\) Indeed, \(x_0-3\cdot(1-t)\) is very close to the maximum likelihood estimate for \(\mu.\) These seem to be reasons not to reject \(H_0.\)

On the other hand, the overall mass of the spike is small, and the entire spike is well out on the tail of \(p.\) Thus it is in some sense unlikely to get an \(x_0\) out there, assuming \(H_0\) is true. Since the cumulative distribution looks very much like the normal cumulative distribution for \(\phi(x,\mu,1),\) it makes sense to apply the usual test and reject \(H_0.\)

I find that not rejecting in this case is the right thing to do, but I am not sure what others might think. I also wonder how often such a question is relevant.

 

What is an Unusual Event?

My random thoughts on random events posted as a blog comment:

Suppose we have disjoint events \( (E_1,E_2,\dots,E_N),\) and corresponding probabilities for these events \( (p_1,p_2,\dots,p_N),\) where \( p_n=\mbox{Prob}(E_n)\) and \(\sum_{n=1}^N p_n=1.\) If a particular event \(E_k\) occurs, what would make us think this was in some sense "unusual" or perhaps "suspicious"? It's not enough that \(p_k\) be small, since for large \(N\), even a uniform distribution on the \(E_n\) will have \(p_k=1/N\) small. Nor is it enough that \(p_k\) be much less than \( \max_n p_n,\) since it is possible that all \(p_n\) are equal except for one event having many times larger yet still tiny probability. It's not enough if \(p_k\) is less than nearly all the other \(p_n\), because all the \(p_n\) could be very nearly equal.

What does seem to work in the cases I can think of is to choose some factor \(R \gt 1\), calculate \(\sum\{p_n: p_n \gt R p_k\}\), and see if this is close to \(1.\) To work this into a hypothesis test, we could reject the null hypothesis \(H_0\) if $$\sum\{p_n: p_n\gt R p_k\} \gt (1-1/R),$$ though the expression on the right-hand side is rather arbitrary. With this setup, what value should \(R\) be? Let \(x_0\) be a sample we have collected, and consider the standard normal and the \(p=.05\) rule, where \(\mbox{Prob}(|x_0|\gt 1.96)=0.05.\) Then \(R=3.71,\) since \(3.71\cdot \phi(1.96) = \phi(1.1),\) and \(\mbox{Prob}(|x_0|\gt 1.1)=1/3.71.\) If we wanted \(R=20,\) we would need to use a cutoff \( |x_0|\gt 3.135749,\) which corresponds to a very small standard \(p\)-value of \(0.001714.\)

Clearly, given any \(p\) cutoff, a.k.a \(\alpha\), we can find a corresponding factor \(R,\) and vice-versa. Since the \(p=.05\) rule is arbitrary, I don't see what difference it makes for the most common cases. Thus, \(p\)-value analysis seems generally ok to me in practice. My concern here is with its justification.

Color Me Skeptical

I am mulling over the following, which I posted as a comment here.

I have a question about Iowahawk’s analysis. In his follow-up post he describes Simpson’s paradox…:

“Hitter B has a higher batting average against both righties and lefties, but Hitter A has a higher overall average by dint of facing a different mix of pitchers. Now comes the question: it’s the bottom of the 9th, two out, and you need a base hit. Who would you insert as a pinch hitter, A or B? The detailed data suggests Hitter B, irrespective of whether the pitcher was right- or left-handed. The overall average, in this context, is worse than meaningless – it leads you to exactly the wrong conclusion.”

If this is legit analysis, it opens a can of worms, I think. Suppose that instead of handedness, we had pitchers with and without some “X Factor.” Suppose we didn’t even know what the X Factor is. Suppose we didn’t even know what the proportion of pitchers in the league have the X Factor. But suppose some third party did know, and simply provided us with a breakdown of our two hitters’ averages against the two types of pitchers, and they formed a Simpson’s paradox. Would it still be “wrong” to put in hitter A? Is it better to put in B, because he is better at hitting against both X and not-X pitchers? Color me skeptical.

Public Debt vs Vote for the O

I was reading this post by Chuck DeVore on Big Government:
According to Moody’s, the average state per capita debt of the 28 Obama states is $1,728 while the average debt in the 22 McCain states is less than half, at $749. This information alone says a lot about voters and their attitude towards government and debt. Voters with a propensity to elect politicians who burden future generations who can’t yet vote with huge debts voted for Obama while fiscally responsible voters generally voted for McCain.

And I thought I would do a scatter plot of each state's public debt versus its vote for Obama (click on image for bigger version):



It seems that while not all low-debt states were McCain states, all high-debt states were Obama states. And, the higher the debt, the greater the vote for Obama.

Hypothesis Testing - II

A comment I posted on a blog:

[previous commenter] has a good point, and I see this sort of thing all the time in medical research. Do you think the life expectancy of people who drink Coke is *exactly* equal to the life expectancy of people who drink Pepsi? Exactly equal? Of course not. If you sample enough people, eventually you will detect a "statistically significant" difference. Then you can publish your paper saying e.g. "we found that people who drink Coke live significantly shorter lives than people who drink Pepsi!" Or, as it would appear in the newspaper and on tv: "Coke Kills!!"

Hypothesis Testing

I think the way hypothesis testing is presented to students and justified makes no sense. Suppose we know that x has been drawn from a normal distribution with unit variance, and we want to test H_0: mu=0. Suppose x=2. Then they say “given H_0, the chance of seeing |x|>1.96 is less than 5%, so since x=2>1.96 we reject H_0.” Why does this make sense? Given H_0, the chance of seeing |x|<.01 is less than 5% too. Would you reject the null if our sample were x=0? No! So then they talk about x being “extreme,” i.e. far from the mean. What exactly does distance from the mean have to do with it? Suppose we knew that x was drawn from a uniform distribution on some interval [mu-1/2,mu+1/2], and again we wanted to test H0: mu=0. If x=.4999 would you reject? No, that makes no sense, because given H0, x=.4999 is no less likely than x=0. You could reject if x=.5001, but not if x is in the interval (-1/2,1/2). You could easily find an example of a bimodal distribution where the pdf at the mean is zero. Then you should reject if the sample is near the mean! Distance from the mean is not in general relevant.

I am being pedantic, but hypothesis testing works e.g. for the standard normal distribution f because if x>1.96, then f(x) is much less than values of f near x=0, not because of areas at the tails or distance from the mean.

It's science!

I love this sentence:
The problem is that 71.3% of what passes as peer reviewed climate science is simply junk science, as false as the percentage cited in this sentence. [ Willis Eschenbach ]
That reminds me of this graph:

Bayes, Laplace and the Sun

William M Briggs mentions Laplace's Rule of Succession in a recent blog post. Briggs' is a blog about statistics and related matters that I highly recommend. Laplace used the rule, which relies on Bayes' Theorem, to calculate the probability that the sun will rise tomorrow. It is an elegant and fascinating bit of analysis. According to Wikipedia, Laplace's method give odds 1826250:1 in favour of the sun rising tomorrow.

But I beg to differ! Using Bayesian analysis, I calculate that the probability that the sun will rise tomorrow is 1/2!

Here is my argument. Let N+1 be the number of days from the start of history through tomorrow. Suppose at first that we know nothing about the sun, neither the related physics nor the past history of its rising. This was Laplace's assumption as well. Suppose we only know that the sun rises on some subset of the N+1 days in question. With only this knowledge, we assume a uniform prior probability distribution on this subset. Thus we assume that all 2^(N+1) possible subsets of the N+1 days are equally likely to be the sun-rising subset, each having probability 2^(-N-1). Now suppose we are given additional knowledge, specifically that the sun rose on the first N days. There are then only two possibilities for the sun-rising subset: the set of all N+1 days, and the set that contains only the first N days. By our prior assumption, and a trivial application of Bayes Rule, we see that each of these possibilities now has the posterior probability of 1/2. Thus the probability of the sun rising tomorrow is 1/2.

Put that in your pipe and smoke it! Related post here.

Save Halloween

A little Halloween common sense from Leonore Skenazy's book and blog "Free Range Kids":
Was there ever really a rash of candy killings? Joel Best, a professor of sociology and criminal justice at the University of Delaware, took it upon himself to find out. He studied crime reports from Halloween dating back as far as 1958, and guess exactly how many kids he found poisoned by a stranger’s candy?

A hundred and five? A dozen? Well, one, at least?

“The bottom line is that I cannot find any evidence that any child has ever been killed or seriously hurt by a contaminated treat picked up in the course of trick-or-treating,” says the professor. The fear is completely unfounded.

Now, one time, in 1974, a Texas dad did kill his own son with a poisoned Pixie Stix. “He had taken out an insurance policy on his son’s life shortly before Halloween, and I think he probably did this on the theory that there were so many poison candy deaths, no one would ever suspect him,” says Best. “In fact, he was very quickly tried and put to death long ago.” That’s Texas for you.

A little perspective

From an article on Pajamas Media, written by Soeren Kern, here's some European reaction to the US health care debate:
[A]nother Independent article titled “Republicans, religion and the triumph of unreason” says: “Here’s what’s actually happening. The U.S. is the only major industrialized country that does not provide regular health care to all its citizens. Instead, they are required to provide for themselves — and 50 million people can’t afford the insurance. As a result, 18,000 U.S. citizens die every year needlessly, because they can’t access the care they require. That’s equivalent to six 9/11s, every year, year on year.”
Shall we put those numbers in a little perspective?

(First of all, my guess is that they are playing fast and loose with the word "citizen", considering a large number of the uninsured are illegal aliens.)

In a nation of 300 million people, that amounts to 6 of 100,000 people.

In 2003, a heat wave struck Europe. It resulted in more than 37,000 deaths across a comparable population--meaning more than 12 in every 100,000 people died. US heat-wave mortality for the years 1999-2003 was 3,442, for an annual average of 688.

The UK had 2,139 deaths due to the heat-wave. In a population of about 60 million, that's every 3.5 out of 100,000 people died.

In 2007 hospital-acquired c-diff infections in the UK claimed the lives of 8,324, or more than 13 of every 100,000 people. (The numbers fell by 29% in 2008.)

This last statistic I have is rather shocking. Looking at cancer deaths in the US, we had 559,303 in 2005, for an death rate of 186 per 100,000. In the UK in 2007, they had 155,484 cancer deaths, for a death rate of 259 per 100,000. A death rate nearly 40% higher than in the US. That means an excess of 43,800 Britons are dying each year due to sub-standard health care.

So, where exactly is there a triumph of unreason, and whose religion is clouding their judgement?

Uninsured

Powerline posts a now-common break down of the "46 million uninsured" number. They mention people eligible but not yet signed up for government insurance, the young and healthy, the relatively well off, and the non-citizen, but they still miss something.

This is the final number that they give for the uninsured:
This leaves about 15.5 million (one-third of Obama's 46 million) who actually are uninsured, cannot become insured simply by enrolling in a free program, are U.S. citizens, and cannot easily afford to purchase insurance. About 5 million members of this cohort are childless adults.
The remaining problem is this: if you make a list of the names of each of those 15.5 million and look back in on them in 3, or 6, or 12 months, how many of them would still be uninsured? And how many of them were only uninsured briefly as they were looking for work or between jobs? The 47 million number, just like the number of people living in poverty, gives the illusion of a static cohort, when in fact people move in and out of insurance and in and out of poverty all of the time. Being without insurance for a couple of months is a very different thing from being a long-termed, unwilling uninsured person, or person who can't get health insurance because of pre-existing conditions.

Interestingly, over on Hot Air, the Captain posted an excerpt from the proposed bill (Here's a link to the text from Thomas.gov -- the Congress's own computers):
(a) Tax Imposed- In the case of any individual who does not meet the requirements of subsection (d) at any time during the taxable year, there is hereby imposed a tax equal to 2.5 percent of the excess of–
‘(1) the taxpayer’s modified adjusted gross income for the taxable year, over
That should be a pretty nasty provision. Go without health insurance for one day--while changing jobs, of if you employer screws up and accidentally doesn't pay premiums on time, or if you're short of cash for a couple of months and fall behind on your premiums, you're socked with a tax of 2.5% on your gross income.