"But the full hypothesis is 'Given the data, are girls better than boys at this exam?' and clearly, the prior probability is relevant."
No, prior probabilities have nothing to do with it.
We state our 'null' hypothesis that the boys and girls do equally well. This hypothesis has nothing to do with a belief of prior probabilities or belief of any probabilities at all. Instead, we state this hypothesis as something that will give us some mathematical assumptions to do some calculations to reject it and, then, conclude that it was false.
Generally in hypothesis testing we don't believe the null hypothesis as prior probabilities; indeed, likely we don't believe it at all and are stating it to reject it and conclude it is false.
In more detail, we assume that 20 boys and 18 girls are 38 independent samples from some one distribution. It turns out, we don't need to say anything about that distribution because we are being 'distribution-free'. In particular, we get to ignore the Gaussian distribution. GOOD.
Independent? Okay: Suppose we DO give you the true distribution of the data and the first 37 scores. Now you get to guess score 38. Do the 37 scores help you beyond just the distribution? No. Same for any subset of the scores. Then, we have independence.
With this null hypothesis, the average of the scores of the 20 boys and the average of the scores of the 18 girls should be 'close'. How close? Well, under the null hypothesis and with the values we observed, we have a way to proceed: The distribution of the difference in the scores, with everything we do know given and fixed, we can find. For this distribution, basically we look at all the 33 billion or so differences obtained by taking all combinations of 38 things taken 18 at a time. Justification? If work at it mathematically, then under the null hypothesis can show that each of those 33 billion cases was equally probable.
Then we pick a small number, say, 1% for the size of our Type I error, that is, the probability of rejecting the null hypothesis when it is true.
Then we find the differences in the 1% tail of the 33 billion differences.
Then we look at the difference from our actual data. That difference will be one of the 33 billion. We see if that difference is in the 1% tail.
If the difference is in the 1% tail, then one of two things is true:
(A) The null hypothesis is true, the boys and girls are the same, that is, independent samples from the same distribution, and with our actual data the difference is relatively large, out in a tail, and we have observed something that happens only 1% of the time.
(B) The null hypothesis is false, that is, in some way the boys and girls are different. That is, we still believe the independence assumption, so what is false is just that the mean for the boys is different from the mean for the girls.
If the 1% is so small we don't believe (A), then we conclude (B).
Variance has nothing to do with it.
Welcome to distribution-free 'two sample' hypothesis testing 101.
I've been reading Jaynes again this week, and he's just very, very convincing. And so I'm trying to read everything you wrote through these Bayesian glasses, but sadly, I'm not successful. Jaynes is rather critical of Fisher's hypothesis testing, on the ground that you can't accept or reject an hypothesis on its own; you need an alternative to compare it to, and that alternative needs to make definite predictions. I don't see what the alternative to your null hypothesis is (the negation of the null hypothesis does not make definite predictions)
No, prior probabilities have nothing to do with it.
We state our 'null' hypothesis that the boys and girls do equally well. This hypothesis has nothing to do with a belief of prior probabilities or belief of any probabilities at all. Instead, we state this hypothesis as something that will give us some mathematical assumptions to do some calculations to reject it and, then, conclude that it was false.
Generally in hypothesis testing we don't believe the null hypothesis as prior probabilities; indeed, likely we don't believe it at all and are stating it to reject it and conclude it is false.
In more detail, we assume that 20 boys and 18 girls are 38 independent samples from some one distribution. It turns out, we don't need to say anything about that distribution because we are being 'distribution-free'. In particular, we get to ignore the Gaussian distribution. GOOD.
Independent? Okay: Suppose we DO give you the true distribution of the data and the first 37 scores. Now you get to guess score 38. Do the 37 scores help you beyond just the distribution? No. Same for any subset of the scores. Then, we have independence.
With this null hypothesis, the average of the scores of the 20 boys and the average of the scores of the 18 girls should be 'close'. How close? Well, under the null hypothesis and with the values we observed, we have a way to proceed: The distribution of the difference in the scores, with everything we do know given and fixed, we can find. For this distribution, basically we look at all the 33 billion or so differences obtained by taking all combinations of 38 things taken 18 at a time. Justification? If work at it mathematically, then under the null hypothesis can show that each of those 33 billion cases was equally probable.
Then we pick a small number, say, 1% for the size of our Type I error, that is, the probability of rejecting the null hypothesis when it is true.
Then we find the differences in the 1% tail of the 33 billion differences.
Then we look at the difference from our actual data. That difference will be one of the 33 billion. We see if that difference is in the 1% tail.
If the difference is in the 1% tail, then one of two things is true:
(A) The null hypothesis is true, the boys and girls are the same, that is, independent samples from the same distribution, and with our actual data the difference is relatively large, out in a tail, and we have observed something that happens only 1% of the time.
(B) The null hypothesis is false, that is, in some way the boys and girls are different. That is, we still believe the independence assumption, so what is false is just that the mean for the boys is different from the mean for the girls.
If the 1% is so small we don't believe (A), then we conclude (B).
Variance has nothing to do with it.
Welcome to distribution-free 'two sample' hypothesis testing 101.