Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This can be explained without bringing in accusations of fraud or incompetence [1].

Basically, if you have a bazillion experiments running and only publish the results which are significant, you still have a huge predictability problem, because most of those statistically significant results will be due to chance alone.

For example, assume that only 1% of cancer experiments produce something valid and interesting, and that standard for statistical significance is the industry standard 95%, or 1 in 20. If 100,000 hypothesis are tested, you'll get 100,000 * 1% = 1000 valid results and 100,000 * 5% = 5000 results from chance alone. In this example, only 16% of the published results will be reproducible, because only 1000 / 6000 are valid.

[1] http://sciencehouse.wordpress.com/2009/05/08/why-most-publis...



this basic fact is understood by scientists and children alike. it is still either incompetence (lack of middle school statistics) or fraud (ignoring the above fact).




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: