https://statmodeling.stat.columbia.edu/2023/05/20/if-you-do-not-know-what-you-would-have-done-under-all-possible-scenarios-then-you-cannot-know-the-type-i-error-rate-for-your-analysis/ Skip to primary content Statistical Modeling, Causal Inference, and Social Science Search [ ] [Search] Main menu * Home * Authors * Blogs We Read * Sponsors Post navigation What happened in the 2022 elections Is omicron natural or not - a probabilistic theory? "If you do not know what you would have done under all possible scenarios, then you cannot know the Type I error rate for your analysis." Posted on May 20, 2023 9:23 AM by Andrew Jose Iparraguirre writes: Just finished "Understanding Statistics and Experimental Design. How to Not Lie with Statistics", by Michael Herzog, Gregory Francis, and Aaron Clarke (Springer, 2019). Near the end (p. 128), I read the following regarding "optional stopping": ...suppose a scientist notes a marginal (p = 0.07) result in Experiment 1 and decides to run a new Experiment 2 to check on the effect. It may sound like the scientist is doing careful work, however, this is not necessarily true. Suppose Experiment 1 produced a significant effect (p = 0.03), would the scientist still have run Experiment 2 as a second check? If not, then the scientist is essentially performing optional stopping across experiments, and the Type I error rate for any given experiment (or across experiments) is unknown. Indeed, the problem with optional stopping is not the actual behavior preformed by the scientist (e.g., the study with a planned sample size gives p = 0.02) but with what he would have done if the result turned out differently (e.g., if the study with a planned sample size gives p = 0.1, he would have added 20 more subjects). More precisely, if you do not know what you would have done under all possible scenarios, then you cannot know the Type I error rate for your analysis. It is the last statement, "if you do not know what you would have done under all possible scenarios, then you cannot know the Type I error rate for your analysis", that kept me wondering and prompted me to write to you and ask you for your comments. My reply: 1. Yes, they're correct that if you do not know what you would have done under all possible scenarios, then you cannot know the Type I error rate for your analysis. We make this point in section 1.2 of our Garden of Forking Paths paper and this is part of the definition of these error rates; the statement should not be controversial. 2. I'm pretty much not interested in type 1 or type 2 error rates; mostly the only reason I think it's worth thinking about them is to respond to confused researchers who think that a low p-value represents strong evidence for their preferred theories. 3. I think that stopping the experiment based on the data is just fine in practice and is not cheating, as some people might think. See my discussion here. I also sent to Greg Francis, who wrote: I too don't think the statement is controversial, even though it is surprising to many practicing scientists. In part I think that is because textbooks on hypothesis testing give examples with very specific settings (always with a fixed sample size). Even though (usually) the textbooks properly describe theorems (e.g., for defining a sampling distribution), they ignore some realities of data collection (a fixed sample size is not the norm) and so do not consider what happens when you deviate from the theorems. Contrary to Andrew, I think a lot of scientists genuinely do care about Type I and Type II error rates. However, getting control of those error rates is more difficult than many people realize, and I think that difficulty is good motivation to consider Bayesian approaches. For a bit more discussion about error rates across experiments, you might look at this paper on a "reverse Bonferroni" method. I'm not sure I would actually recommend the method for any practical situation, but it highlights what to consider when you try to control error rates across experiments. This entry was posted in Decision Analysis, Miscellaneous Statistics by Andrew. Bookmark the permalink. 23 thoughts on ""If you do not know what you would have done under all possible scenarios, then you cannot know the Type I error rate for your analysis."" 1. [71d3fdf5]David J. Marcus on May 20, 2023 10:03 AM at 10:03 am said: If only people were taught Bayesian statistics, they would stop caring about Type I and Type II errors. I recommend Inference for a Bernoulli Process (A Bayesian View) by D. V. Lindley and L. D. Phillips The American Statistician, Vol. 30, No. 3 (Aug., 1976), pp. 112-119 http://www.jstor.org/stable/2683855 and the book "The Likelihood Principle" by James O. Berger and Robert L. Wolpert. Reply | + [d582]Andrew on May 20, 2023 10:29 AM at 10:29 am said: David: Sure, but just remember that it not necessary that Bayesian methods conform to the likelihood principle. Reply | o [71d3]David J. Marcus on May 20, 2023 6:03 PM at 6:03 pm said: I wasn't suggesting that we replace teaching hypothesis tests with teaching the likelihood principle. I was suggesting that people who have been taught the usual school (frequentist) statistics read the article and book. Reply | 2. [d7e3e90d]paul alper on May 20, 2023 10:05 AM at 10:05 am said: Change "I'm pretty much no interested in type 1 or type 2 error rates;" to "I'm pretty much not interested in type 1 or type 2 error rates;" As to "optional stopping," what sporting event allows that? Does there exist a sporting event which keeps on going only until my side is eventually triumphant? In other words, my side never loses because the game, by definition, may never end. Reply | + [d582]Andrew on May 20, 2023 10:29 AM at 10:29 am said: Typos fixed; thanks. Reply | o [ad1f]iainjgallagher on May 20, 2023 11:30 AM at 11:30 am said: You changed the wrong typo... it was better as 'Scottish'... "Ah'm no interested in type 1 or type 2 error rates". Credentials: Ah'm Scoattish! Reply | # [3ac1]McMe on May 21, 2023 1:09 AM at 1:09 am said: Pure dead brilliant. But 'sgotty be "Type Wan" 'n' "Type Twa", dya no hink? + [538c]Dale Lehman on May 20, 2023 10:52 AM at 10:52 am said: It depends which side you're on. In college softball there is the rule limit - a lead of 8 runs ends the game after 5 innings. In golf, there is "stroke control" which limits your score on a hole to 10 strokes. So, these are forms of stopping rules that limit how bad things can get (or, alternatively, they limit the possibility of a comeback in the case of softball: unfortunately, there is no way to lower a golf score on a hole by taking additional strokes). Reply | + [2685]Jonathan (another one) on May 20, 2023 11:15 AM at 11:15 am said: Poker is a perpetual game of optional stopping by both sides, so long as somebody will advance you credit. Reply | + [28b7]John N-G on May 20, 2023 12:11 PM at 12:11 pm said: Fergie Time. Refers to a perception in English football that referees allowed a game to extend longer if a particular team needed to score. Reply | 3. [1e06a63a]Anoneuoid on May 20, 2023 11:12 AM at 11:12 am said: Suppose Experiment 1 produced a significant effect (p = 0.03), would the scientist still have run Experiment 2 as a second check? Check out this paper about the pioneer anomaly: https://arxiv.org /abs/1204.2507 The scientists bend over backwards to find ways to not see a significant effect, because the null hypothesis corresponds to their hypothesis. This is exactly opposite the NHST incentive. Reply | + [1e06]Anoneuoid on May 20, 2023 11:27 AM at 11:27 am said: Most here have read this paper already, but Paul Meehl called it a paradox: https://meehl.umn.edu/sites/meehl.umn.edu/files/files/ 074theorytestingparadox.pdf How can there be something we call "science", but the exact opposite procedures are also called science? It is like bizarro-science. Reply | 4. [e2e9dd2f]samuel on May 20, 2023 11:38 AM at 11:38 am said: So getting control of Type 1 error just requires pre-specifying all possible experiments/studies you might want to run over the course of your career? Sounds easy! Reply | 5. [cbf7434d]Carlos Ungil on May 20, 2023 12:10 PM at 12:10 pm said: > the scientist is essentially performing optional stopping across experiments, and the Type I error rate for any given experiment (or across experiments) is unknown I don't get it. How do "across experiments" decisions affect the type I error rate for any given experiment? Or does the "(or across experiments)" remark indicate that "for any given experiment" doesn't mean each experiment taken independently? If I throw a die the probability of getting a 6 is 1/6. If I can throw again, and stop when I want, I will eventually get a 6. But the probablility of getting a 6 for any given throw is 1/6 even if I perform optional stopping across throws. Reply | + [0dfa]John N-G on May 20, 2023 12:33 PM at 12:33 pm said: But if your hypothesis is that the probability of getting a 6 is at least 1/5*, and you choose when to stop, most experiments will confirm your hypothesis. *or some such number larger than but sufficiently similar to 1/6. Reply | o [cbf7]Carlos Ungil on May 20, 2023 12:47 PM at 12:47 pm said: > suppose a scientist notes a marginal (p = 0.07) result in Experiment 1 and decides to run a new Experiment 2 to check on the effect In this case each throw would be a different experiment. It would not correct to say in this case that "most experiments will confirm your hypothesis." Or maybe when they say "the Type I error rate for any given experiment is unknown" they don't mean "the Type I error rate for Experiment 1 is unknown" and "the Type I error rate for Experiment 1 is unknown" . Reply | # [0a75]Josh on May 20, 2023 5:56 PM at 5:56 pm said: Carlos, you seem to be assuming the experiments are independent. They are not because whether the second experiment happens depends on the outcome of the first experiment. It's like deciding to flip a coin until it comes up heads and then for the first flip that does come up heads saying, this is a fair coin so this was a fifty-fifty result. # [cbf7]Carlos Ungil on May 21, 2023 4:24 AM at 4:24 am said: I'm not assuming anything, I'm trying to understand what they wrote: "Suppose Experiment 1 produced a significant effect (p = 0.03), would the scientist still have run Experiment 2 as a second check? If not, then the scientist is essentially performing optional stopping across experiments, and the Type I error rate for any given experiment (or across experiments) is unknown." If that's not wrong at least it seems very ambiguous because after discussing "Experiment 1" and "Experiment 2" it's far from obvious that the claim "for any given experiment ..." must not be interpreted as "for Experiment 1 ..." and "for Experiment 2 ...". # [71d3]David J. Marcus on May 21, 2023 8:45 AM at 8:45 am said: This is the problem with frequentist analysis: You have to decide what the plan was rather than just look at what data you have. Were they always planning to do a second experiment to confirm their results? (This reminds me of the recent intent to treat discussion.) For a clear example of how ridiculous it is to let textbook statistics box you into this corner, see the article " Inference for a Bernoulli Process (A Bayesian View)" by D. V. Lindley and L. D. Phillips, The American Statistician, Vol. 30, No. 3 (Aug., 1976), pp. 112-119, http://www.jstor.org/stable/ 2683855 And for people who think you can ignore priors, consider the example of the lady tasting tea. 6. [4a7d4f01]Michael Lew on May 20, 2023 6:01 PM at 6:01 pm said: The conundrum presented by the scenario just shows how badly the Neyman-Pearsonian hypothesis test approach marries with scientific approaches to gaining knowledge. The accept/reject dichotomy is necessary to the accounting required for type I and type II error calculations, but is inimical to efficient learning from experimental results. If you get a P-value from one experiment that is neither small enough nor large enough to convince you to discontinue a line of investigation, then you SHOULD continue in one way or another. A fresh dataset cannot not influence the evidential meaning of the original dataset and so a fresh experiment analysed independently of the original dataset would constitute good science whether the accountants of error like it or not. Type I and type II error rates are related to the properties of the analysis, not to the evidence in the data. You can read a longer account here: https://link.springer.com/chapter/10.1007/ 164_2019_286 Reply | + [e2e9]samuel on May 20, 2023 7:14 PM at 7:14 pm said: +1, and thanks for the reference. Reply | + [1e06]Anoneuoid on May 20, 2023 9:19 PM at 9:19 pm said: Those graphical power functions show clearly the three-way relationship between sample size, effect size and the risk of a false negative outcome (i.e. one minus the power). Actually, your figure shows a four-way relationship. Alpha (significance threshold) is also important. What happens in practice is that sample size is a function of how expensive it is to run the experiment. Eg, you will rarely see studies of more than 10 primates per group but this is common for rodents. For data like microarrays or particle collisions it can run into the hundreds, thousands, or more. Then alpha is chosen so that there are enough "discoveries", but not so many that it seems too easy. This corresponds to a power of 25-50%, so you need to at least try out a few different things before getting something to publish. Ie, alpha is a function of the typical effect size for that type of data, along with the achievable sample size. It caps out at 0.1, and below that gets rounded to some easily remembered number. Obviously 0.05 is a very popular one, most likely what gets considered "too expensive" by funding agencies is also a function of that standard choice. Reply | 7. [4bab043b]Psyoskeptic on May 21, 2023 10:49 AM at 10:49 am said: So... what about a replication from another lab done because the initial study was statistically significant. What's the Type I error rate for that? What about the many labs studies replicating an experiment many times? Reply | Leave a Reply Cancel reply Your email address will not be published. Required fields are marked * [ ] [ ] [ ] [ ] [ ] [ ] [ ] Comment * [ ] Name [ ] Email [ ] Website [ ] [Post Comment] [ ] [ ] [ ] [ ] [ ] [ ] [ ] D[ ] * Art * Bayesian Statistics * Causal Inference * Decision Analysis * Economics * Jobs * Literature * Miscellaneous Science * Miscellaneous Statistics * Multilevel Modeling * Papers * Political Science * Public Health * Sociology * Sports * Stan * Statistical computing * Statistical graphics * Teaching * Zombies 1. Anoneuoid on Principal stratification for vaccine efficacy (causal inference)May 23, 2023 5:56 PM These were the results: Of the participants not meeting infection criteria and deemed uninfected, low-level non-consecutive viral detections were observed... 2. Anonymous on Bayesian Believers in a House of Pain?May 23, 2023 5:10 PM I've tried to write some lyrics fitting to this post. No robots or computers were involved (as far as I'm... 3. Daniel Lakeland on What exactly is a "representative sample"?May 23, 2023 3:51 PM A sample that can be analyzed usefully is any sample you know enough about how it was generated to write... 4. Unanon on Principal stratification for vaccine efficacy (causal inference)May 23, 2023 3:25 PM 2) High specificity tests: For example, PCRs for COVID-19 have very high specificity, but tend to have sensitivities in the... 5. Raphael Nishimura on What exactly is a "representative sample"? May 23, 2023 1:44 PM That's precisely the type of problem that I have with this term: in sampling, we never talk about a single... 6. Raphael Nishimura on What exactly is a "representative sample"? May 23, 2023 1:26 PM In general terms, I completely agree with you: everyday language is full of terms like that and I have absolutely... 7. Andrew on Going for it on 4th down: What's striking is not so much that we were wrong, but that we had so little imagination that we didn't even consider the possibility that we might be wrong. (I wonder what Gerd Gigerenzer, Daniel Kahneman, Josh "hot hand" Miller, and other experts on cognitive illusions think about this one.)May 23, 2023 1:15 PM Aeforty: With probabilities you always want to use expectation---that is, arithmetic mean---because the expectation of an expectation is an expectation.... 8. aeforty2 on Going for it on 4th down: What's striking is not so much that we were wrong, but that we had so little imagination that we didn't even consider the possibility that we might be wrong. (I wonder what Gerd Gigerenzer, Daniel Kahneman, Josh "hot hand" Miller, and other experts on cognitive illusions think about this one.)May 23, 2023 12:26 PM I believe the 4th Down charts use the arithmetic mean of win probability added (WPA) to assess which scenarios are... 9. Matt Skaggs on What exactly is a "representative sample"?May 23, 2023 12:08 PM "I know it when I see it!" Wins thread, IMO. Language is full of terms like "representative sample" that cannot... 10. Andrew on What exactly is a "representative sample"?May 23, 2023 11:24 AM Christian: Yes, part of the point of my post was to emphasize that I do think that "representative sample" is... 11. Anoneuoid on Principal stratification for vaccine efficacy (causal inference)May 23, 2023 11:22 AM I didn't follow the entire argument but they seem to make three questionable assumptions: 1) That vaccine effectiveness can be... 12. OliP on What exactly is a "representative sample"?May 23, 2023 11:15 AM This is a nice summary of the issue, and relates to the Meng (2018) definition, where one needs to know... 13. Andrew on Principal stratification for vaccine efficacy (causal inference)May 23, 2023 10:24 AM Link added; thanks. 14. Rodney Sparapani on Principal stratification for vaccine efficacy (causal inference)May 23, 2023 10:05 AM Looks like this post is referring to this article... https:// arxiv.org/abs/2211.16502 15. gec on What exactly is a "representative sample"?May 23, 2023 9:56 AM I have discovered a truly marvelous example of such a number, but this comment is too small to contain it. 16. Raphael Nishimura on What exactly is a "representative sample"? May 23, 2023 8:56 AM Thank you for your reply! I'm sorry to keep insisting, but I don't see how useful can be a concept... 17. Raphael Nishimura on What exactly is a "representative sample"? May 23, 2023 8:40 AM Ideas of Sampling is a little book very well-written, but it is very introductory. It's a good start if you... 18. Christian Hennig on What exactly is a "representative sample"?May 23, 2023 8:33 AM Just that a word is around doesn't mean it has an agreed correct definition. As with some other terms, if... 19. Andrew on What exactly is a "representative sample"?May 23, 2023 4:33 AM Raphael: I agree that a sample can be nonrepresentative and still be useful (the point of your fourth paragraph above).... 20. Andrew on What exactly is a "representative sample"?May 23, 2023 4:28 AM Josh: No, you're givings a definition of a sampling procedure. I'm asking about the sample itself. There's just the sample,... 21. OliP on What exactly is a "representative sample"?May 23, 2023 3:51 AM Didn't Meng (2018) define this as when the "data defect correlation", rho{R,Y}, where R is an indicator variable defining in/out... 22. Werner W. Wittmann on What exactly is a "representative sample"? May 23, 2023 3:28 AM A sample is fully representative if there exists a single real number,which when multiplied with the sample leads to the... 23. Anoneuoid on What exactly is a "representative sample"?May 23, 2023 12:00 AM can be represented by sufficiently long periodic samples 0, 1, 2, 3, 0, 1, 2, 3, 0, 1, 2, 3,... 24. Anon Y Mouse on What exactly is a "representative sample"?May 22, 2023 10:19 PM I appreciate the information! Wasn't aware there was this much research on sampling. Is The Idea of Sampling the best... 25. Daniel Lakeland on What exactly is a "representative sample"?May 22, 2023 9:31 PM Algorithmic randomness gives us a definition of a random sequence as any sequence that passes some uniformly most powerful computable... 26. Daniel Lakeland on What exactly is a "representative sample"?May 22, 2023 9:24 PM Also would need proof but I suspect for any given epsilon, there exists a finite sample length N such that... 27. theCamel on What exactly is a "representative sample"?May 22, 2023 6:25 PM I know it when I see it! 28. Daniel Lakeland on What exactly is a "representative sample"?May 22, 2023 5:40 PM Representativeness of a sample is an absolute requirement for accurate Frequentist inference... you have an estimator the frequency properties of... 29. somebody on What exactly is a "representative sample"?May 22, 2023 5:05 PM Also, we have to be careful about what we're sampling and whether samples are sets (unordered) or sequences (ordered). An... 30. Howard Edwards on What exactly is a "representative sample"?May 22, 2023 4:50 PM Diaconis also made some videos on card shuffling: https:// youtu.be/AxJubaijQbI 31. Howard Edwards on What exactly is a "representative sample"?May 22, 2023 4:42 PM Persi Diaconis did a lot of work in the 1980s answering this question. Diaconis was a professional magician who became... 32. somebody on What exactly is a "representative sample"?May 22, 2023 4:36 PM Woah, typos There are infinitely many functions to integrate against (expectations to take), so almost every finite sample will not... 33. somebody on What exactly is a "representative sample"?May 22, 2023 4:33 PM My take: From a measure theoretic perspective, a probability distribution is something that gets integrated against a function If all... 34. Chris Wilson on What exactly is a "representative sample"?May 22, 2023 3:42 PM Great question! I personally am convinced that entropy measures capture what we want :) IOW, once the entropy saturates with... 35. Dale Lehman on What exactly is a "representative sample"?May 22, 2023 2:13 PM I don't see how any distance measure can adequately be used as a definition. Distances can only be measured for... 36. James Hanley on What exactly is a "representative sample"?May 22, 2023 1:34 PM See series of articles in 1979 by Kruskal and Mosteller (Int Stat Review) 37. Kyle C on What exactly is a "representative sample"?May 22, 2023 1:24 PM For a taste of what it can be like to be a judge or lawyer, read these comments and then... 38. Carlos Ungil on What exactly is a "representative sample"?May 22, 2023 1:06 PM Related: What's a well-shuffled deck of cards? 39. Anoneuoid on Is omicron natural or not - a probabilistic theory? May 22, 2023 12:15 PM You can passage a virus through human cells and select for variants with whatever property. If you are really cynical,... 40. somebody on What exactly is a "representative sample"?May 22, 2023 11:36 AM The issue with a process based definition is how to justify "rerandomization" and enforced covariate balance in experimental designs with... 41. Noah Motion on What exactly is a "representative sample"?May 22, 2023 11:25 AM Your second paragraph sounds a lot like Andrew's definition, and it invites at least a couple challenging questions: How "same"... 42. Raphael Nishimura on What exactly is a "representative sample"? May 22, 2023 11:00 AM This is my favorite pet peeve and something I've been arguing for over a decade (and many other sampling statisticians... 43. Josh on What exactly is a "representative sample"?May 22, 2023 10:20 AM I suppose a necessary but not sufficient condition would be for the sample allow us to make estimates that are... 44. Anonymous on What exactly is a "representative sample"?May 22, 2023 10:15 AM Why I cannot define it like the following? Suppose the sample was drawn superpopulation with distribution F(x). Through a sample,... 45. Roman on What exactly is a "representative sample"?May 22, 2023 10:11 AM I do believe there is at least one "official" defintion. Statistically representative sample is a subset generated from a sampling... 46. Chris on Is omicron natural or not - a probabilistic theory?May 22, 2023 8:08 AM It is an interesting observation (Omicron variant has 29 non-synonymous mutations and only 1 synonymous one). Clearly a variant with... 47. Gregory C. Mayer on Is omicron natural or not - a probabilistic theory?May 22, 2023 7:21 AM Here's some of the background substance. A non-synonymous substitution in a protein-coding gene changes the protein's structure; a synonymous substitution... 48. Kevin Nelson on What happened in the 2022 electionsMay 22, 2023 6:26 AM I just finished reading the Catalist report. It's not bad, but it could have been a lot better. Instead of... 49. Kevin Nelson on What happened in the 2022 electionsMay 22, 2023 4:58 AM The "midterm penalty" is one of the most consistent patterns in American politics, going way back to the nineteenth century.... 50. Anonymous on The Puzzle of Paul Meehl: An intellectual history of research criticism in psychologyMay 22, 2023 4:57 AM Quote from above: "And it doesn't matter concerning these 3 things that some finding is not true or turns out... Proudly powered by WordPress