[HN Gopher] 68-95-99.7 Rule
       ___________________________________________________________________
        
       68-95-99.7 Rule
        
       Author : turrini
       Score  : 113 points
       Date   : 2023-04-09 11:39 UTC (11 hours ago)
        
 (HTM) web link (en.wikipedia.org)
 (TXT) w3m dump (en.wikipedia.org)
        
       | H8crilA wrote:
       | Remember that a lot of data is not actually normally distributed,
       | even though it looks so at first. The horror of "risk management"
       | units of financial institutions, when they get a "7-sigma" event
       | twice in a month.
       | 
       | Can you imagine there was a time when the entire US options
       | market was running on flat volatility vs strike price (implying
       | that financial asset prices are lognormal). Must have been cool
       | to just buy the cheapest, most OTM options and collect big
       | results every now and then.
        
         | bombcar wrote:
         | Also if it's once in 300 million to happen every day, it could
         | happen to someone every day in the USA (assuming it's something
         | people do daily).
         | 
         | If the outfall from a 7-sigma is disastrous, you have to try to
         | prevent it, because it WILL happen eventually given enough
         | goes.
        
           | H8crilA wrote:
           | That's true but those people were sampling closer to once per
           | day, and still got N-sigma in their models for some high
           | value of N.
           | 
           | The easiest way to get something that's not actually normal
           | in the typical central limit theorem construction is to have
           | the variables be just slightly not independent. Which becomes
           | obvious in retrospect after you start realizing the
           | dependence and shit hits the fan.
        
             | [deleted]
        
             | bombcar wrote:
             | Banks failing and CDOs exploding would be two good
             | examples; they're rare if not nonexistent but if they start
             | they may all flare up.
             | 
             | Also natural distribution assumes you don't have someone
             | trying to game the extremes to make money. If being seven
             | feet tall resulted in billions we'd have people researching
             | how to grow taller people.
             | 
             | You can see stuff like this in distributions of tax returns
             | - many people will be unexpectedly clustered right below
             | certain cut-off amounts.
        
               | looping__lui wrote:
               | Worse is correlation break-down in a sense - e.g., that
               | black swan event happening across an entire asset class
               | because of liquidity squeeze, margin calls, etc.
        
         | paulpauper wrote:
         | _Can you imagine there was a time when the entire US options
         | market was running on flat volatility vs strike price (implying
         | that financial asset prices are lognormal). Must have been cool
         | to just buy the cheapest, most OTM options and collect big
         | results every now and then._
         | 
         | This may have been possible pre-1987. Since then, out of money
         | put options have a very steep skew, so not possible. Even if
         | OTM put options are very cheap nominally, but by having much
         | higher implied volatility makes them much more expensive
         | compared to how much they would otherwise cost without the
         | skew. This from a Kelly perspective makes them much less
         | lucrative if one was to try to construct a strategy with this.
         | It cuts your ROI big time.
        
           | User23 wrote:
           | In an efficient market the average ROI on options should be
           | zero.
        
             | looping__lui wrote:
             | But it isn't efficient; banks like overcharge you absurdly
             | especially for OTM options.
        
         | jrm4 wrote:
         | One time for Nassim Nicholas Taleb who's ALL OVER this in a
         | deeply interesting and entertaining way. "The Black Swan" is
         | great at going in on such a weirdly obvious-but-not point.
        
           | paulpauper wrote:
           | Unlike other hedge funds, Taleb's/Universa's records, like
           | CAGR, are sketchy. Good records do not exist except what has
           | been 'leaked' to the media by PR. It's not like Bridgewater
           | in this regard.
           | 
           | His barbell OTM put strategy almost certainly did badly in
           | 2022, because of the failure of the Vix index to spike and
           | the decline of the S&P 500 being orderly.
           | 
           | https://greyenlightenment.com/2022/10/08/tail-hedging-
           | strate...
           | 
           | So both parts of the barbell lost money, which is not
           | supposed to happen. This is why you have to be warry of
           | individuals and strategies that are overhyped. If a strategy
           | is getting a lot of positive press, it likely means it's
           | saturated. If a lot of quants are buying the same options for
           | the same hedging purposes, who is going to sell them?
        
             | jrm4 wrote:
             | Judging Taleb by a hedge fund's performance, even if it's
             | one of his, is really missing the point entirely.
        
       | djaychela wrote:
       | The table of numerical values is pretty interesting for a layman.
       | I've often heard scientists (such as Brian Cox) talk about
       | x-sigma probabilities, but putting it the context in this table
       | is much more meaningful for someone like me who only has self-
       | study reading as scientific understanding.
        
       | Eddy_Viscosity2 wrote:
       | I like how they call this a 'rule' instead just a list of three
       | hard to remember numbers in a row.
        
       | tedunangst wrote:
       | Memorizing a set of numbers doesn't really feel like a shorthand
       | way to remember those numbers. Did you know about the 3.14 rule?
       | It's a way to remember the first three digits of pi are 3.14.
        
         | bumbledraven wrote:
         | An more useful thing to remember is Shah's approximation, which
         | estimates the area under the standard normal curve from 0 to a
         | point z on the x axis as follows:                 * z(4.4 -
         | z)/10 when 0 <= z <= 2.2       * 0.49 when 2.2 < z < 2.6
         | * 0.50 when 2.6 <= z
         | 
         | Shah's approximation is accurate to within roughly +-1/2%.
         | 
         | Examples:
         | 
         | - The area under the standard normal curve within 1 standard
         | deviation of zero is approximately 2(1(4.4 - 1))/10 = 0.68.
         | (The factor of 2 is there to get the area on both sides of
         | zero).
         | 
         | - The area within 2 standard deviations of zero is
         | approximately 2(2(4.4 - 2))/10 = 0.96.
         | 
         | - The area within 3 standard deviations of zero is
         | approximately 1 (because 2.6 <= 3).
         | 
         | - The probability of sampling from the standard normal
         | distribution and getting a value less than 1.5 deviations above
         | the mean is approximately 0.5 + 1.5(4.4-1.5) = 0.935. (The 0.5
         | term is there because we need to include the area to the left
         | of 0, and Shah's approximation only counts the area to the
         | right of zero.)
         | 
         | Reference:
         | 
         | Arvind K. Shah (1985) A Simpler Approximation for Areas under
         | the Standard Normal Curve, The American Statistician, 39:1, 80,
         | DOI: 10.1080/00031305.1985.10479396 (https://twitter.com/jordan
         | curve/status/958026273149915136/ph...)
        
         | sroussey wrote:
         | Which is, I don't know -- ironic? -- because pi rounded to two
         | decimals is 3.15.
        
           | karmakaze wrote:
           | The 3.2 rounding of Pi in an Indiana bill[0] is also quite
           | infamous, which confusingly had other values of Pi for use in
           | different contexts. [I'm a Tau touter myself.]
           | 
           | [0] https://www.straightdope.com/21341975/did-a-state-
           | legislatur...
        
           | elijaht wrote:
           | No it's not- it's 3.141, which rounds to 3.14
        
             | midasuni wrote:
             | It's
             | 
             | 3.14159
             | 
             | 3.1416
             | 
             | 3.142
             | 
             | 3.14
             | 
             | 3.1
             | 
             | 3
        
           | euroderf wrote:
           | The approximation 355/113 has nice pairs of odd digits.
        
           | scythmic_waves wrote:
           | No pi is 3.14159... which is 3.14 when rounded.
        
       | abakker wrote:
       | I always like the inverse of this, which is that if you calculate
       | that less than 68% is in the first standard deviation, you can
       | guess that the data isn't normally distributed.
       | 
       | It's really a heuristic on where to start with analysis, though.
       | Not a result in and of itself.
        
         | qsort wrote:
         | The gist of the idea is of course correct, but if you actually
         | do this, _please use an actual normality test_.
         | 
         | https://en.wikipedia.org/wiki/Shapiro-Wilk_test
         | 
         | https://en.wikipedia.org/wiki/Anderson-Darling_test
         | 
         | https://en.wikipedia.org/wiki/Kolmogorov-Smirnov_test
        
           | disgruntledphd2 wrote:
           | The trouble with normality tests is that you either have too
           | little data (in which case they're useless) or too much (in
           | which case their also useless).
           | 
           | I always think of normality as a nice theory for simple
           | stuff, but you want to be running sensitivity analyses for
           | all of your assumptions, including normality (which I don't
           | think I've ever seen professionally, tbh).
        
       | paulpauper wrote:
       | This is why also claims of IQs above 180 or so are almost always
       | BS, at least on normed tests. 6sd is about the maximum
        
         | scythmic_waves wrote:
         | It's also literally impossible for many modern tests. For
         | example the WAIS-IV, a very popular measure, has a maximum
         | possible FSIQ of 160. With a mean of 100 and an SD of 15,
         | that's "only" 4 SDs from the mean (99.997 percentile).
         | 
         | And as you say, modern tests just aren't calibrated in that
         | way. The WAIS-IV is calibrated for the 70-130 range [0]. It's
         | used to provide broad ranges (e.g. 85-100 vs. 115-130), not to
         | provide useful differentiations on the high end like 158 vs.
         | 159.
         | 
         | In fact, the error increases with the score [1]. There is less
         | certainty for a score of 145 than there is for 115. It's just
         | not designed for separating geniuses from super geniuses or
         | however you want to put it.
         | 
         | [0] https://www.quora.com/Whats-the-highest-IQ-score-you-can-
         | get...
         | 
         | (Not a direct source I'm afraid, but the materials are
         | proprietary.)
         | 
         | [1] https://en.wikipedia.org/wiki/IQ_classification#Giftedness
         | 
         | (There are several citations in this section, pertinent
         | discussion below the table.)
        
         | ben_w wrote:
         | Anything outside 70-130 is suspect in practice, as the tests
         | aren't well calibrated beyond +-2s.
         | 
         | I'm not sure what intellectual performance is even _supposed_
         | to be implied by any given IQ score, given that the scores
         | commonly used today are just a linear mapping from standard
         | deviations above or below mean.
         | 
         | That said, on the low end at -3s I would expect a serious
         | deviation from a normal distribution just from the effect of
         | dementia.
        
           | qsort wrote:
           | Sure, but the point is that 180 wouldn't make sense even if
           | the test were perfect. There aren't enough humans on earth
           | for that many standard deviations to exist in the first
           | place. 180 IQ could as well be 2000 IQ, the sky is the limit
           | when the scale makes no sense.
        
             | computerphage wrote:
             | That's just wrong? 6 sigma would be IQ 100 + 6 * 15 = 190.
             | 6 sigma is 1 in 506,797,346. So there should be 16 people
             | with IQs that extreme if intelligence is normally
             | distributed.
             | 
             | (There's a factor of two somewhere in here for the two
             | tailed nature, but the point stands)
        
             | delecti wrote:
             | Either I'm misunderstanding the math, or you are. An IQ of
             | 180 is 5 1/3 SD above mean, which is a bit more frequent
             | than 1 in 26,330,254 (which would be 5.5 SD). Half that
             | common because we're only talking about _above_ , and 1 in
             | 50m people on earth is still 160 people. That's not many,
             | but I don't think it's negligible either.
             | 
             | https://en.wikipedia.org/wiki/68%E2%80%9395%E2%80%9399.7_ru
             | l...
        
       ___________________________________________________________________
       (page generated 2023-04-09 23:02 UTC)