Review and intuition why we divide by n-1 for the unbiased sample | Khan Academy



Courses on Khan Academy are always 100% free. Start practicing—and saving your progress—now:

Reviewing the population mean, sample mean, population variance, sample variance and building an intuition for why we divide by n-1 for the unbiased sample variance

Practice this lesson yourself on KhanAcademy.org right now:

Watch the next lesson:

Missed the previous lesson?

Probability and statistics on Khan Academy: We dare you to go through a day in which you never consider or use probability. Did you check the weather forecast? Busted! Did you decide to go through the drive through lane vs walk in? Busted again! We are constantly creating hypotheses, making predictions, testing, and analyzing. Our lives are full of probabilities! Statistics is related to probability because much of the data we use when determining probable outcomes comes from our understanding of statistics. In these tutorials, we will cover a range of topics, some which include: independent events, dependent probability, combinatorics, hypothesis testing, descriptive statistics, random variables, probability distributions, regression, and inferential statistics. So buckle up and hop on for a wild ride. We bet you’re going to be challenged AND love it!

About Khan Academy: Khan Academy offers practice exercises, instructional videos, and a personalized learning dashboard that empower learners to study at their own pace in and outside of the classroom. We tackle math, science, computer programming, history, art history, economics, and more. Our math missions guide learners from kindergarten to calculus using state-of-the-art, adaptive technology that identifies strengths and learning gaps. We’ve also partnered with institutions like NASA, The Museum of Modern Art, The California Academy of Sciences, and MIT to offer specialized content.

For free. For everyone. Forever. #YouCanLearnAnything

Subscribe to KhanAcademy’s Probability and Statistics channel:

Subscribe to KhanAcademy:

source

49 Comments

  1. The proof from 8:11 onwards does not seem rigorous. Correct me if I'm wrong.
    The sum of squares of the distances the 3 samples have to pop. mean, divided by 3, could potentially be an over-estimation of the pop. variance. Because you could have picked 3 samples that are all far from the pop. mean, whereas in the population there are points much closer to the pop. mean. 
    Which means: even the sum of the squares of the distances the 3 samples have to the sample mean, divided by 3, is always smaller than the sum of squares of the distances the 3 samples have to pop. mean, divided by 3, it is still not enough to show that the sum of the squares of the distances the 3 samples have to the sample mean, divided by 3, is always smaller than the pop. variance.

  2. Thank you very much for your video, it was very very good at explaining. But I have one more question, If descriptive statistics do not try to generalize to a population (since there is no uncertainty in descriptive statistics), then why does the sample standard deviation try to best estimate the population mean? Yet it is still considered a descriptive statistic

  3. So instead of the sample lying somewhere much lower than the true population mean, what if it's lying much higher? Would it be correct to use n+1 instead of n-1 in order to deliberately make the sample variance smaller?

  4. After 3 videos, I finally understood this n-1. Basically when we consider a sample from our population and calculate the mean for it, it may or may not be as close to the overall population mean (which is thr mean that matters) so to lower the possibility of a highly distinct sample mean/variance we use n-1 to reach at least near the population mean…

  5. A very interesting and important discussion. I made a break in the middle and thought about it by myself. I have a rather short explanation: If the sample size n is very small, such as 3, the variance calculated for the sample has more chance to be very different from the actual variance. The smaller the n is, the more effect has this '-1' on the result.
    Why do we use '-1' and not some other values like '-2', I think it is just a tradition. For the smallest sample size of 2, this unbiased variance can still be calculated. However, it is not really purely 'unbiased', just relatively 'unbiased'.

  6. I get the math…. What I don't get is how you're able to write with the drawing/annotation feature so freakin' nicely?!?!? Either you missed your calling as a steady-handed microsurgeon or there is some sort of stabilization assistance with the program you're using.

  7. 8:40 I think you should not represent the true variance and the sample variance on the same number line you drew for the population points. Also the consequence of your putting them together is you're visualizing the distance between the sample variance and the population variance on the same number line, resulting in your conclusion that because the sample points are far from the population mean, the variance is far too. Ponder over it, you'll realize.
    Love your lectures BTW 😃

  8. Let's say a report comes out that mentions standard deviation. How are we supposed to know which formula was used to calculate that standard deviation.

  9. The analogy you’re using is probably not very convincing/intuitive enough. Because there’s also a likelihood that the sample is over-estimating the population mean, so why don’t we divide it by n+1?

  10. 9:08 – You are just as likely to be overestimating, you just chose to pick the bottom points rather than the top ones. This offers literally NO explanation, let alone an intuitive one, as to why I should expect there to be a downwards bias.

  11. So this means that the n-1 of the sample variance equation was just an arbitrarily chosen value because it's empirically closer to the actual population variance? Or is there any equation or a logical path in deriving the n-1? I kinda see that it's the former but kinda feel that there might be a theory that could explain why n-1 is the most appropriate and not any other value and that it's just a natural consequence of our math. Anyone who does have one, please tell me!
    Thank you for the video Khan Academy! It was very informative!

  12. I can't understand why we would underestimate variance in general this way. Let's take population [0, 10, 20] and its sample [0, 20]. They have the same mean 10, and variance of the population is (100 + 100 + 0) / 3, while variance of the sample is (100 + 100) / 2, so we overestimate the variance.

  13. still dont get it. yes you would be underestimating it if u take the sample cluster below the mean. but if the cluster is above the mean? you would be overestimating it! seems arbitrary to me.

Leave a Reply

Your email address will not be published. Required fields are marked *

You might like

© 2026 Cantinho do Vídeo - WordPress Video Theme by WPEnjoy