Confidence Intervals, Type I and II Errors, p-value, t-distribution
One of the major objectives of statistics is to make inferences about the distribution of the elements in a population based on information contained in a sample.
Numerical summaries that characterize the population distribution are called parameters.
The population mean \(\mu\) and population variance \(\sigma^2\) are two important parameters. Others are median, range, mode, etc.
Methods for making inferences are basically designed to answer one of two types of questions:
Statisticians answer the first question by estimating the parameter using the sample. The second case might require a test of a hypothesis.
Point estimation of \(\mu\) does not require that \(\sigma\) be known or a large \(n\)
Point estimates by themselves do not tell how much \(\bar{y}\) might differ from \(\mu\), that is, the accuracy or precision of the estimate
A measure of accuracy is the difference between sample mean \(\bar{y}\) and population mean \(\mu\) is called sampling error
An interval estimate, called a confidence interval, incorporates information about the amount of sampling error in \(\bar{y}\)
A confidence interval for \(\mu\) takes the form \[(\bar{y} - E,\, \bar{y} + E),\] for a number \(E\)
An associated number called the confidence coefficient helps assess how likely it is for \(\mu\) to be in the interval.
To derive the specific form of the confidence interval for \(\mu\), for the case when \(\sigma\) is known and \(n\) is large, the CLT result must be used.
By the CLT, \(Z = (\bar{Y} - \mu)/\sigma_{\bar{y}}\) is approximately \(\sim N(0, 1)\).
Let \(Z\) have a \(N(0, 1)\) distribution (exactly).
Let \(z_{\alpha/2}\) denote the \(1 - \alpha/2\) quantile of the standard normal distribution, for a given number \(\alpha\), \(0 < \alpha < 1\).
Then the following probability statement is true: \[P\left (-z_{\alpha/2}\, \leq\, \frac{(\bar{Y} - \mu)}{ \sigma_{\bar{y}}}\, \leq\, z_{\alpha/2}\right ) = 1 - \alpha\]
Manipulating the inequalities, without changing values, we have
\[P\left (-z_{\alpha/2}\, \leq\, \frac{(\bar{Y} - \mu)}{ \sigma_{\bar{y}}}\, \leq\, z_{\alpha/2}\right ) = 1 - \alpha\]
Using this statement, a (1-\(\alpha\))100% confidence interval for \(\mu\), with confidence coefficient \(1-\alpha\) can be calculated using a random sample of data \(y_1,y_2,\ldots,y_n\)
It is usually written in the form \[\left( \bar{y} - \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2}, \ \bar{y} + \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2}\right)\]
Example: Suppose \(n\) = 36, \(\sigma=12\), \(\bar{y} = 24.8\)
And for \(\alpha=0.05\), \(z_{\alpha/2} \equiv z_{0.025} = 1.96\)
Thus a 95% C.I. for \(\mu\) is \[\left( 24.8 - \frac{12}{\sqrt{36}} \times 1.96, \quad 24.8 + \frac{12}{\sqrt{36}} \times 1.96 \right)\]
The interval for \(\mu\) is: (20.88, 28.72) with confidence coefficient 0.95.
We might say that we are 95% confident that the population mean \(\mu\) is between 20.88 and 28.72.
Reminder: \(z_{\alpha/2}\) is the value of \(z\) at which the \(P(Z>z)=\alpha/2\), so we need to flip the direction for qnorm()
Before the sample is drawn, the probability is (\(1-\alpha\)) that the random interval \((\bar{Y} - \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2}\), \(\bar{Y} + \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2}\)) will contain \(\mu\) (because \(\bar{Y}\) is a random variable).
However, once the sample is drawn, and \(\bar{y}\) is calculated, the interval ceases to be random. It is a numerical interval \((\bar{y} - \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2}\), \(\bar{y} + \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2}\)), calculated specifically for the drawn sample; thus we cannot associate a probability with it.
We may say that the process which led us to this interval will, on the average, produce an interval containing \(\mu,\ 100 (1-\alpha)\)% of the time.
It is not correct to say that the interval \((\bar{y} - \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2},\, \bar{y} + \frac{\sigma}{\sqrt{n}} \, z_{\alpha/2})\) contains \(\mu\) with a specified probability.
If the process of sampling is repeated many times and 100(1-\(\alpha\))% confidence intervals calculated for each sample, then we are confident that 100(1-\(\alpha\))% of those intervals will contain \(\mu\).
As the confidence coefficient increases the confidence interval becomes wider. A wider interval estimates \(\mu\) less precisely. Thus a 99% confidence interval is less accurate than a 95% confidence interval but one has more confidence that \(\mu\) is in the first interval.
alpha confidence lcl ybar ucl
1 0.10 0.90 37.38862 40 42.61138
2 0.05 0.95 36.88834 40 43.11166
3 0.01 0.99 35.91059 40 44.08941
Also note that increasing the sample size \(n\) results in a narrower interval (more accurate estimate) for the same confidence coefficient.
The caffeine content (in mg) was examined for a random sample of 50 cups of black coffee dispensed by a new machine. The mean and standard deviation were 110 mg and 7.1 mg, respectively. Use these data to construct a 98% confidence interval for \(\mu\), the mean caffeine content for cups dispensed by the machine.
As the sample size is large we can use the CLT and also approximate \(\sigma\), the sample standard deviation of the population by \(s=7.1\)
\(\left[\bar{y} - z_{0.01} \frac{\sigma}{\sqrt{n}}, \
\bar{y} + z_{0.01} \frac{\sigma}{\sqrt{n}}\right]\)
\(\left[110 - (2.33)\left(\frac{7.1}{\sqrt{50}}\right), \
110 + (2.33)\left(\frac{7.1}{\sqrt{50}}\right) \right]\)
\((107.66,\ 112.34)\)
We are \(98\%\) confident that the mean caffeine content for cups dispensed by the machine is between 107.66 and 112.34 mg.
qnorm() measures \(P(Z\leq z)\) when \(Z\sim N(0,1)\). \(z_{\alpha/2}\) refers to the the value of \(z\) that produces \(P(Z\gt z)=\alpha/2\).Estimation (either point or interval estimation) was used to help answer a question like “Approximately what is the value of \(\mu\)?”
The other kind of question mentioned earlier is “Is it likely that \(\mu\) is less than (or greater than) the value \(\mu_0\) (a pre-determined value)?”.
Every test of hypothesis features
a Null Hypothesis \(H_0\) which describes a characteristic of the population (it is believed to be) as it currently exists,
a Research Hypothesis (or Alternative Hypothesis) \(H_a\), which is a proposal about this characteristic, by the person(s) conducting the statistical study.
The idea is that the null hypothesis is presumed to hold unless there is overwhelming evidence in the data in support of the research hypothesis.
To determine whether the mean yield per acre (in bushels), \(\mu\), of a variety of soybeans increased the current year over the last two years when \(\mu\) is believed to be 520 bushels per acre, the following might be tested. \[H_0: \mu \leq 520 \ \ \ \mbox{vs.}\ \ \ H_a: \mu > 520\]
If \(\bar{y}\) is in the rejection region then reject \(H_0\) and say the evidence favors \(H_a\).
For example, in the above example if \(\bar{y} < 520\) we will say there is not sufficient evidence to reject \(H_0\).
Even if \(\bar{y} > 520\), we might still say there is not sufficient evidence to reject \(H_0\)
This is ok as long as \(\bar{y}\) is not too much greater than \(520\).
“How much is too large?” To decide this, use the probability distribution of \(\bar{Y}\). Begin by picking a small probability \(\alpha\) like \(\alpha\) = 0.05, or \(\alpha = 0.01\), or \(\alpha =0.001\).
Then reject \(H_0\) only when the probability of obtaining a value of \(\bar{y}\) larger than the observed value \(\bar{y}\) is \(\leq \alpha\) when \(\mu\) is indeed \(\leq 520\), i.e., when \(H_0\) is true.
That is, if \(H_0\) is true, the chance of observing a sample that results in a \(\bar{y}\) as large as the one calculated should be very small.
By the CLT, \(\bar{Y}\) is approximately a normal random variable.
We will ignore the approximation and assume \(\bar{Y} \sim N(\mu, \frac{\sigma^2}{n})\).
Then, for a given \(\alpha\), \[P \left( \frac{\bar{Y}-\mu}{\sigma/\sqrt{n}} > z_{\alpha} \right) = \alpha\] where \(z_\alpha\) is the \(1-\alpha\) quantile of the standard normal distribution.
So when \(H_0: \mu \leq \mu_0\) is true , \[P \left( \frac{\bar{Y}-\mu_0}{\sigma/\sqrt{n}} > z_{\alpha} \right) \leq \alpha\]
That is, \(P(\bar{Y} > \mu_0 + z_\alpha \frac{\sigma}{\sqrt{n}})\leq \alpha.\) This suggests that we reject \(H_0: \mu \leq \mu_0\) in favor of, say \(H_a: \mu > \mu_0\) when \[\bar{y} > \mu_0 + z_\alpha \frac{\sigma}{\sqrt{n}}\] (i.e., \(\bar{y}\) is too much when \(\bar{y}\) exceeds the number \(\mu_0 +
z_\alpha \frac{\sigma}{\sqrt{n}})\)
All possible \(\bar{y}\) values satisfying \(\bar{y} > \mu_0 + z_\alpha \frac{\sigma}{\sqrt{n}}\) constitute the rejection region.
Instead of comparing \(\bar{y}\) to \(\mu_0 + z_\alpha \frac{\sigma}{\sqrt{n}}\), it is easier to first calculate \[z_c=\frac{\bar{y}-\mu_0}{\sigma/\sqrt{n}}\]
We see that comparing \(\bar{y}\) to \(\mu_0 + z_\alpha \frac{\sigma}{\sqrt{n}}\) is the same as comparing \(z_c\) to \(z_\alpha\).
That is, we reject the null hypothesis if \(z_c > z_\alpha\) whis is the same thing as doing so if \(\bar{y} > \mu_0 + z_\alpha \frac{\sigma}{\sqrt{n}}\)
This quantity \(z\) is called the test statistic and it’s value can be calculated using the data and the value \(\mu_0\) specified in the null hypothesis \(H_0\).
The notation \(z_c\) is used denote the “computed value” of the test statistic, that is when we plug in the observed value of \(\bar{y}\) and obtain a numerical value for \(z\).
The \(\alpha\) corresponds to the probability that the null hypothesis is rejected when actually it is true and is called the probability of committing a Type I error.
Since the experimenter selects the value of \(\alpha\) used in the test procedure, she is able to specify or control the Type I error rate or “how much” of this type of error is permitted in the testing procedure.
Suppose from a sample of 36 1-acre plots, the yield of corn this year was measured and \(\bar{y}=573\) and \(s=124\) calculated. Can we conclude that the mean yield of corn for all farms exceeded 520 bushels/acre this year ? Here we are going to assume that \(\sigma\) can be approximated by \(s\). Use \(\alpha=.025\)
Set-up five parts of the testing procedure:
\(H_0: \mu \leq 520\)
\(H_a: \mu > 520\)
T.S.: \(z = \frac{\bar{y}-\mu_0}{\sigma/\sqrt{n}}\)
R.R: Reject \(H_0:\) in favor of \(H_a:\) if \(z > z_{.025}\)
i.e. the R.R. is \(z>1.96\) since \(z_{.025}=1.96\)
Compute the observed value of the test statistic:
\(z_c =\frac{573-520}{124/{\sqrt{36}}}= 2.56\)
Since \(z_c> 1.96\), \(z_c\) is in the R.R. Thus \(H_0:\mu \leq 520\) is rejected in favor of the research hypothesis \(H_a: \mu > 520\). It is concluded that the mean yield this year exceeds 520 bushels/acre.
Hypotheses:
Case 1: \(H_0: \mu \leq \mu_0\ \mbox{vs.}\ H_a: \mu >\mu_0\) (right-tailed test)
Case 2: \(H_0: \mu \geq \mu_0\ \mbox{vs.}\ H_a: \mu <\mu_0\) (left-tailed test)
Case 3: \(H_0: \mu = \mu_0\ \mbox{vs.}\ H_a: \mu \not =\mu_0\) (two-tailed test)
T.S: \(z_c = \frac{\bar{y}-\mu_0}{\sigma/\sqrt{n}}\)
R.R: For Type I error probability of \(\alpha\):
Case 1: Reject \(H_0:\) if \(z \geq z_{\alpha}\)
Case 2: Reject \(H_0:\) if \(z \leq -z_{\alpha}\)
Case 3: Reject \(H_0:\) if \(|z| \geq z_{\alpha/2}\)
If the computed value of the test statistic, \(z_c\), is in the R.R., then we will reject the null hypothesis \(H_0:\) at the specified \(\alpha\) value. Otherwise, we say we fail to reject \(H_0:\) at the specified \(\alpha\) value.
Example A corporation maintains a large fleet of company cars for its salespeople. To check the average number of miles driven per month per car, a random sample of \(n = 40\) cars is examined. The mean and standard deviation for the sample are 2,752 miles and 350 miles, respectively. Records for previous years indicate that the average number of miles driven per car per month was 2,600. Use the sample data to test the research hypothesis that the current mean \(\mu\) differs from 2,600. Set \(\alpha\) = .05 and assume that \(\sigma\) can be replaced by \(s\).
The null hypothesis for this statistical test is \(H_0: \mu\) = 2,600 and the research hypothesis is \(H_a: \mu \neq\) 2,600. Using \(\alpha\) = .05, the two-tailed rejection region for this test is \(|z|> z_{.025}\) or \(|z|> 1.96\) and is located as shown below.
To determine how many standard errors our test statistics \(\bar{y}\) lies away from \(\mu\) = 2,600, compute \[z_c = \frac{\bar{y}-\mu_0}{\sigma/\sqrt{n}} =
\frac{2,752 - 2,600}{350/\sqrt{40}} = 2.75 \, .\] Thus \(|z_c|=2.75>1.96\) and therefore we reject \(H_0: \mu = 2,600\) at \(\alpha=.05\). It follows that the observed value for \(\bar{y}\) lies more than 1.96 standard errors above the mean \(\mu = 2,600\), so we reject the null hypothesis in favor of the alternative \(H_a: \mu \neq
2,600\).
Since \(\bar{y}>2,600\) we conclude that the mean number of miles driven is greater than \(2,600\).
As an alternative to the formal test where one uses a rejection region based on a specified the Type I error rate \(\alpha\), many researchers compute and report the level of significance or the p-value for the test.
This is the probability, when \(H_0\) is true, of observing a statistic as extreme as the one actually observed. Here “extreme” is means “large” or “small” according to the alternative hypothesis \(H_a\).
Specifically, for a given \(\sigma^2\) and \(\mu_0\) we find the p-value by:
First computing the test statistic, \(z_c\), \[z_c = (\bar{y} - \mu_0)/(\sigma/\sqrt{n})\]
A p-value smaller than the pre-specified \(\alpha\) value is evidence in favor of rejecting \(H_0\).
In the earlier example we tested \(H_0: \mu \leq 380\ \mbox{vs.}\ H_a: \mu > 380\) Calculating the \(z-\)statistic, \[z_c = \frac{\bar{y}-380}{\sigma/\sqrt{n}}= \frac{390-380}{35.2/\sqrt{50}} = 2.01\]
The level of significance, or p-value for this test is \[p= P(Z > 2.01)= 1- P(Z < 2.01)= 0.0222\]
We fail to reject \(H_0: \mu \leq 380\) at \(\alpha=.01\) since p-value is not less than \(.01\)
STAT 801A