Inference for Quantitative Data: Means
Unit 4 of AP Statistics, worth 10–20% of the exam. 13 questions below, each with the working. Every answer was checked by a second pass before it was published.
t-distributions, confidence intervals and tests for one mean, paired data and two means, conditions and interpretation.
How this unit is tested
What you have to know
13 practice questions
-
A random sample of 25 college students has a mean study time of 15.2 hours/week with sample standard deviation 3.4 hours. Construct a 95% confidence interval for the true mean study time (t* = 2.064, df = 24).
Show the answer
Answer. (13.80, 16.60) hours
SE = 3.4/√25 = 0.68. Margin of error = 2.064 × 0.68 ≈ 1.40. CI = 15.2 ± 1.40, giving approximately (13.80, 16.60) hours. -
Interpret the confidence interval (13.80, 16.60) hours from the study-time problem.
Show the answer
Answer. We are 95% confident that the true mean weekly study time of all college students (in the population sampled) is between 13.80 and 16.60 hours.
A correct CI interpretation always refers to the population parameter, stated in context, not to individual students or to a probability about this specific interval. -
A cereal box claims a mean weight of 16 oz. A quality control manager weighs a random sample of 20 boxes and finds mean 15.7 oz, s = 0.6 oz. Test at α = 0.05 whether the true mean weight differs from 16 oz. Give the test statistic and conclusion (df = 19, two-tailed critical t ≈ 2.093).
Show the answer
Answer. t ≈ −2.24, p ≈ 0.037 < 0.05, so reject H0: there is evidence the true mean weight differs from 16 oz.
SE = 0.6/√20 ≈ 0.1342. t = (15.7−16)/0.1342 ≈ −2.236. Since |t| = 2.236 > 2.093 (or p ≈ 0.037 < 0.05), reject the null hypothesis. -
List the three conditions that must be checked before running a one-sample t confidence interval or test for a mean.
Show the answer
Answer. Random, Independent (10% condition), and Normal/Large Sample.
Random ensures the sample represents the population; Independence (via the 10% condition when sampling without replacement) ensures observations don't affect each other; Normal/Large Sample ensures the sampling distribution of the mean is approximately normal, either because the population is normal or because n is large enough (CLT) with no strong skew or outliers. -
A physical therapist measures the flexibility (in degrees) of the same 18 patients before and after a 6-week stretching program to see if flexibility improved. What inference procedure should she use?
Show the answer
Answer. A paired t-test (one-sample t-test applied to the before/after differences).
Because the same 18 patients are measured twice, the two sets of measurements are dependent; the correct approach is to compute each patient's difference and run a one-sample t-procedure on those differences. -
For the 18 patients above, the mean of the after-minus-before differences is 5.2 degrees with standard deviation 6.4 degrees. Test H0: μ_diff = 0 vs Ha: μ_diff > 0 at α = 0.01 (df = 17, one-tail critical t ≈ 2.567).
Show the answer
Answer. t ≈ 3.45 > 2.567, so reject H0: there is evidence the program increases flexibility.
SE = 6.4/√18 ≈ 1.508. t = 5.2/1.508 ≈ 3.449. Since 3.449 exceeds the critical value 2.567, reject the null hypothesis in favor of the one-sided alternative. -
Independent random samples of runners give: Group A (n1=16): mean 52.3 min, s1=4.1 min; Group B (n2=18): mean 55.0 min, s2=5.0 min. Test at α = 0.05 whether the true mean finish times differ, using conservative df = 15 (two-tailed critical t ≈ 2.131).
Show the answer
Answer. t ≈ −1.73, |t| < 2.131, so fail to reject H0: insufficient evidence of a difference in mean finish times.
SE = √(4.1²/16 + 5²/18) = √(1.051+1.389) ≈ 1.562. t = (52.3−55.0)/1.562 ≈ −1.728, which does not exceed the critical value, so we cannot conclude the means differ. -
Using the runner data above (Group A mean 52.3, Group B mean 55.0, SE ≈ 1.562), construct a 90% confidence interval for the difference in means (A − B), using t* = 1.753 (df=15).
Show the answer
Answer. Approximately (−5.44, 0.04) minutes.
Margin of error = 1.753 × 1.562 ≈ 2.738. CI = (52.3−55.0) ± 2.738 = −2.7 ± 2.738, giving roughly (−5.44, 0.04) minutes. -
Why do statistical software packages typically report a non-integer 'Welch' degrees of freedom for a two-sample t-procedure instead of using the conservative df = min(n1−1, n2−1)?
Show the answer
Answer. The Welch formula usually gives a larger, more accurate df, producing a narrower, more precise confidence interval and a more powerful test; the conservative df underestimates precision.
The conservative approach is a safe shortcut for hand calculations that guarantees a valid (though wider) interval or a valid (though less powerful) test, while technology's exact formula better reflects the true sampling variability. -
A student says, 'There is a 95% probability that the true mean lies in this particular confidence interval.' What is wrong with this statement, and what is the correct interpretation?
Show the answer
Answer. The parameter is fixed, not random, so no probability applies to one specific already-computed interval; correctly, about 95% of intervals produced by repeated random sampling using this method would capture the true mean.
Confidence level describes the long-run success rate of the procedure across many samples, not the chance that this one interval, once calculated, happens to contain the parameter. -
In the cereal box test (t ≈ −2.24, p ≈ 0.037), interpret the p-value in context.
Show the answer
Answer. Assuming the true mean weight is really 16 oz, there is about a 3.7% chance of getting a sample mean at least as far from 16 oz as 15.7 oz (in either direction) just by random sampling variation from 20 boxes.
A p-value is a conditional probability: the chance of the observed (or more extreme) result given that the null hypothesis is true, not the probability that the null hypothesis itself is true. -
A sample of size n = 8 shows moderate skewness on a dotplot with no outliers. Should you proceed with a one-sample t-procedure? Explain using the standard sample-size guidelines.
Show the answer
Answer. Proceed with caution; for very small samples (roughly n < 15), t-procedures require the data to look close to normal, and moderate skew at n = 8 is a real concern, so the normal/large-sample condition is not clearly satisfied.
Guidelines: for n < 15, use t only if data appear close to normal; for 15 ≤ n < 30, t is fine unless there is strong skewness or outliers; for n ≥ 30, t is generally fine unless skew/outliers are extreme. At n=8 with moderate skew, the condition is borderline, so any conclusion should be stated with caution or additional data should be collected. -
A study compares average commute times of 30 randomly selected people who drive to work with 30 different randomly selected people who take public transit. Which procedure applies?
Show the answer
Answer. A two-sample t-test/interval for independent means.
The two groups consist of different, unrelated individuals, so the samples are independent rather than paired; this calls for the two-sample t procedure comparing μ1 and μ2.
What people get wrong
- Using z* instead of t* whenever σ is unknown (which is almost always). Instead, always use the t-distribution with the correct degrees of freedom when estimating the standard deviation from the sample.
- Treating two independent samples as if they should use a pooled-variance formula. AP Statistics convention is to always use the unpooled two-sample t procedure for means, not the pooled t-test taught in some other courses.
- Confusing paired and independent designs — using the two-sample formula on data that are actually paired (same subjects measured twice) inflates the standard error and gives wrong conclusions. Instead, recognize paired designs and reduce to differences first.
- Interpreting a confidence interval as 'there is a 95% probability the true mean is in this specific interval' or '95% of the data lies in this interval.' Instead, state that in repeated sampling, about 95% of intervals constructed this way would capture the true population mean.
- Skipping the explicit check of conditions (jumping straight to computation) or claiming normality is satisfied just because n is large without checking for skew or outliers in small-to-moderate samples.
- Forgetting that for paired data, the independence condition applies to the pairs themselves (different subjects), not to the two measurements within a pair, which are expected to be dependent.
Drill this unit until it sticks
These questions come back on a schedule built from what you get wrong, alongside the rest of AP Statistics. Free, and no account needed to start.