Chapter page 33 / 3824 Inference for linear regression with a single predictor
English

24  Inference for linear regression with a single predictor

We now bring together ideas of inferential analyses with the descriptive models seen in Chapter 7. In particular, we will use the least squares regression line to test whether there is a relationship between two continuous variables. Additionally, we will build confidence intervals which quantify the slope of the linear regression line. The setting is now focused on predicting a numeric response variable (for linear models) or a binary response variable (for logistic models), we continue to ask questions about the variability of the model from sample to sample. The sampling variability will inform the conclusions about the population that can be drawn.

Many of the inferential ideas are remarkably similar to those covered in previous chapters. The technical conditions for linear models are typically assessed graphically, although independence of observations continues to be of utmost importance.

We encourage the reader to think broadly about the models at hand without putting too much dependence on the exact p-values that are reported from the statistical software. Inference on models with multiple explanatory variables can suffer from data snooping which results in false positive claims. We provide some guidance and hope the reader will further their statistical learning after working through the material in this text.

24.1 Case study: Sandwich store

24.1.1 Observed data

We start the chapter with a hypothetical example describing the linear relationship between dollars spent advertising for a chain sandwich restaurant and monthly revenue. The hypothetical example serves the purpose of illustrating how a linear model varies from sample to sample. Because we have made up the example and the data (and the entire population), we can take many many samples from the population to visualize the variability. Note that in real life, we always have exactly one sample (that is, one dataset), and through the inference process, we imagine what might have happened had we taken a different sample. The change from sample to sample leads to an understanding of how the single observed dataset is different from the population of values, which is typically the fundamental goal of inference.

Consider the following hypothetical population of all of the sandwich stores of a particular chain seen in Figure 24.1. In this made-up world, the CEO actually has all the relevant data, which is why they can plot it here. The CEO is omniscient and can write down the population model which describes the true population relationship between the advertising dollars and revenue. There appears to be a linear relationship between advertising dollars and revenue (both in $1,000).

Scatterplot with advertising amount on the x-axis and revenue on the y-axis. A linear model is superimposed. The points show a reasonably strong and positive linear trend.
Figure 24.1: Revenue as a linear model of advertising dollars for a population of sandwich stores, in thousands of dollars.

You may remember from Chapter 7 that the population model is: \[y = \beta_0 + \beta_1 x + \varepsilon.\]

Again, the omniscient CEO (with the full population information) can write down the true population model as: \[\texttt{expected revenue} = 11.23 + 4.8 \times \texttt{advertising}.\]

24.1.2 Variability of the statistic

Unfortunately, in our scenario, the CEO is not willing to part with the full set of data, but they will allow potential franchise buyers to see a small sample of the data in order to help the potential buyer decide whether set up a new franchise. The CEO is willing to give each potential franchise buyer a random sample of data from 20 stores.

As with any numerical characteristic which describes a subset of the population, the estimated slope of a sample will vary from sample to sample. Consider the linear model which describes revenue (in $1,000) based on advertising dollars (in $1,000).

The least squares regression model uses the data to find a sample linear fit: \[\hat{y} = b_0 + b_1 x.\]

Two random samples of 20 stores shows different least squares regression lines in Figure 24.2 (a) and Figure 24.2 (b), depending on which observations are selected. Both trends are similar to those seen in Figure 24.1, which describes the population.

For two random samples of 20 stores, scatterplots with advertising amount on the x-axis and revenue on the y-axis. Linear models are superimposed. The points in each plot show reasonably strong and positive linear trends.
(a) First sample.
For two random samples of 20 stores, scatterplots with advertising amount on the x-axis and revenue on the y-axis. Linear models are superimposed. The points in each plot show reasonably strong and positive linear trends.
(b) Second sample.
Figure 24.2: Two random samples of 20 stores from the entire population. A linear trend between advertising and revenue is observed in both.

Figure 24.3 shows the two samples and the least squares regressions from Figure 24.2 on the same plot. We can see that the two lines are different. That is, there is variability in the regression line from sample to sample. The concept of the sampling variability is something you’ve seen before, but in this lesson, you will focus on the variability of the line often measured through the variability of a single statistic: the slope of the line.

For two different random samples, superimposed onto the same plot, scatterplot with advertising amount on the x-axis and revenue on the y-axis. Two linear models are plotted to demonstrate that the lines are very similar, yet they are not the same.
Figure 24.3: The linear models from the two random samples are quite similar, but not exactly the same.

Figure 24.4 shows least squares lines fit to many more random samples of 20 from the population.

An x-y coordinate system with least squares regression lines from many random samples of size 20 (no points are plotted). The lines vary around the true population line. On the x-axis is advertising amount; on the y-axis is revenue.
Figure 24.4: If repeated samples of size 20 are taken from the entire population, each linear model will be slightly different. The red line provides the linear fit to the entire population.

You might notice in Figure 24.4 that the \(\hat{y}\) values given by the lines are much more consistent in the middle of the dataset than at the ends. The reason is that the data itself anchors the lines in such a way that the line must pass through the center of the data cloud. The effect of the fan-shaped lines is that predicted revenue for advertising close to $4,000 will be much more precise than the revenue predictions made for $1,000 or $7,000 of advertising.

The distribution of slopes (for samples of size \(n=20\)) can be seen in Figure 24.5.

Histogram of the slope values from many random samples of size 20. The slope estimates vary from about 2.1 to 8. The histogram is reasonably bell-shaped.
Figure 24.5: Variability of slope estimates from many different samples of stores, each of size 20.

Recall, the example described in this introduction is hypothetical. That is, we created an entire population in order to demonstrate how the slope of a line would vary from sample to sample. The tools in this textbook are designed to evaluate only one single sample of data. With actual studies, we do not have repeated samples, so we are not able to use repeated samples to visualize the variability in slopes. We have seen variability in samples throughout this text, so it should not come as a surprise that different samples will produce different linear models. However, it is nice to visually consider the linear models produced by different slopes. Additionally, as with measuring the variability of previous statistics (e.g., \(\overline{X}_1 - \overline{X}_2\) or \(\hat{p}_1 - \hat{p}_2\)), the histogram of the sample statistics can provide information related to inferential considerations.

In the following sections, the distribution (i.e., histogram) of \(b_1\) (the estimated slope coefficient) will be constructed in the same three ways that, by now, may be familiar to you. First (in Section 24.2), the distribution of \(b_1\) when \(\beta_1 = 0\) is constructed by randomizing (permuting) the response variable. Next (in Section 24.3), we can bootstrap the data by taking random samples of size \(n\) from the original dataset. And last (in Section 24.4), we use mathematical tools to describe the variability using the \(t\)-distribution that was first encountered in Section 19.2.

24.2 Randomization test for the slope

Consider data on 100 randomly selected births gathered originally from the US Department of Health and Human Services. Some of the variables are plotted in Figure 24.6.

The scientific research interest at hand will be in determining the linear relationship between weight of baby at birth (in lbs) and number of weeks of gestation. The dataset is quite rich and deserves exploring, but for this example, we will focus only on the weight of the baby.

The births14 data can be found in the openintro R package. We will work with a random sample of 100 observations from these data.

Four different scatterplots, all with weight of baby on the y-axis. On the x-axis are weight gained by mother, mother's age, number of hospital visits, and weeks gestation. Weeks gestation and weight of baby show the strongest linear relationship (which is positive).
Figure 24.6: Weight of baby at birth (in lbs) as plotted by four other birth variables (mother’s weight gain, mother’s age, number of hospital visits, and weeks gestation).

As you have seen previously, statistical inference typically relies on setting a null hypothesis which is hoped to be subsequently rejected. In the linear model setting, we might hope to have a linear relationship between weeks and weight in settings where weeks gestation is known and weight of baby needs to be predicted.

The relevant hypotheses for the linear model setting can be written in terms of the population slope parameter. Here the population refers to a larger population of births in the US.

  • \(H_0: \beta_1= 0\), there is no linear relationship between weight and weeks.
  • \(H_A: \beta_1 \ne 0\), there is some linear relationship between weight and weeks.

Recall that for the randomization test, we permute one variable to eliminate any existing relationship between the variables. That is, we set the null hypothesis to be true, and we measure the natural variability in the data due to sampling but not due to variables being correlated. Figure 24.7 (a) shows the observed data and Figure 24.7 (b) shows one permutation of the weight variable. The careful observer can see that each of the observed values for weight (and for weeks) exist in both the original data plot as well as the permuted weight plot, but the weight and weeks gestation are no longer matched for a given birth. That is, each weight value is randomly assigned to a new weeks gestation.

By repeatedly permuting the response variable, any pattern in the linear model that is observed is due only to random chance (and not an underlying relationship). The randomization test compares the slopes calculated from the permuted response variable with the observed slope. If the observed slope is inconsistent with the slopes from permuting, we can conclude that there is some underlying relationship (and that the slope is not merely due to random chance).

Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. The left panel is the original data. The right panel is data where the weight of the baby has been permuted across the observations.
(a) Original data.
Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. The left panel is the original data. The right panel is data where the weight of the baby has been permuted across the observations.
(b) Permuted data.
Figure 24.7: Permutation removes the linear relationship between weight and weeks. Repeated permutations allow for quantifying the variability in the slope under the condition that there is no linear relationship (i.e., that the null hypothesis is true).

24.2.1 Observed data

We will continue to use the births data to investigate the linear relationship between weight and weeks gestation. Note that the least squares model (see Chapter 7) describing the relationship is given in Table 24.1. The columns in Table 24.1 are further described in Section 24.4.

Table 24.1: The least squares estimates of the intercept and slope are given in the estimate column. The observed slope is 0.335.
term estimate std.error statistic p.value
(Intercept) -5.72 1.61 -3.54 6e-04
weeks 0.34 0.04 8.07 <0.0001

24.2.2 Variability of the statistic

After permuting the data, the least squares estimate of the line can be computed. Repeated permutations and slope calculations describe the variability in the line (i.e., in the slope) due only to the natural variability and not due to a relationship between weight and weeks gestation. Figure 24.8 shows two different permutations of weight and the resulting linear models.

Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. Each plot includes data where the weight of the baby has been permuted across the observations. The two different permutations produce slightly different least squares regression lines.
(a) First permutation.
Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. Each plot includes data where the weight of the baby has been permuted across the observations. The two different permutations produce slightly different least squares regression lines.
(b) Second permutation.
Figure 24.8: Two permutations of weight with slightly different least squares regression lines.

As you can see, sometimes the slope of the permuted data is positive, sometimes it is negative. Because the randomization happens under the condition of no underlying relationship (because the response variable is completely mixed with the explanatory variable), we expect to see the center of the randomized slope distribution to be zero.

24.2.3 Observed statistic vs. null statistics

Histogram of slopes describing the linear model from permuted weight regressed on weeks gestation. The permuted slopes range from -0.15 to +0.15 and are nowhere near the observed slope value of 0.335.
Figure 24.9: Histogram of slopes given different permutations of the weight variable. The vertical red line is at the observed value of the slope, 0.335.

As we can see from Figure 24.9, a slope estimate as extreme as the observed slope estimate (the red line) never happened in many repeated permutations of the weight variable. That is, if indeed there were no linear relationship between weight and weeks, the natural variability of the slopes would produce estimates between approximately -0.15 and +0.15. We reject the null hypothesis. Therefore, we believe that the slope observed on the original data is not just due to natural variability and indeed, there is a linear relationship between weight of baby and weeks gestation for births in the US.

24.3 Bootstrap confidence interval for the slope

As we have seen in previous chapters, we can use bootstrapping to estimate the sampling distribution of the statistic of interest (here, the slope) without the null assumption of no relationship (which was the condition in the randomization test). Because interest is now in creating a CI, there is no null hypothesis, so there won’t be any reason to permute either of the variables.

24.3.1 Observed data

Returning to the births data, we may want to consider the relationship between mage (mother’s age) and weight. Is mage a good predictor of weight? And if so, what is the relationship? That is, what is the slope that models average weight of baby as a function of mage (mother’s age)? The linear model regressing weight on mage is provided in Table 24.2.

Scatterplot with mother's age on the x-axis and baby's weight on the y-axis. A linear model is superimposed. The points show a weak positive linear trend.
Figure 24.10: Using the original data, the weight of baby as a linear model of mother’s age. Notice that the relationship between mother’s age and weight of baby is not as strong as the relationship we saw previously between weeks gestation and weight of baby.
Table 24.2: The least squares estimates of the intercept and slope are given in the estimate column. The observed slope is 0.036.
term estimate std.error statistic p.value
(Intercept) 6.23 0.71 8.79 <0.0001
mage 0.04 0.02 1.50 0.1362

24.3.2 Variability of the statistic

Because the focus here is not on a null distribution, we sample with replacement \(n = 100\) observations from the original dataset. Recall that with bootstrapping the resample always has the same number of observations as the original dataset in order to mimic the process of taking a sample from the population. When sampling in the linear model case, consider each observation to be a single dot. If the dot is resampled, both the weight and the mage measurement are observed. The measurements are linked to the dot (i.e., to the birth in the sample).

Two scatterplots, both with mother's age on the x-axis and baby's weight on the y-axis. The left plot is the original data. The right plot is the bootstrapped data. Comparing the bootstrapped points to the original points, we can see that some observations were sampled more than once, and some observations were not selected for the bootstrap sample at all.
(a) Original data.
Two scatterplots, both with mother's age on the x-axis and baby's weight on the y-axis. The left plot is the original data. The right plot is the bootstrapped data. Comparing the bootstrapped points to the original points, we can see that some observations were sampled more than once, and some observations were not selected for the bootstrap sample at all.
(b) Bootstrapped data.
Figure 24.11: Original and one bootstrap sample of the births data. It is difficult to differentiate between the two plots, as (within a single bootstrap sample) the observations which have been resampled twice are plotted as points on top of one another. The red circles represent points in the original data which were not included in the bootstrap sample. The blue circles represent a data point that was repeatedly resampled (and is therefore darker) in the bootstrap sample. The green circles represent a particular structure to the data which is observed in both the original and bootstrap samples.

Figure 24.11 (a) shows the original data as compared with a single bootstrap sample in Figure 24.11 (b), resulting in (slightly) different linear models. The red circles represent points in the original data which were not included in the bootstrap sample. The blue circles represent a point that was repeatedly resampled (and is therefore darker) in the bootstrap sample. The green circles represent a particular structure to the data which is observed in both the original and bootstrap samples. By repeatedly resampling, we can see dozens of bootstrapped slopes on the same plot in Figure 24.12.

An x-y coordinate system with least squares regression lines from many bootstrap samples (no points are plotted). The lines vary around the observed population line. On the x-axis is mother's age; on the y-axis is baby's weight.
Figure 24.12: Repeated bootstrap resamples of size 100 are taken from the original data. Each of the bootstrapped linear models is slightly different.

Recall that in order to create a confidence interval for the slope, we need to find the range of values that the statistic (here the slope) takes on from different bootstrap samples. Figure 24.13 is a histogram of the relevant bootstrapped slopes. We can see that a 95% bootstrap percentile interval for the true population slope is given by (-0.01, 0.081). We are 95% confident that for the model describing the population of births, predicting weight of baby from mother’s age, a one unit increase in mage (in years) is associated with an increase in predicted average baby weight of between -0.01 and 0.081 pounds. Notice that the CI contains zero, so the true relationship might be null!

Histogram of the slopes computed from many bootstrapped samples. The bootstrap samples range from -0.05 (with the 2.5 percentile at -0.01) to +0.1 (with the 97.5 percentile at 0.081). The bootstrapped slopes form a histogram that is reasonably symmetric and bell-shaped.
Figure 24.13: The original births data on baby’s weight and mother’s age is bootstrapped 1,000 times. The histogram provides a sense for the variability of the slope of the linear model from sample to sample.

Using Figure 24.13, calculate the bootstrap estimate for the standard error of the slope. Using the bootstrap standard error, find a 95% bootstrap SE confidence interval for the true population slope, and interpret the interval in context.


Notice that most of the bootstrapped slopes fall between -0.01 and +0.08 (a range of 0.09). Using the empirical rule (that with bell-shaped distributions, most observations are within two standard errors of the center), the standard error of the slopes is approximately 0.0225. The critical value for a 95% confidence interval is \(z^\star = 1.96\) which leads to a confidence interval of \(b_1 \pm 1.96 \cdot SE \rightarrow 0.036 \pm 1.96 \cdot 0.0225 \rightarrow (-0.0081, 0.0801).\) The bootstrap SE confidence interval is almost identical to the bootstrap percentile interval. In context, we are 95% confident that for the model describing the population of births, predicting weight of baby from mother’s age, a one unit increase in mage (in years) is associated with an increase in predicted average baby weight of between -0.0081 and 0.0801 pounds.

24.4 Mathematical model for testing the slope

When certain technical conditions apply, it is convenient to use mathematical approximations to test and estimate the slope parameter. The approximations will build on the t-distribution which was described in Chapter 19. The mathematical model is often correct and is usually easy to implement computationally. The validity of the technical conditions will be considered in detail in Section 24.6.

In this section, we discuss uncertainty in the estimates of the slope and y-intercept for a regression line. Just as we identified standard errors for point estimates in previous chapters, we start by discussing standard errors for the slope and y-intercept estimates.

24.4.1 Observed data

Midterm elections and unemployment

Elections for members of the United States House of Representatives occur every two years, coinciding every four years with the U.S. Presidential election. The set of House elections occurring during the middle of a Presidential term are called midterm elections. In America’s two-party system (the vast majority of House members through history have been either Republicans or Democrats), one political theory suggests the higher the unemployment rate, the worse the President’s party will do in the midterm elections. In 2020 there were 232 Democrats, 198 Republicans, and 1 Libertarian in the House.

To assess the validity of the claim related to unemployment and voting patterns, we can compile historical data and look for a connection. We consider every midterm election from 1898 to 2018, with the exception of the elections during the Great Depression. The House of Representatives is made up of 435 voting members.

The midterms_house data can be found in the openintro R package.

Figure 24.14 shows these data and the least-squares regression line:

\[ \begin{aligned} &\texttt{percent change in House seats for President's party} \\ &\qquad\qquad= -7.36 - 0.89 \times \texttt{(unemployment rate)} \end{aligned} \]

We consider the percent change in the number of seats of the President’s party (e.g., percent change in the number of seats for Republicans in 2018) against the unemployment rate.

Examining the data, there are no clear deviations from linearity or substantial outliers (see Section 7.1.3 for a discussion on using residuals to visualize how well a linear model fits the data). While the data are collected sequentially, a separate analysis was used to check for any apparent correlation between successive observations; no such correlation was found.

Scatterplot with percent unemployed on the x-axis and percent change in House seats for the President's party on the y-axis. Each point represents a different President's midterm and is colored according to their political party (Democrat or Republican). The relationship is moderate and negative.
Figure 24.14: The percent change in House seats for the President’s party in each election from 1898 to 2010 plotted against the unemployment rate. The two points for the Great Depression have been removed, and a least squares regression line has been fit to the data.

The data for the Great Depression (1934 and 1938) were removed because the unemployment rate was 21% and 18%, respectively. Do you agree that they should be removed for this investigation? Why or why not?1

There is a negative slope in the line shown in Figure 24.14. However, this slope (and the y-intercept) are only estimates of the parameter values. We might wonder, is this convincing evidence that the “true” linear model has a negative slope? That is, do the data provide strong evidence that the political theory is accurate, where the unemployment rate is a useful predictor of the midterm election? We can frame this investigation into a statistical hypothesis test:

  • \(H_0\): \(\beta_1 = 0\). The true linear model has slope zero.
  • \(H_A\): \(\beta_1 \neq 0\). The true linear model has a slope different than zero. The unemployment is predictive of whether the President’s party wins or loses seats in the House of Representatives.

We would reject \(H_0\) in favor of \(H_A\) if the data provide strong evidence that the true slope parameter is different than zero. To assess the hypotheses, we identify a standard error for the estimate, compute an appropriate test statistic, and identify the p-value.

24.4.2 Variability of the statistic

Just like other point estimates we have seen before, we can compute a standard error and test statistic for \(b_1\). We will generally label the test statistic using a \(T\), since it follows the \(t\)-distribution.

We will rely on statistical software to compute the standard error and leave the explanation of how this standard error is determined to a second or third statistics course. Table 24.3 shows software output for the least squares regression line in Figure 24.14. The row labeled unemp includes all relevant information about the slope estimate (i.e., the coefficient of the unemployment variable, the related SE, the T statistic, and the corresponding p-value).

Table 24.3: Output from statistical software for the regression line modeling the midterm election losses for the President’s party as a response to unemployment.
term estimate std.error statistic p.value
(Intercept) -7.36 5.16 -1.43 0.16
unemp -0.89 0.83 -1.07 0.30

What do the first and second columns of Table 24.3 represent?


The entries in the first column represent the least squares estimates, \(b_0\) and \(b_1\), and the values in the second column correspond to the standard errors of each estimate. Using the estimates, we could write the equation for the least square regression line as

\[ \hat{y} = -7.36 - 0.89 x \]

where \(\hat{y}\) in this case represents the predicted change in the number of seats for the president’s party, and \(x\) represents the unemployment rate.

We previously used a \(t\)-test statistic for hypothesis testing in the context of numerical data. Regression is very similar. In the hypotheses we consider, the null value for the slope is 0, so we can compute the test statistic using the T score formula:

\[ T \ = \ \frac{\text{estimate} - \text{null value}}{\text{SE}} = \ \frac{-0.89 - 0}{0.835} = \ -1.07 \]

The T score we calculated corresponds to the third column of Table 24.3.

Use Table 24.3 to determine the p-value for the hypothesis test.


The last column of the table gives the p-value for the two-sided hypothesis test for the coefficient of the unemployment rate 0.2961. That is, the data do not provide convincing evidence that a higher unemployment rate has any correspondence with smaller or larger losses for the President’s party in the House of Representatives in midterm elections. If there was no linear relationship between the two variables (i.e., if \(\beta_1 = 0)\), then we would expect to see linear models as or more extreme that the observed model roughly 30% of the time.

24.4.3 Observed statistic vs. null statistics

As the final step in a mathematical hypothesis test for the slope, we use the information provided to make a conclusion about whether the data could have come from a population where the true slope was zero (i.e., \(\beta_1 = 0\)). Before evaluating the formal hypothesis claim, sometimes it is important to check your intuition. Based on everything we have seen in the examples above describing the variability of a line from sample to sample, ask yourself if the linear relationship given by the data could have come from a population in which the slope was truly zero.

Examine Figure 7.13, which relates the Elmhurst College aid and student family income. Are you convinced that the slope is discernibly different from zero? That is, do you think a formal hypothesis test would reject the claim that the true slope of the line should be zero?


While the relationship between the variables is not perfect, there is an evident decreasing trend in the data. Such a distinct trend suggests that the hypothesis test will reject the null claim that the slope is zero.

The elmhurst data can be found in the openintro R package.

The tools in this section help you go beyond a visual interpretation of the linear relationship toward a formal mathematical claim about whether the slope estimate is meaningfully different from 0 to suggest that the true population slope is different from 0.

Table 24.4: Summary of least squares fit for the Elmhurst College data, where we are predicting the gift aid by the university based on the family income of students.
term estimate std.error statistic p.value
(Intercept) 24319.33 1291.45 18.83 <0.0001
family_income -0.04 0.01 -3.98 2e-04

Table 24.4 shows statistical software output from fitting the least squares regression line shown in Figure 7.13. Use the output to formally evaluate the following hypotheses.2

  • \(H_0\): The true coefficient for family income is zero.
  • \(H_A\): The true coefficient for family income is not zero.

Inference for regression.

We usually rely on statistical software to identify point estimates, standard errors, test statistics, and p-values in practice. However, be aware that software will not generally check whether the method is appropriate, meaning we must still verify conditions are met. See Section 24.6.

24.5 Mathematical model, interval for the slope

Similar to how we can conduct a hypothesis test for a model coefficient using regression output, we can also construct confidence intervals for the slope and intercept coefficients.

Confidence intervals for coefficients.

Confidence intervals for model coefficients (e.g., the intercept or the slope) can be computed using the \(t\)-distribution:

\[ b_i \ \pm\ t_{df}^{\star} \times SE_{b_{i}} \]

where \(t_{df}^{\star}\) is the appropriate \(t^{\star}\) cutoff corresponding to the confidence level with the model’s degrees of freedom, \(df = n - 2\).

Compute the 95% confidence interval for the coefficient using the regression output from Table 24.4.


The point estimate is -0.0431 and the standard error is \(SE = 0.0108\). When constructing a confidence interval for a model coefficient, we generally use a \(t\)-distribution. The degrees of freedom for the distribution are noted in the regression output, \(df = 48\), allowing us to identify \(t_{48}^{\star} = 2.01\) for use in the confidence interval.

We can now construct the confidence interval in the usual way:

\[ \begin{aligned} \text{point estimate} &\pm t_{48}^{\star} \times SE \\ -0.0431 &\pm 2.01 \times 0.0108 \\ (-0.0648 &, -0.0214) \end{aligned} \]

We are 95% confident that for an additional one unit (i.e., $1000 increase) in family income, the university’s gift aid is predicted to decrease on average by $21.40 to $64.80.

On the topic of intervals in this book, we have focused exclusively on confidence intervals for model parameters. However, there are other types of intervals that may be of interest (and are outside the scope of this book), including prediction intervals for a response value and confidence intervals for a mean response value in the context of regression.

24.6 Checking model conditions

In the previous sections, we used randomization and bootstrapping to perform inference when the mathematical model was not valid due to violations of the technical conditions. In this section, we’ll provide details for when the mathematical model is appropriate and a discussion of technical conditions needed for the randomization and bootstrapping procedures. Recall from Section 7.1.3 that residual plots can be used to visualize how well a linear model fits the data.

24.6.1 What are the technical conditions for the mathematical model?

When fitting a least squares line, we generally require the following:

  • Linearity. The data should show a linear trend. If there is a nonlinear trend (e.g., first panel of Figure 24.15) an advanced regression method from another book or later course should be applied.

  • Independent observations. Be cautious about applying regression to data that are sequential observations in time such as a stock price each day. Such data may have an underlying structure that should be considered in a different type of model and analysis. An example of a dataset where successive observations are not independent is shown in the fourth panel of Figure 24.15. There are also other instances where correlations within the data are important, which is further discussed in Chapter 25.

  • Nearly normal residuals. Generally, the residuals should be nearly normal. When this condition is found to be unreasonable, it is often because of outliers or concerns about influential points, which we’ll talk about more in Section 7.3. An example of a residual that would be potentially concerning is shown in the second panel of Figure 24.15, where one observation is clearly much further from the regression line than the others. Outliers should be treated extremely carefully. Do not automatically remove an outlier if it truly belongs in the dataset. However, be honest about its impact on the analysis. A strategy for dealing with outliers is to present two analyses: one with the outlier and one without the outlier. Additionally, a type of violation of normality happens when the positive residuals are smaller in magnitude than the negative residuals (or vice versa). That is, when the residuals are not symmetrically distributed around the line \(y=0.\)

A grid of 2 by 4 scatterplots with fabricated data. The top row of plots contains original x-y data plots with a least squares regression line. The bottom row of plots is a series of residual plot with predicted value on the x-axis and residual on the y-axis. The first column of plots gives an example of points that have a quadratic relationship instead of a linear relationship. The second column of plots gives an example where a single outlying point does not fit the linear model. The third column of points gives an example where the points have increasing variability as the value of x increases. The last column of points gives an example where the points are correlated with one another, possibly as part of a time series.
Figure 24.15: Four examples showing when the methods in this chapter are insufficient to apply a linear model to the data. The top set of graphs represents the \(x\) and \(y\) relationship. The bottom set of graphs is a residual plot. First panel – linearity fails. Second panel – there are outliers, most especially one point that is very far away from the line. Third panel – the variability of the errors is related to the value of \(x\). Fourth panel – a time series dataset is shown, where successive observations are highly correlated.
  • Constant or equal variability. The variability of points around the least squares line remains roughly constant. An example of non-constant variability is shown in the third panel of Figure 24.15, which represents the most common pattern observed when this condition fails: the variability of \(y\) is larger when \(x\) is larger.

Should we have concerns about applying least squares regression to the Elmhurst data in Figure 7.14?3

The technical conditions are often remembered using the LINE mnemonic. The linearity, normality, and equality of variance conditions usually can be assessed through residual plots, as seen in Figure 24.15. A careful consideration of the experimental design should be undertaken to confirm that the observed values are indeed independent.

  • L: linear model
  • I: independent observations
  • N: points are normally distributed around the line
  • E: equal variability around the line for all values of the explanatory variable

24.6.2 Why do we need technical conditions?

As with other inferential techniques we have covered in this text, if the technical conditions above do not hold, then it is not possible to make concluding claims about the population. That is, without the technical conditions, the T score will not have the assumed t-distribution. That said, it is almost always impossible to check the conditions precisely, so we look for large deviations from the conditions. If there are large deviations, we will be unable to trust the calculated p-value or the endpoints of the resulting confidence interval.

The model based on linearity

The linearity condition is among the most important if your goal is to understand a linear model between \(x\) and \(y\). For example, the value of the slope will not be at all meaningful if the true relationship between \(x\) and \(y\) is quadratic, as in Figure 7.3. Not only should we be cautious about the inference, but the model itself is also not an accurate portrayal of the relationship between the variables. However, an extended discussion on the different methods for modeling functional forms other than linear is outside the scope of this text.

The importance of independence

The technical condition describing the independence of the observations is often the most crucial but also the most difficult to diagnose. It is also extremely difficult to gather a dataset which is a true random sample from the population of interest. A true randomized experiment from a fixed set of individuals is much easier to implement, and indeed, randomized experiments are done in most medical studies these days.

Dependent observations can bias results in ways that produce fundamentally flawed analyses. That is, if you hang out at the gym measuring height and weight, your linear model is surely not a representation of all students at your university. At best it is a model describing students who use the gym, but also who are willing to talk to you, that use the gym at the times you were there measuring, etc.

In lieu of trying to answer whether your observations are a true random sample, you might instead focus on whether you believe your observations are representative of a population of interest. Humans are notoriously bad at implementing random procedures, so you should be wary of any process that used human intuition to balance the data with respect to, for example, the demographics of the individuals in the sample.

Some thoughts on normality

The normality condition requires that points vary symmetrically around the line, spreading out in a bell-shaped fashion. You should consider the “bell” of the normal distribution as sitting on top of the line (coming off the paper in a 3-D sense) so as to indicate that the points are dense close to the line and disperse gradually as they get farther from the line.

The normality condition is less important than linearity or independence for a few reasons. First, the linear model fit with least squares will still be an unbiased estimate of the true population model. However, the distribution of the estimate will be unknown. Fortunately the Central Limit Theorem (described in Section 19.2) tells us that most of the analyses (e.g., SEs, p-values, confidence intervals) done using the mathematical model (with the \(t\)-distribution) will still hold (even if the data are not normally distributed around the line) as long as the sample size is large enough. One analysis method that does require normality, regardless of sample size, is creating intervals which predict the response of individual outcomes at a given \(x\) value, using the linear model. One additional reason to worry slightly less about normality is that neither the randomization test nor the bootstrapping procedures require the data to be normal around the line.

Equal variability for prediction in particular

As with normality, the equal variability condition (that points are spread out in similar ways around the line for all values of \(x\)) will not cause problems for the estimate of the linear model. That said, the inference on the model (e.g., computing p-values) will be incorrect if the variability around the line is extremely heterogeneous. Data that exhibit non-equal variance across the range of x-values will have the potential to seriously mis-estimate the variability of the slope which will have consequences for the inference results (i.e., hypothesis tests and confidence intervals).

In many cases, the inference results for both a randomization test or a bootstrap confidence interval are also robust to the equal variability condition, so they provide the analyst a set of methods to use when the data are heteroskedastic (that is, exhibit unequal variability around the regression line). Although randomization tests and bootstrapping allow us to analyze data using fewer conditions, some technical conditions are required for all methods described in this text (e.g., independent observation). When the equal variability condition is violated and a mathematical analysis (e.g., p-value from T score) is needed, there are other existing methods (outside the scope of this text) which can handle the unequal variance (e.g., weighted least squares analysis).

24.6.3 What if all the technical conditions are met?

When the technical conditions are met, the least squares regression model and inference is provided by virtually all statistical software. In addition to being ubiquitous, however, an additional advantage to the least squares regression model (and related inference) is that the linear model has important extensions (which are not trivial to implement with bootstrapping and randomization tests). In particular, random effects models, repeated measures, and interaction are all linear model extensions which require the above technical conditions. When the technical conditions hold, the extensions to the linear model can provide important insight into the data and research question at hand. We will discuss some of the extended modeling and associated inference in Chapter 25 and Chapter 26. Many of the techniques used to deal with technical condition violations are outside the scope of this text, but they are taught in universities in the very next class after this one. If you are working with linear models or are curious to learn more, we recommend that you continue learning about statistical methods applicable to a larger class of datasets.

24.7 Chapter review

24.7.1 Summary

Recall that early in the text we presented graphical techniques which communicated relationships across multiple variables. We also used modeling to formalize the relationships. Many chapters were dedicated to inferential methods which allowed claims about the population to be made based on samples of data. Not only did we present the mathematical model for each of the inferential techniques, but when appropriate, we also presented bootstrapping and permutation methods.

In Chapter 24 we brought all of those ideas together by considering inferential claims on linear models through randomization tests, bootstrapping, and mathematical modeling. We continue to emphasize the importance of experimental design in making conclusions about research claims. In particular, recall that variability can come from different sources (e.g., random sampling vs. random allocation, see Figure 2.8).

24.7.2 Terms

The terms introduced in this chapter are presented in Table 24.5. If you’re not sure what some of these terms mean, we recommend you go back in the text and review their definitions. You should be able to easily spot them as bolded text.

Table 24.5: Terms introduced in this chapter.
bootstrap CI for the slope randomization test for the slope technical conditions linear regression
inference with single predictor regression t-distribution for slope variability of the slope

24.8 Exercises

Answers to odd-numbered exercises can be found in Appendix A.24.

  1. Body measurements, randomization test. Researchers studying anthropometry collected body and skeletal diameter measurements, as well as age, weight, height and sex for 507 physically active individuals. A linear model is built to predict height based on shoulder girth (circumference of shoulders measured over deltoid muscles), both measured in centimeters.4 (Heinz et al. 2003) Shown below are the linear model output for predicting height from shoulder girth and the histogram of slopes from 1,000 randomized datasets (1,000 times, hgt was permuted and regressed against sho_gi). The red vertical line is drawn at the observed slope value which was produced in the linear model output.

    term estimate std.error statistic p.value
    (Intercept) 105.832 3.27 32.3 <0.0001
    sho_gi 0.604 0.03 20.0 <0.0001

    1. What are the null and alternative hypotheses for evaluating whether the slope of the model predicting height from shoulder girth is differen than 0.

    2. Using the histogram which describes the distribution of slopes when the null hypothesis is true, find the p-value and conclude the hypothesis test in the context of the problem (use words like shoulder girth and height).

    3. Is the conclusion based on the histogram of randomized slopes consistent with the conclusion from the mathematical model? Explain your reasoning.

  2. Baby’s weight and father’s age, randomization test. US Department of Health and Human Services, Centers for Disease Control and Prevention collect information on births recorded in the country. The data used here are a random sample of 1000 births from 2014. Here, we study the relationship between the father’s age and the weight of the baby.5 (ICPSR 2014) Shown below are the linear model output for predicting baby’s weight (in pounds) from father’s age (in years) and the histogram of slopes from 1000 randomized datasets (1000 times, weight was permuted and regressed against fage). The red vertical line is drawn at the observed slope value which was produced in the linear model output.

    term estimate std.error statistic p.value
    (Intercept) 7.101 0.199 35.674 <0.0001
    fage 0.005 0.006 0.757 0.4495

    1. What are the null and alternative hypotheses for evaluating whether the slope of the model for predicting baby’s weight from father’s age is different than 0?

    2. Using the histogram which describes the distribution of slopes when the null hypothesis is true, find the p-value and conclude the hypothesis test in the context of the problem (use words like father’s age and weight of baby). What does the conclusion of your test say about whether the father’s age is a useful predictor of baby’s weight?

    3. Is the conclusion based on the histogram of randomized slopes consistent with the conclusion from the mathematical model? Explain your reasoning.

  1. Body measurements, mathematical test. The scatterplot and least squares summary below show the relationship between weight measured in kilograms and height measured in centimeters of 507 physically active individuals. (Heinz et al. 2003)

    term estimate std.error statistic p.value
    (Intercept) -105.01 7.54 -13.9 <0.0001
    hgt 1.02 0.04 23.1 <0.0001
    1. Describe the relationship between height and weight.

    2. Write the equation of the regression line. Interpret the slope and intercept in context.

    3. Do the data provide convincing evidence that the true slope parameter is different than 0? State the null and alternative hypotheses, report the p-value (using a mathematical model), and state your conclusion.

    4. The correlation coefficient for height and weight is 0.72. Calculate \(R^2\) and interpret it in context.

  1. Baby’s weight and father’s age, mathematical test. Is the father’s age useful in predicting the baby’s weight? The scatterplot and least squares summary below show the relationship between baby’s weight (measured in pounds) and father’s age for a random sample of babies. (ICPSR 2014)

    term estimate std.error statistic p.value
    (Intercept) 7.1042 0.1936 36.698 <0.0001
    fage 0.0047 0.0061 0.779 0.4359
    1. What is the predicted weight of a baby whose father is 30 years old?

    2. Do the data provide convincing evidence that the model for predicting baby weights from father’s age has a slope different than 0? State the null and alternative hypotheses, report the p-value (using a mathematical model), and state your conclusion.

    3. Based on your conclusion, is father’s age a useful predictor of baby’s weight?

  1. Body measurements, bootstrap percentile interval. In order to estimate the slope of the model predicting height based on shoulder girth (circumference of shoulders measured over deltoid muscles), 1,000 bootstrap samples are taken from a dataset of body measurements from 507 people. A linear model predicting height based on shoulder girth is fit to each bootstrap sample, and the slope is estimated. A histogram of these slopes is shown below. (Heinz et al. 2003)

    1. Using the bootstrap percentile method and the histogram above, find a 98% confidence interval for the slope parameter.

    2. Interpret the confidence interval in the context of the problem.

  1. Baby’s weight and father’s age, bootstrap percentile interval. US Department of Health and Human Services, Centers for Disease Control and Prevention collect information on births recorded in the country. The data used here are a random sample of 1000 births from 2014. Here, we study the relationship between the father’s age and the weight of the baby. Below is the bootstrap distribution of the slope statistic from 1,000 different bootstrap samples of the data. (ICPSR 2014)

    1. Using the bootstrap percentile method and the histogram above, find a 95% confidence interval for the slope parameter.

    2. Interpret the confidence interval in the context of the problem.

  1. Body measurements, standard error bootstrap interval. A linear model is built to predict height based on shoulder girth (circumference of shoulders measured over deltoid muscles), both measured in centimeters. (Heinz et al. 2003) Shown below are the linear model output for predicting height from shoulder girth and the bootstrap distribution of the slope statistic from 1,000 different bootstrap samples of the data.

    term estimate std.error statistic p.value
    (Intercept) 105.832 3.27 32.3 <0.0001
    sho_gi 0.604 0.03 20.0 <0.0001
    1. Using the histogram, approximate the standard error of the slope statistic (that is, quantify the variability of the slope statistic from sample to sample).

    2. Find a 98% bootstrap SE confidence interval for the slope parameter.

    3. Interpret the confidence interval in the context of the problem.

     

  1. Baby’s weight and father’s age, standard error bootstrap interval. US Department of Health and Human Services, Centers for Disease Control and Prevention collect information on births recorded in the country. The data used here are a random sample of 1000 births from 2014. Here, we study the relationship between the father’s age and the weight of the baby. (ICPSR 2014) Shown below are the linear model output for predicting baby’s weight (in pounds) from father’s age (in years) and the the bootstrap distribution of the slope statistic from 1000 different bootstrap samples of the data.

    term estimate std.error statistic p.value
    (Intercept) 7.101 0.199 35.674 <0.0001
    fage 0.005 0.006 0.757 0.4495
    1. Using the histogram, approximate the standard error of the slope statistic (that is, quantify the variability of the slope statistic from sample to sample).

    2. Find a 95% bootstrap SE confidence interval for the slope parameter.

    3. Interpret the confidence interval in the context of the problem.

     

  1. Body measurements, conditions. The scatterplot below shows the residuals (on the y-axis) from the linear model of weight vs. height from a dataset of body measurements from 507 physically active individuals. The x-axis is the height of the individuals, in cm. (Heinz et al. 2003)

    1. For these data, \(R^2\) is 51.84%. What is the value of the correlation coefficient? How can you tell if it is positive or negative? (Hint: you may need to look at a previous exercise.)

    2. Examine the residual plot. What do you observe? Is a simple least squares fit appropriate for these data? Which of the LINE conditions are met or not met?

  1. Baby’s weight and father’s age, conditions. The scatterplot below shows the residuals (on the y-axis) from the linear model of baby’s weight (measured in pounds) vs. father’s age for a random sample of babies. Father’s age is on the x-axis. (ICPSR 2014)

    1. For these data, \(R^2\) is 0.09%. What is the value of the correlation coefficient? How can you tell if it is positive or negative? (Hint: you may need to look at a previous exercise.)

    2. Examine the residual plot. What do you observe? Is a simple least squares fit appropriate for these data? Which of the LINE conditions are met or not met?

  1. Murders and poverty, randomization test. The following regression output is for predicting annual murders per million (annual_murders_per_mil) from percentage living in poverty (perc_pov) in a random sample of 20 metropolitan areas. Shown below are the linear model output for predicting annual murders per million from percentage living in poverty for metropolitan areas and the histogram of slopes from 1000 randomized datasets (1000 times, annual_murders_per_mil was permuted and regressed against perc_pov). The red vertical line is drawn at the observed slope value which was produced in the linear model output.

    term estimate std.error statistic p.value
    (Intercept) -29.90 7.79 -3.84 0.0012
    perc_pov 2.56 0.39 6.56 <0.0001
    1. What are the null and alternative hypotheses for evaluating whether the slope of the model for predicting annual murder rate from poverty percentage is different than 0?

    2. Using the histogram which describes the distribution of slopes when the null hypothesis is true, find the p-value and conclude the hypothesis test in the context of the problem (use words like murder rate and poverty).

    3. Is the conclusion based on the histogram of randomized slopes consistent with the conclusion which would have been obtained using the mathematical model? Explain your reasoning.

     

  2. Murders and poverty, mathematical test. The table below shows the output of a linear model annual murders per million (annual_murders_per_mil) from percentage living in poverty (perc_pov) in a random sample of 20 metropolitan areas.

    term estimate std.error statistic p.value
    (Intercept) -29.90 7.79 -3.84 0.0012
    perc_pov 2.56 0.39 6.56 <0.0001
    1. What are the hypotheses for evaluating whether the slope of the model predicting annual murder rate from poverty percentage is different than 0?

    2. State the conclusion of the hypothesis test from part (a) in context. What does this say about whether poverty percentage is a useful predictor of annual murder rate?

    3. Calculate a 95% confidence interval for the slope of poverty percentage, and interpret it in context.

    4. Do your results from the hypothesis test and the confidence interval agree? Explain your reasoning.

  1. Murders and poverty, bootstrap percentile interval. Data on annual murders per million (annual_murders_per_mil) and percentage living in poverty (perc_pov) is collected from a random sample of 20 metropolitan areas. Using these data we want to estimate the slope of the model predicting annual_murders_per_mil from perc_pov. We take 1,000 bootstrap samples of the data and fit a linear model predicting annual_murders_per_mil from perc_pov to each bootstrap sample. A histogram of these slopes is shown below.

    1. Using the percentile bootstrap method and the histogram above, find a 90% confidence interval for the slope parameter.

    2. Interpret the confidence interval in the context of the problem.

  2. Murders and poverty, standard error bootstrap interval. A linear model is built to predict annual murders per million (annual_murders_per_mil) from percentage living in poverty (perc_pov) in a random sample of 20 metropolitan areas. Shown below are the standard linear model output for predicting annual murders per million from percentage living in poverty for metropolitan areas and the bootstrap distribution of the slope statistic from 1000 different bootstrap samples of the data.

    term estimate std.error statistic p.value
    (Intercept) -29.90 7.79 -3.84 0.0012
    perc_pov 2.56 0.39 6.56 <0.0001

    1. Using the histogram, approximate the standard error of the slope statistic (that is, quantify the variability of the slope statistic from sample to sample).

    2. Find a 90% bootstrap SE confidence interval for the slope parameter.

    3. Interpret the confidence interval in the context of the problem.

  1. Murders and poverty, conditions. The scatterplot below shows the annual murders per million vs. percentage living in poverty in a random sample of 20 metropolitan areas. The second figure plots residuals on the y-axis and percent living in poverty on the x-axis.

    1. For these data, \(R^2\) is 70.56%. What is the value of the correlation coefficient? How can you tell if it is positive or negative?

    2. Examine the residual plot. What do you observe? Is a simple least squares fit appropriate for the data? Which of the LINE conditions are met or not met?

  1. I heart cats. Researchers collected data on heart and body weights of 144 domestic adult cats. The table below shows the output of a linear model predicting heart weight (measured in grams) from body weight (measured in kilograms) of these cats.6

    term estimate std.error statistic p.value
    (Intercept) -0.357 0.692 -0.515 0.6072
    Bwt 4.034 0.250 16.119 <0.0001
    1. What are the hypotheses for evaluating whether body weight is positively associated with heart weight in cats?

    2. State the conclusion of the hypothesis test from part (a) in context.

    3. Calculate a 95% confidence interval for the slope of body weight, and interpret it in context.

    4. Do your results from the hypothesis test and the confidence interval agree? Explain your reasoning.

  1. Beer and blood alcohol content. Many people believe that weight, drinking habits, and many other factors are much more important in predicting blood alcohol content (BAC) than simply considering the number of drinks a person consumed. Here we examine data from sixteen student volunteers at Ohio State University who each drank a randomly assigned number of cans of beer. These students were evenly divided between men and women, and they differed in weight and drinking habits. Thirty minutes later, a police officer measured their blood alcohol content (BAC) in grams of alcohol per deciliter of blood. The scatterplot and regression table summarize the findings. 7 (Malkevitch and Lesser 2008)

    term estimate std.error statistic p.value
    (Intercept) -0.0127 0.0126 -1.00 0.332
    beers 0.0180 0.0024 7.48 <0.0001
    1. Describe the relationship between the number of cans of beer and BAC.

    2. Write the equation of the regression line. Interpret the slope and intercept in context.

    3. Do the data provide convincing evidence that drinking more cans of beer is associated with an increase in blood alcohol? State the null and alternative hypotheses, report the p-value, and state your conclusion.

    4. The correlation coefficient for number of cans of beer and BAC is 0.89. Calculate \(R^2\) and interpret it in context.

    5. Suppose we visit a bar in our own town, ask people how many drinks they have had, and also measure their BAC. Would the relationship between number of drinks and BAC would be as strong as the relationship found in the Ohio State study? Why?

  1. Urban homeowners, conditions. The scatterplot below shows the percent of families who own their home vs. the percent of the population living in urban areas. (US Census Bureau 2010) There are 52 observations, each corresponding to a state in the US. Puerto Rico and District of Columbia are also included. The second figure plots residuals on the y-axis and percent of the population living in urban areas on the x-axis.

    1. For these data, \(R^2\) is 29.16%. What is the value of the correlation coefficient? How can you tell if it is positive or negative?

    2. Examine the residual plot. What do you observe? Is a simple least squares fit appropriate for the data? Which of the LINE conditions are met or not met?

  1. I heart cats, LINE conditions. Researchers collected data on heart and body weights of 144 domestic adult cats. The figure below shows the output of the predicted values and residuals generated from a linear model predicting heart weight (measured in grams) from body weight (measured in kilograms) of these cats.

    1. Examine the residual plot. Notice that for the small predicted values the residuals have a smaller magnitude than the larger residuals seen with the larger predicted values. The change in magnitude of the residuals across the predicted values is an indication of violation of which LINE technical condition?

    2. If the LINE condtion described in part (a) is violated, might it lead to an incorrect conclusion about the model (i.e., the least squares regression line itself), the inference of the model (i.e., the p-value associated with the least squares regression line), neither, or both? Explain your reasoning.

  1. Beer and blood alcohol content, LINE conditions. The figure below shows the output of the predicted values and residuals generated from a linear model predicting the blood alcohol content (BAC) from number of cans of beer drunk by sixteen student volunteers at Ohio State University.8 (Malkevitch and Lesser 2008)

    1. Examine the residual plot. Notice that it is difficult to identify any convincing patterns for or against violation of the LINE technical conditions. What is it about the residual plot that makes it difficult to assess the LINE technical conditions?

    2. Is there anything about the residual plot which would make you hesitate about using the linear model for inference about all students? Is there anything about the experimental design of the study which would make you hesitate about using the linear model for inference about all students?


  1. The answer to this question relies on the idea that statistical data analysis is somewhat of an art. That is, in many situations, there is no “right” answer. As you do more and more analyses on your own, you will come to recognize the nuanced understanding which is needed for a particular dataset. In terms of the Great Depression, we will provide two contrasting considerations. Each of these points would have very high leverage on any least-squares regression line, and years with such high unemployment may not help us understand what would happen in other years where the unemployment is only modestly high. On the other hand, the Depression years are exceptional cases, and we would be discarding important information if we exclude them from a final analysis.↩︎

  2. We look in the second row corresponding to the family income variable. We see the point estimate of the slope of the line is -0.0431, the standard error of this estimate is 0.0108, and the \(t\)-test statistic is \(T = -3.98\). The p-value corresponds exactly to the two-sided test we are interested in: 0.0002. The p-value is so small that we reject the null hypothesis and conclude that family income and financial aid at Elmhurst College for freshman entering in the year 2011 are negatively correlated and the true slope parameter is indeed less than 0, just as we believed in our analysis of Figure 7.13.↩︎

  3. The trend appears to be linear, the data fall around the line with no obvious outliers, the variance is roughly constant. The data do not come from a time series or other obvious violation to independence. Least squares regression can be applied to these data.↩︎

  4. The bdims data used in this exercise can be found in the openintro R package.↩︎

  5. The births14 data used in this exercise can be found in the openintro R package.↩︎

  6. The cats data used in this exercise can be found in the MASS R package.↩︎

  7. The bac data used in this exercise can be found in the openintro R package.↩︎

  8. The bac data used in this exercise can be found in the openintro R package.↩︎

中文

24  单预测变量线性回归的推断

现在,我们将推断分析的思想与 第 7中看到的描述性模型结合起来。特别地,我们将使用最小二乘回归线来检验两个连续变量之间是否存在关系。此外,我们还将构建置信区间来量化线性回归线的斜率。现在的设定聚焦于预测一个数值型响应变量(对于线性模型)或二元响应变量(对于逻辑模型),我们继续探讨模型从样本到样本的变异性的问题。抽样变异性将为关于总体的结论提供依据。

许多推断思想与前面章节中介绍的内容非常相似。线性模型的技术条件通常通过图形来评估,尽管观测值的独立性仍然至关重要。

我们鼓励读者从更宏观的角度看待手头的模型,而不要过分依赖统计软件报告的精确 p 值。对包含多个解释变量的模型进行推断时,可能会受到数据窥探的影响,从而导致假阳性结论。我们提供了一些指导,并希望读者在完成本书内容的学习后能继续深化统计学学习。

24.1 案例研究:三明治店

24.1.1 观测数据

我们以一个假设的例子开始本章,该例子描述了某连锁三明治餐厅的广告支出(美元)与月收入之间的线性关系。这个假设例子的目的是说明线性模型如何随样本而变化。由于我们虚构了这个例子和数据(以及整个总体),我们可以从总体中抽取许多许多样本,以直观展示变异性。请注意,在现实生活中,我们总是恰好只有一个样本(即一个数据集),而通过推断过程,我们设想如果抽取了不同的样本可能会发生什么。样本之间的变化有助于理解单个观测数据集与总体数值之间的差异,而这通常是推断的根本目标。

考虑以下假设总体,即某连锁品牌的全部三明治门店,如 图 24.1. 在这个虚构的世界中,CEO 实际上拥有所有相关数据,这就是为什么他们可以在这里绘制它。CEO 是全知的,可以写出描述广告费用与收入之间真实总体关系的总体模型。广告费用与收入(均以千美元计)之间似乎存在线性关系。

Scatterplot with advertising amount on the x-axis and revenue on the y-axis. A linear model is superimposed. The points show a reasonably strong and positive linear trend.
图 24.1:对于三明治店总体,收入作为广告费用的线性模型,单位为千美元。

你可能还记得 第 7 章 中,总体模型是: \[y = \beta_0 + \beta_1 x + \varepsilon.\]

同样,全知的 CEO(拥有完整的总体信息)可以写出真实的总体模型为: \[\texttt{expected revenue} = 11.23 + 4.8 \times \texttt{advertising}.\]

24.1.2 统计量的变异性

不幸的是,在我们的场景中,CEO 不愿意交出完整的数据集,但他们允许潜在加盟商购买者查看一小部分数据样本,以帮助潜在购买者决定是否开设新加盟店。CEO 愿意给每个潜在加盟商购买者一个来自 20 家门店的随机数据样本。

与任何描述总体子集的数值特征一样,样本的估计斜率会因样本而异。考虑基于广告费用(千美元)描述收入(千美元)的线性模型。

最小二乘回归模型使用数据来找到样本线性拟合: \[\hat{y} = b_0 + b_1 x.\]

两个 20 家门店的随机样本在 图 24.2 (a)图 24.2 (b)中显示出不同的最小二乘回归线,具体取决于选择了哪些观测值。两个趋势都与 图 24.1中描述总体的趋势相似。

For two random samples of 20 stores, scatterplots with advertising amount on the x-axis and revenue on the y-axis. Linear models are superimposed. The points in each plot show reasonably strong and positive linear trends.
(a) 第一个样本。
For two random samples of 20 stores, scatterplots with advertising amount on the x-axis and revenue on the y-axis. Linear models are superimposed. The points in each plot show reasonably strong and positive linear trends.
(b) 第二个样本。
图 24.2:来自整个总体的两个 20 家门店的随机样本。两者中均观察到广告与收入之间的线性趋势。

图 24.3 展示了两个样本及其来自 图 24.2 的最小二乘回归,绘制在同一张图上。我们可以看到这两条线是不同的。也就是说,回归线在不同样本之间存在 变异性 。抽样变异性的概念你之前已经见过,但在本课中,你将重点关注这条线的变异性,通常通过单个统计量的变异性来衡量: 直线的斜率.

For two different random samples, superimposed onto the same plot, scatterplot with advertising amount on the x-axis and revenue on the y-axis. Two linear models are plotted to demonstrate that the lines are very similar, yet they are not the same.
图 24.3:来自两个随机样本的线性模型非常相似,但不完全相同。

图 24.4 展示了对来自总体的更多个容量为 20 的随机样本拟合的最小二乘直线。

An x-y coordinate system with least squares regression lines from many random samples of size 20 (no points are plotted). The lines vary around the true population line. On the x-axis is advertising amount; on the y-axis is revenue.
图 24.4:如果从整个总体中反复抽取容量为 20 的样本,每个线性模型都会略有不同。红线表示对整个总体的线性拟合。

你可能会在 图 24.4 中注意到,这些直线给出的 \(\hat{y}\) 值在数据集的中部比在两端更加一致。原因在于数据本身将直线“锚定”,使得直线必须穿过数据云的中心。扇形分布的直线带来的效果是,对于接近 4,000 美元的广告投入,预测的收入将比针对 1,000 美元或 7,000 美元广告投入的收入预测精确得多。

斜率(对于容量为 \(n=20\)的样本)的分布可以在 图 24.5.

Histogram of the slope values from many random samples of size 20. The slope estimates vary from about 2.1 to 8. The histogram is reasonably bell-shaped.
中看到。

图 24.5:来自许多不同容量为 20 的商店样本的斜率估计值的变异性。 \(\overline{X}_1 - \overline{X}_2\)\(\hat{p}_1 - \hat{p}_2\)回顾一下,本引言中描述的例子是假设性的。也就是说,我们创建了一个完整的总体,以演示直线的斜率在不同样本之间会如何变化。本教科书中的工具仅用于评估单个数据样本。在实际研究中,我们没有重复的样本,因此无法利用重复样本来直观地展示斜率的变异性。我们在本书中已经见过样本的变异性,因此不同的样本会产生不同的线性模型并不令人意外。不过,直观地观察由不同斜率产生的线性模型还是很有帮助的。此外,与测量之前统计量(例如

)的变异性一样,样本统计量的直方图可以提供与推断相关的信息。 \(b_1\) (估计的斜率系数)将以你如今可能已经熟悉的三种方式来构建。首先(在 Section 24.2中),通过随机化(置换)响应变量来构建 \(b_1\) 的概率,当 \(\beta_1 = 0\) 的分布。接下来(在 Section 24.3中),我们可以通过从原始数据集中抽取大小为 \(n\) 的随机样本来对数据进行自助法(bootstrap)重采样。最后(在 Section 24.4中),我们使用数学工具,通过 \(t\)分布来描述变异性,该分布最早在 第 19.2 节.

24.2 中出现过。

斜率的随机化检验 图 24.6.

考虑最初从美国卫生与公众服务部收集的 100 例随机抽取的出生数据。其中一些变量绘制在

births14 数据可以在 openintro 中。这里的科学研究兴趣在于确定婴儿出生体重(以磅为单位)与妊娠周数之间的线性关系。该数据集相当丰富,值得探索,但在本例中,我们将只关注婴儿的体重。

Four different scatterplots, all with weight of baby on the y-axis. On the x-axis are weight gained by mother, mother's age, number of hospital visits, and weeks gestation. Weeks gestation and weight of baby show the strongest linear relationship (which is positive).
R 包。我们将使用这些数据中 100 个观测值的随机样本。

正如您之前所见,统计推断通常依赖于设定一个希望随后被拒绝的原假设。在线性模型的设定中,我们可能希望 weeksweightweeks 之间存在线性关系,在已知怀孕期(gestation)而需要预测婴儿 weight 的场景中。

线性模型设定中的相关假设可以用总体斜率参数来表述。这里的总体指的是美国更大范围的出生人口。

  • \(H_0: \beta_1= 0\),则 weightweeks.
  • \(H_A: \beta_1 \ne 0\)之间不存在线性关系 weightweeks.

,则 之间存在某种线性关系 图 回想一下,在随机化检验中,我们对一个变量进行置换,以消除变量之间任何已有的关系。也就是说,我们将原假设设为真,并测量数据中由抽样引起的自然变异,但 由变量之间相关引起的变异。 图 24.7 (a) 展示了观测数据,而 weight 24.7 (b) weight 展示了 weeks变量的一次置换。细心的观察者可以发现,每个观测到的 weight 图,但是 weightweeks 妊娠期不再与给定的出生相匹配。也就是说,每个 weight 值被随机分配给一个新的 weeks 妊娠期。

通过反复对响应变量进行置换,线性模型中观察到的任何模式都仅由随机机会造成(而不是某种潜在关系)。随机化检验将根据置换后的响应变量计算出的斜率与观察到的斜率进行比较。如果观察到的斜率与置换得到的斜率不一致,我们可以得出结论:存在某种潜在关系(即斜率并非仅仅由随机机会造成)。

Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. The left panel is the original data. The right panel is data where the weight of the baby has been permuted across the observations.
(a) 原始数据。
Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. The left panel is the original data. The right panel is data where the weight of the baby has been permuted across the observations.
(b) 置换后的数据。
图 24.7:置换消除了 weightweeks之间的线性关系。反复置换可以在没有线性关系(即原假设为真)的条件下量化斜率的变异性。

24.2.1 观测数据

我们将继续使用出生数据来研究 weightweeks 妊娠期之间的线性关系。注意,描述该关系的最小二乘模型(见 第 7 章)在 表 24.1中给出。 表 24.1 中的各列在 Section 24.4.

表 24.1:截距和斜率的最小二乘估计值在 estimate 列中给出。观察到的斜率为 0.335。
term 估计 的线性模型摘要 统计量 std.error
p.value -5.72 1.61 -3.54 6e-04
weeks 0.34 0.04 8.07 <0.0001

24.2.2 统计量的变异性

对数据进行置换后,可以计算该直线的最小二乘估计。重复进行置换和斜率计算,可以描述仅由自然变异性(而非由以下变量之间的关系)引起的直线(即斜率)的变异性: weightweeks 妊娠期。 图 24.8 展示了 weight 的两种不同置换以及由此得到的线性模型。

Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. Each plot includes data where the weight of the baby has been permuted across the observations. The two different permutations produce slightly different least squares regression lines.
(a) 第一次置换。
Two scatterplots, both with length of gestation on the x-axis and weight of baby on the y-axis. Each plot includes data where the weight of the baby has been permuted across the observations. The two different permutations produce slightly different least squares regression lines.
(b) 第二次置换。
图 24.8: weight 的两种置换,其最小二乘回归线略有不同。

如你所见,置换后数据的斜率有时为正,有时为负。由于随机化是在不存在潜在关系的条件下进行的(因为响应变量与解释变量被完全混合),我们预期随机化斜率分布的中心为零。

24.2.3 观测统计量与零假设统计量的对比

Histogram of slopes describing the linear model from permuted weight regressed on weeks gestation. The permuted slopes range from -0.15 to +0.15 and are nowhere near the observed slope value of 0.335.
图 24.9:对 weight 变量进行不同置换后所得斜率的直方图。红色竖线位于斜率的观测值 0.335 处。

图 24.9可以看出,在对 weight 变量进行的大量重复置换中,从未出现过与观测斜率估计值(红线)同样极端的斜率估计。也就是说,如果 weightweeks之间确实不存在线性关系,斜率的自然变异性将产生大约介于 -0.15 和 +0.15 之间的估计值。我们拒绝原假设。因此,我们认为在原始数据上观测到的斜率并非仅仅由自然变异性造成,事实上, weight 婴儿体重与 weeks 美国出生数据的孕周。

24.3 斜率的自助法置信区间

正如我们在前几章中所看到的,我们可以使用自助法来估计目标统计量(这里是斜率)的抽样分布,而无需“无关系”这一原假设(那是随机化检验中的条件)。由于现在的目标是构建置信区间,不存在原假设,因此也就没有任何理由对任一变量进行置换。

24.3.1 观测数据

回到出生数据,我们可能想考虑 mage (母亲年龄)与 weight之间的关系。 mageweight的良好预测变量吗?如果是,二者是什么关系?也就是说,建模婴儿平均 weight 作为 mage (母亲年龄)的函数的斜率是多少?以 weightmage 为响应的线性模型见 表 24.2.

Scatterplot with mother's age on the x-axis and baby's weight on the y-axis. A linear model is superimposed. The points show a weak positive linear trend.
图 24.10:使用原始数据,将婴儿体重作为母亲年龄的线性模型。注意,母亲年龄与婴儿体重之间的关系不如我们之前看到的孕周与婴儿体重之间的关系强。
表 24.2:截距和斜率的最小二乘估计在 estimate 列中给出。观测到的斜率为 0.036。
term 估计 的线性模型摘要 统计量 std.error
p.value 6.23 0.71 8.79 <0.0001
mage 0.04 0.02 1.50 0.1362

24.3.2 统计量的变异性

由于这里的重点不是 原假设分布,我们采用有放回抽样 \(n = 100\) 来自原始数据集的观测值。回想一下,在使用自助法(bootstrapping)时,重抽样样本总是与原始数据集具有相同数量的观测值,以模拟从总体中抽取样本的过程。在线性模型情形下进行抽样时,将每个观测值视为一个点。如果该点被重抽样,则两个 weightmage 测量值都会被观测到。这些测量值与该点(即样本中的出生记录)相关联。

Two scatterplots, both with mother's age on the x-axis and baby's weight on the y-axis. The left plot is the original data. The right plot is the bootstrapped data. Comparing the bootstrapped points to the original points, we can see that some observations were sampled more than once, and some observations were not selected for the bootstrap sample at all.
(a) 原始数据。
Two scatterplots, both with mother's age on the x-axis and baby's weight on the y-axis. The left plot is the original data. The right plot is the bootstrapped data. Comparing the bootstrapped points to the original points, we can see that some observations were sampled more than once, and some observations were not selected for the bootstrap sample at all.
(b) 自助法重抽样数据。
图 24.11:出生数据的原始数据和一个自助法样本。很难区分这两幅图,因为(在单个自助法样本内)被重抽样两次的观测值会以重叠的点绘制。红色圆圈表示原始数据中未包含在自助法样本中的点。蓝色圆圈表示在自助法样本中被反复重抽样的数据点(因此颜色更深)。绿色圆圈表示在原始数据和自助法样本中均可观察到的一种特定数据结构。

图 24.11 (a) 展示了原始数据, 图 24.11 (b)中展示了一个自助法样本作为对比,从而得到(略有不同的)不同线性模型。红色圆圈表示原始数据中未包含在自助法样本中的点。蓝色圆圈表示在自助法样本中被反复重抽样的点(因此颜色更深)。绿色圆圈表示在原始数据和自助法样本中均可观察到的一种特定数据结构。通过反复重抽样,我们可以在同一幅图上看到数十个自助法斜率,如 图 24.12.

An x-y coordinate system with least squares regression lines from many bootstrap samples (no points are plotted). The lines vary around the observed population line. On the x-axis is mother's age; on the y-axis is baby's weight.
图 24.12:从原始数据中反复抽取容量为 100 的自助法重抽样样本。每个自助法线性模型都略有不同。

回想一下,为了构建斜率的置信区间,我们需要找出该统计量(此处为斜率)在不同自助法样本中所取值的范围。 图 24.13 是相关自助法斜率的直方图。我们可以看到,真实总体斜率的 95% 自助法百分位置信区间为 (-0.01, 0.081)。我们有 95% 的把握认为,对于描述出生总体的模型(由母亲年龄预测婴儿体重), mage (以年为单位)每增加一个单位,预测的平均婴儿 weight 会增加 -0.01 到 0.081 磅之间。请注意,该置信区间包含零,因此真实关系 可能 是零!

Histogram of the slopes computed from many bootstrapped samples. The bootstrap samples range from -0.05 (with the 2.5 percentile at -0.01) to +0.1 (with the 97.5 percentile at 0.081). The bootstrapped slopes form a histogram that is reasonably symmetric and bell-shaped.
图 24.13:对原始的婴儿体重与母亲年龄出生数据进行了 1,000 次 bootstrap 重抽样。直方图展示了线性模型斜率在样本之间的变异程度。

使用 图 24.13,计算斜率标准误的 bootstrap 估计值。利用 bootstrap 标准误,求出真实总体斜率的 95% bootstrap SE 置信区间,并结合具体情境解释该区间。


请注意,大多数 bootstrap 斜率落在 -0.01 到 +0.08 之间(范围为 0.09)。根据经验法则(对于钟形分布,大多数观测值位于中心两侧两个标准误之内),斜率的标准误约为 0.0225。95% 置信区间的临界值为 \(z^\star = 1.96\) ,由此得到置信区间 \(b_1 \pm 1.96 \cdot SE \rightarrow 0.036 \pm 1.96 \cdot 0.0225 \rightarrow (-0.0081, 0.0801).\) bootstrap SE 置信区间与 bootstrap 百分位数区间几乎完全相同。结合具体情境,我们有 95% 的把握认为,在描述总体出生数据的模型中,用母亲年龄预测婴儿体重时, mage (以年为单位)每增加一个单位,预测的平均婴儿 weight 将增加 -0.0081 到 0.0801 磅。

24.4 用于检验斜率的数学模型

当满足某些技术条件时,使用数学近似来检验和估计斜率参数会很方便。这些近似将建立在 t 分布的基础上,t 分布已在 第 19 章中介绍过。该数学模型通常正确,且在计算上通常易于实现。技术条件的有效性将在 第 24.6 节.

中详细讨论。

24.4.1 观测数据

在本节中,我们讨论回归线的斜率和 y 截距估计值的不确定性。正如在前几章中为点估计确定标准误一样,我们首先讨论斜率和 y 截距估计值的标准误。

中期选举与失业率

美国众议院议员选举每两年举行一次,每四年与美国总统大选同时进行。在总统任期中期举行的众议院选举被称为中期选举。在美国的两党制下(历史上绝大多数众议院议员要么是共和党人,要么是民主党人),有一种政治理论认为,失业率越高,总统所属政党在中期选举中的表现就越差。2020 年,众议院中有 232 名民主党议员、198 名共和党议员和 1 名自由意志党议员。

midterms_house 数据可以在 openintro R 包中找到。

图 24.14 为评估与失业率和投票模式相关的这一说法的有效性,我们可以汇编历史数据并寻找其中的联系。我们考察 1898 年至 2018 年的每一次中期选举,但大萧条期间的选举除外。众议院由 435 名有投票权的议员组成。展示了这些数据以及最小二乘回归线:

\[ \begin{aligned} &\texttt{percent change in House seats for President's party} \\ &\qquad\qquad= -7.36 - 0.89 \times \texttt{(unemployment rate)} \end{aligned} \]

我们考察总统所属政党席位数的变化百分比(例如,2018年共和党席位数的变化百分比)与失业率之间的关系。

检查数据后,未发现明显偏离线性或显著离群点的情况(参见 第7.1.3节 中关于使用残差来直观展示线性模型对数据拟合程度的讨论)。虽然数据是按时间顺序收集的,但我们进行了单独的分析来检查相邻观测值之间是否存在明显的相关性;结果未发现此类相关性。

Scatterplot with percent unemployed on the x-axis and percent change in House seats for the President's party on the y-axis. Each point represents a different President's midterm and is colored according to their political party (Democrat or Republican). The relationship is moderate and negative.
图 24.14:1898 年至 2010 年每次选举中总统所属政党在众议院席位变化百分比与失业率的关系图。大萧条时期的两个数据点已被移除,并对数据拟合了一条最小二乘回归线。

大萧条时期(1934 年和 1938 年)的数据被移除,因为失业率分别高达 21% 和 18%。你是否同意在本次研究中移除这些数据?为什么?1

图中所示的直线斜率为负(见 图 24.14)。然而,这个斜率(以及 y 轴截距)只是参数值的估计值。我们可能会问:这是否足以令人信服地证明“真实”线性模型的斜率为负?也就是说,数据是否提供了强有力的证据表明该政治理论是准确的,即失业率是中期选举结果的有效预测指标?我们可以将这项研究构建为一个统计假设检验:

  • \(H_0\): \(\beta_1 = 0\)。真实线性模型的斜率为零。
  • \(H_A\): \(\beta_1 \neq 0\)。真实线性模型的斜率不为零。即失业率可以预测总统所属政党在众议院是赢得还是失去席位。

如果数据提供了强有力的证据表明真实斜率参数不为零,我们将拒绝 \(H_0\) 而支持 \(H_A\) 。为了评估这些假设,我们需要确定估计值的标准误,计算适当的检验统计量,并确定 p 值。

24.4.2 统计量的变异性

与我们之前见过的其他点估计一样,我们可以为 \(b_1\)计算标准误和检验统计量。我们通常将检验统计量标记为 \(T\),因为它服从 \(t\)分布。

我们将依靠统计软件来计算标准误,并把这一标准误如何确定的解释留给第二或第三门统计学课程。 表 24.3 展示了 图 24.14中最小二乘回归线的软件输出。标记为 unemp 的行包含了关于斜率估计的所有相关信息(即失业变量的系数、相应的标准误、T 统计量以及对应的 p 值)。

表 24.3:统计软件输出的回归线结果,该回归线以失业率为自变量,对总统所属政党在中期选举中的席位损失进行建模。
term 估计 的线性模型摘要 统计量 std.error
p.value -7.36 5.16 -1.43 0.16
unemp -0.89 0.83 -1.07 0.30

的第一列和第二列分别代表什么? 表 24.3 第一列中的数值代表最小二乘估计值


,第二列中的数值对应每个估计值的标准误。利用这些估计值,我们可以将最小二乘回归线的方程写为 \(b_0\)\(b_1\)其中

\[ \hat{y} = -7.36 - 0.89 x \]

其中 \(\hat{y}\) 表示失业率。 \(x\) 我们之前在数值数据的背景下使用

检验统计量进行假设检验。回归与之非常相似。在我们所考虑的假设中,斜率的原假设值为 0,因此我们可以使用 T 分数公式计算检验统计量: \(t\)我们计算出的 T 分数对应于

\[ T \ = \ \frac{\text{estimate} - \text{null value}}{\text{SE}} = \ \frac{-0.89 - 0}{0.835} = \ -1.07 \]

的第三列。 表 24.3.

使用 表 24.3 以确定假设检验的 p 值。


表格的最后一列给出了失业率系数 0.2961 的双侧假设检验的 p 值。也就是说,数据并没有提供令人信服的证据表明较高的失业率与总统所属政党在中期选举中众议院席位的较大或较小损失之间存在任何对应关系。如果两个变量之间不存在线性关系(即,如果 \(\beta_1 = 0)\),那么我们预计会看到与观测模型同样极端或更极端的线性模型出现的频率约为 30%。

24.4.3 观测统计量与零假设统计量的对比

作为斜率数学假设检验的最后一步,我们利用所提供的信息得出结论:数据是否可能来自真实斜率为零的总体(即, \(\beta_1 = 0\))。在评估正式的假设主张之前,有时检验一下自己的直觉很重要。根据上面示例中关于直线从样本到样本的变异性的所有内容,问问自己:数据所呈现的线性关系是否可能来自斜率真正为零的总体。

考察 图 7.13,它涉及埃尔姆赫斯特学院的资助与学生家庭收入之间的关系。你是否相信斜率明显不同于零?也就是说,你认为正式的假设检验是否会拒绝直线真实斜率应为零这一主张?


虽然变量之间的关系并不完美,但数据中存在明显的下降趋势。如此明显的趋势表明,假设检验将拒绝斜率为零的原假设。

elmhurst 数据可以在 openintro R 包中找到。

本节中的工具帮助你超越对线性关系的直观解读,进而就斜率估计值是否与 0 有意义上的差异做出正式的数学判断,从而表明真实的总体斜率不同于 0。

表 24.4:埃尔姆赫斯特学院数据的最小二乘拟合汇总,其中我们根据学生的家庭收入来预测大学提供的资助。
term 估计 的线性模型摘要 统计量 std.error
p.value 24319.33 1291.45 18.83 <0.0001
family_income -0.04 0.01 -3.98 2e-04

表 24.4 展示了拟合 图 7.13中所示的最小二乘回归线的统计软件输出。请使用该输出正式评估以下假设。2

  • \(H_0\):家庭收入的真实系数为零。
  • \(H_A\):家庭收入的真实系数不为零。

回归推断。

在实践中,我们通常依靠统计软件来识别点估计、标准误、检验统计量和p值。但请注意,软件一般不会检验该方法是否适用,这意味着我们仍必须验证条件是否满足。参见 第 24.6 节.

24.5 数学模型,斜率的区间

类似于我们可以利用回归输出对模型系数进行假设检验,我们也可以为斜率和截距系数构建置信区间。

系数的置信区间。

模型系数(例如截距或斜率)的置信区间可以使用 \(t\)-分布来计算:

\[ b_i \ \pm\ t_{df}^{\star} \times SE_{b_{i}} \]

其中 \(t_{df}^{\star}\) 是对应于置信水平和模型自由度的相应 \(t^{\star}\) 临界值, \(df = n - 2\).

利用 表 24.4.


的回归输出计算该系数的95%置信区间。 \(SE = 0.0108\)点估计为 -0.0431,标准误为 \(t\)。在为模型系数构建置信区间时,我们通常使用 \(df = 48\)-分布。该分布的自由度已在回归输出中注明,即 \(t_{48}^{\star} = 2.01\) ,从而使我们能够确定用于置信区间的

我们现在可以按通常的方式构建置信区间:

\[ \begin{aligned} \text{point estimate} &\pm t_{48}^{\star} \times SE \\ -0.0431 &\pm 2.01 \times 0.0108 \\ (-0.0648 &, -0.0214) \end{aligned} \]

我们有95%的把握认为,家庭收入每增加一个单位(即增加1000美元),大学的助学金预计平均减少21.40美元至64.80美元。

在本书关于区间的话题中,我们一直专注于模型参数的置信区间。然而,还有其他类型的区间可能值得关注(但超出了本书的范围),包括响应值的预测区间,以及回归背景下均值响应值的置信区间。

24.6 检查模型条件

在前面的章节中,当数学模型由于违反技术条件而无效时,我们使用随机化和自助法进行推断。在本节中,我们将详细说明数学模型何时适用,并讨论随机化和自助法程序所需的技术条件。回顾 第7.1.3节 ,残差图可用于直观展示线性模型对数据的拟合程度。

24.6.1 数学模型的技术条件是什么?

在拟合最小二乘直线时,我们通常要求满足以下条件:

  • 线性。 数据应呈现线性趋势。如果存在非线性趋势(例如, 图 24.15的第一幅图),则应采用其他书籍或后续课程中的高级回归方法。

  • 独立观测。 对时间上的顺序观测数据(如每天的股票价格)应用回归时要谨慎。这类数据可能具有某种潜在结构,应在不同类型的模型和分析中加以考虑。第四幅图 图 24.15展示了一个连续观测不独立的数据集示例。还有其他一些数据内部相关性很重要的情形,将在 第 25 章.

  • 中进一步讨论。近似正态的残差。 一般来说,残差应接近正态分布。当发现该条件不合理时,通常是因为存在离群点或对强影响点的担忧,我们将在 第 7.3 节中进一步讨论。第二幅图中展示了一个可能令人担忧的残差示例, 图 24.15其中一个观测值显然比其他观测值离回归线远得多。对离群点必须极其谨慎地处理。如果离群点确实属于数据集,就不要自动将其删除。但是,要诚实地说明它对分析的影响。处理离群点的一种策略是呈现两次分析:一次包含离群点,一次不包含离群点。此外,当正残差的幅度小于负残差(或反之)时,会出现一种违反正态性的情况。也就是说,当残差不是围绕直线对称分布时, \(y=0.\)

A grid of 2 by 4 scatterplots with fabricated data. The top row of plots contains original x-y data plots with a least squares regression line. The bottom row of plots is a series of residual plot with predicted value on the x-axis and residual on the y-axis. The first column of plots gives an example of points that have a quadratic relationship instead of a linear relationship. The second column of plots gives an example where a single outlying point does not fit the linear model. The third column of points gives an example where the points have increasing variability as the value of x increases. The last column of points gives an example where the points are correlated with one another, possibly as part of a time series.
图 24.15:四个示例,说明本章的方法不足以对数据应用线性模型的情况。上面一组图表示 \(x\)\(y\) 关系。下面一组图是残差图。第一幅图——线性不成立。第二幅图——存在离群点,尤其是一个离直线非常远的点。第三幅图——误差的变异性与 \(x\)的值相关。第四幅图——展示了一个时间序列数据集,其中相邻观测值高度相关。
  • 恒定或等方差性。 各点围绕最小二乘直线的变异性大致保持恒定。非恒定变异性的示例见 图 24.15的第三幅图,它展示了该条件不成立时最常见的模式:当 \(y\) 较大时, \(x\) 的变异性也较大。

对于 图 7.14?3

中的 Elmhurst 数据应用最小二乘回归,我们是否应该有所担忧? 这些技术性条件通常可以通过 LINE(线性、独立、正态、等方差)助记词来记忆。线性、正态性和等方差性条件通常可以通过残差图来评估,如 图 24.15。应仔细考虑实验设计,以确认观测值确实是独立的。

  • L: 线性 模型
  • I: 独立 观测值
  • N: 数据点在直线周围呈 正态 分布
  • E: 相等 对于解释变量的所有取值,直线周围的变异性相等

24.6.2 为什么我们需要技术条件?

与本文中介绍的其他推断方法一样,如果上述技术条件不成立,就不可能对总体做出结论性判断。也就是说,如果没有技术条件,T 分数将不具有所假设的 t 分布。话虽如此,精确检验这些条件几乎总是不可能的,因此我们寻找与这些条件的较大偏差。如果存在较大偏差,我们将无法信任计算出的 p 值或所得置信区间的端点。

基于线性关系的模型

如果您的目标是理解……之间的线性模型,线性条件是最重要的条件之一 \(x\)\(y\). 例如,如果 \(x\)\(y\) 之间的真实关系是二次的,如 图 7.3所示,那么斜率的值将完全没有任何意义。我们不仅应该对推断保持谨慎,而且模型 本身 也不能准确刻画变量之间的关系。然而,关于线性以外的其他函数形式建模方法的不同讨论超出了本书的范围。

独立性的重要性

描述观测值独立性的技术条件往往是最关键的,但也是最难诊断的。此外,收集一个真正来自目标总体随机样本的数据集也极其困难。来自固定个体集合的真正随机化实验则容易实施得多,事实上,如今大多数医学研究都采用随机化实验。

相互依赖的观测值可能会以产生根本性错误分析的方式使结果产生偏差。也就是说,如果你在健身房里测量身高和体重,你的线性模型肯定不能代表你所在大学的所有学生。充其量它只是一个描述使用健身房的学生(而且还是愿意与你交谈的学生、在你去测量时使用健身房的学生等)的模型。

与其试图回答你的观测值是否是真正的随机样本,不如转而关注你是否认为你的观测值能够代表某个目标总体。人类在执行随机程序方面出了名地糟糕,因此你应该警惕任何依赖人类直觉来平衡数据的过程,例如在样本中个体的人口统计特征方面。

关于正态性的一些思考

正态性条件要求数据点围绕直线对称变化,并以钟形方式散布。你应该把正态分布的“钟”想象成位于直线之上(从三维意义上讲,从纸面上凸出来),以此表明数据点在靠近直线处很密集,而随着离直线越来越远则逐渐稀疏。

出于几个原因,正态性条件不如线性或独立性重要。首先,用最小二乘法拟合的线性模型仍然是真实总体模型的无偏估计。然而,该估计量的分布将是未知的。幸运的是,中心极限定理(在 第 19.2 节中描述)告诉我们,只要样本量足够大,使用数学模型(带有 \(t\)分布)进行的大多数分析(例如标准误、p 值、置信区间)仍然成立(即使数据不是围绕直线正态分布的)。有一种分析方法 确实 要求正态性,无论样本大小如何,都是为了在使用线性模型时,构建在给定 \(x\) 值处预测单个观测结果响应的区间。稍微不那么担心正态性的另一个原因是,随机化检验和自助法程序都不要求数据在直线周围呈正态分布。

特别是预测时的等方差性

与正态性一样,等方差性条件(即对于所有 \(x\)值,数据点在直线周围的分散方式相似)不会对线性模型的估计造成问题。话虽如此,如果直线周围的变异性极端不均匀,那么对模型的 推断 (例如计算 p 值)将是不正确的。在 x 值范围内呈现不等方差的数据,有可能严重错误估计斜率的变异性,这将对推断结果(即假设检验和置信区间)产生影响。

在许多情况下,随机化检验和自助法置信区间的推断结果对等方差性条件也是稳健的,因此当数据存在异方差性(即在回归线周围呈现不等的变异性)时,它们为分析者提供了一套可用的方法。尽管随机化检验和自助法使我们能够在更少的条件下分析数据,但本文中描述的所有方法都需要一些技术条件(例如独立观测)。当等方差性条件被违反且需要数学分析(例如来自 T 分数的 p 值)时,还有其他现有方法(超出本文范围)可以处理不等方差(例如加权最小二乘分析)。

24.6.3 如果所有技术条件都满足会怎样?

当技术条件满足时,几乎所有统计软件都提供最小二乘回归模型及其推断。然而,除了普遍可用之外,最小二乘回归模型(及相关推断)的另一个优势是线性模型具有重要的扩展(使用自助法和随机化检验实现这些扩展并非易事)。特别是,随机效应模型、重复测量和交互作用都是需要上述技术条件的线性模型扩展。当技术条件成立时,线性模型的这些扩展可以为手头的数据和研究问题提供重要的洞见。我们将在 第 25 章第26章中讨论一些扩展建模及相关推断。许多用于处理技术条件违反的方法超出了本书的范围,但它们会在大学里紧接本书之后的下一门课程中讲授。如果您正在使用线性模型,或者想了解更多,我们建议您继续学习适用于更广泛数据集类别的统计方法。

24.7 本章复习

24.7.1 小结

回想一下,在本书前文中我们介绍了用于传达多个变量之间关系的图形技术。我们还使用建模来形式化这些关系。许多章节专门讨论了推断方法,这些方法使我们能够基于数据样本对总体做出论断。我们不仅介绍了每种推断技术的数学模型,而且在适当的时候,还介绍了自助法和置换方法。

第24章 我们通过随机化检验、自助法和数学建模,将所有这些想法整合在一起,对线性模型进行推断论断。我们继续强调实验设计在得出研究结论中的重要性。特别是,回想一下变异性可能来自不同的来源(例如随机抽样与随机分配,参见 图 2.8).

24.7.2 术语

本章中介绍的术语列于 表 24.5。如果您不确定其中一些术语的含义,我们建议您回到正文中复习它们的定义。您应该能够很容易地发现它们是 粗体文本.

表 24.5:本章介绍的术语。
斜率的自助法置信区间 斜率的随机化检验 技术条件线性回归
单预测变量回归的推断 斜率的t分布 斜率的变异性

24.8 练习

奇数编号习题的答案见 附录 A.24.

  1. 身体测量,随机化检验。 研究人体测量学的研究人员收集了507名身体活跃个体的身体和骨骼直径测量数据,以及年龄、体重、身高和性别。建立一个线性模型,根据肩围(在三角肌上测量的肩部周长)来预测身高,两者均以厘米为单位测量。4 (Heinz 等人,2003) 下面显示的是根据肩围预测身高的线性模型输出,以及来自1,000个随机化数据集的斜率直方图(1,000次,将 hgt 被打乱并与 sho_gi进行回归)。红色竖线画在线性模型输出中得到的观测斜率值处。

    term 估计 的线性模型摘要 统计量 std.error
    p.value 105.832 3.27 32.3 <0.0001
    sho_gi 0.604 0.03 20.0 <0.0001

    1. 评估根据肩围预测身高的模型斜率是否不同于0的原假设和备择假设是什么。

    2. 使用描述原假设为真时斜率分布的直方图,求出p值,并结合问题的背景(使用肩围和身高之类的词语)对该假设检验作出结论。

    3. 基于随机化斜率直方图的结论与基于数学模型的结论是否一致?请解释你的理由。

  2. 婴儿体重与父亲年龄,随机化检验。 美国卫生与公众服务部、疾病控制与预防中心收集该国记录的出生信息。这里使用的数据是2014年1,000例出生的随机样本。在此,我们研究父亲年龄与婴儿体重之间的关系。5 (ICPSR 2014) 下方展示的是用父亲年龄(单位:岁)预测婴儿体重(单位:磅)的线性模型输出,以及来自1000个随机化数据集(1000次,将 weight 被打乱并与 fage进行回归)。红色竖线画在线性模型输出中得到的观测斜率值处。

    term 估计 的线性模型摘要 统计量 std.error
    p.value 7.101 0.199 35.674 <0.0001
    fage 0.005 0.006 0.757 0.4495

    1. 在检验用父亲年龄预测婴儿体重的模型斜率是否不同于0时,零假设和备择假设分别是什么?

    2. 利用该直方图(它描述了当零假设为真时斜率的分布),求出p值,并结合问题情境(使用诸如父亲年龄和婴儿体重之类的词语)对假设检验作出结论。你的检验结论对于父亲年龄是否是婴儿体重的有用预测变量说明了什么?

    3. 基于随机化斜率直方图的结论与基于数学模型的结论是否一致?请解释你的理由。

  1. 身体测量,数学检验。 下方的散点图和最小二乘法汇总显示了507名经常进行体育锻炼的个体的体重(单位:千克)与身高(单位:厘米)之间的关系。 (Heinz 等人,2003)

    term 估计 的线性模型摘要 统计量 std.error
    p.value -105.01 7.54 -13.9 <0.0001
    hgt 1.02 0.04 23.1 <0.0001
    1. 描述身高与体重之间的关系。

    2. 写出回归直线的方程,并结合具体情境解释斜率和截距。

    3. 这些数据是否提供了令人信服的证据,表明真实的斜率参数不同于0?陈述零假设和备择假设,报告p值(使用数学模型),并给出你的结论。

    4. 身高与体重的相关系数为0.72。计算 \(R^2\) 并在上下文中对其进行解释。

  1. 婴儿体重与父亲年龄,数学检验。 父亲的年龄对预测婴儿体重有用吗?下面的散点图和最小二乘法汇总显示了一个随机婴儿样本中婴儿体重(以磅为单位)与父亲年龄之间的关系。 (ICPSR 2014)

    term 估计 的线性模型摘要 统计量 std.error
    p.value 7.1042 0.1936 36.698 <0.0001
    fage 0.0047 0.0061 0.779 0.4359
    1. 父亲为30岁的婴儿的预测体重是多少?

    2. 数据是否提供了令人信服的证据,表明根据父亲年龄预测婴儿体重的模型的斜率不同于0?陈述原假设和备择假设,报告p值(使用数学模型),并给出你的结论。

    3. 根据你的结论,父亲的年龄是预测婴儿体重的有用预测变量吗?

  1. 身体测量,自助法百分位区间。 为了估计根据肩围(在三角肌上方测量的肩部周长)预测身高的模型的斜率,从507人的身体测量数据集中抽取了1,000个自助样本。对每个自助样本拟合一个根据肩围预测身高的线性模型,并估计斜率。这些斜率的直方图如下所示。 (Heinz 等人,2003)

    1. 使用自助法百分位法和上面的直方图,求斜率参数的98%置信区间。

    2. 在问题的背景下解释该置信区间。

  1. 婴儿体重与父亲年龄,自助法百分位区间。 美国卫生与公众服务部、疾病控制与预防中心收集该国出生记录的信息。这里使用的数据是2014年1,000例出生的随机样本。在此,我们研究父亲年龄与婴儿体重之间的关系。下面是来自该数据的1,000个不同自助样本的斜率统计量的自助分布。 (ICPSR 2014)

    1. 使用自助法百分位法和上面的直方图,求斜率参数的95%置信区间。

    2. 在问题的背景下解释该置信区间。

  1. 身体测量,标准误自助区间。 建立一个线性模型,根据肩围(在三角肌上测量的肩部周长)预测身高,两者均以厘米为单位。 (Heinz 等人,2003) 下面显示的是根据肩围预测身高的线性模型输出,以及来自数据的1000个不同自助样本的斜率统计量的自助分布。

    term 估计 的线性模型摘要 统计量 std.error
    p.value 105.832 3.27 32.3 <0.0001
    sho_gi 0.604 0.03 20.0 <0.0001
    1. 使用直方图,近似估计斜率统计量的标准误(即量化斜率统计量从样本到样本的变异性)。

    2. 求斜率参数的98%自助SE置信区间。

    3. 在问题的背景下解释该置信区间。

     

  1. 婴儿体重与父亲年龄,标准误自助区间。 美国卫生与公众服务部、疾病控制与预防中心收集该国记录的出生信息。这里使用的数据是2014年1000例出生的随机样本。在此,我们研究父亲年龄与婴儿体重之间的关系。 (ICPSR 2014) 下面显示的是根据父亲年龄(以年为单位)预测婴儿体重(以磅为单位)的线性模型输出,以及来自数据的1000个不同自助样本的斜率统计量的自助分布。

    term 估计 的线性模型摘要 统计量 std.error
    p.value 7.101 0.199 35.674 <0.0001
    fage 0.005 0.006 0.757 0.4495
    1. 使用直方图,近似估计斜率统计量的标准误(即量化斜率统计量从样本到样本的变异性)。

    2. 求斜率参数的95%自助SE置信区间。

    3. 在问题的背景下解释该置信区间。

     

  1. 身体测量,条件。 下面的散点图显示了来自507名体力活跃个体身体测量数据集中体重与身高线性模型的残差(在y轴上)。x轴是个体的身高,单位为厘米。 (Heinz 等人,2003)

    1. 对于这些数据, \(R^2\) 为 51.84%。相关系数的值是多少?你如何判断它是正的还是负的?(提示:你可能需要查看之前的练习。)

    2. 检查残差图。你观察到了什么?简单最小二乘拟合是否适用于这些数据?LINE 条件中哪些满足,哪些不满足?

  1. 婴儿体重与父亲年龄,条件。 下面的散点图显示了一组随机抽取的婴儿样本中,婴儿体重(以磅为单位)与父亲年龄的线性模型的残差(在 y 轴上)。x 轴为父亲年龄。 (ICPSR 2014)

    1. 对于这些数据, \(R^2\) 为 0.09%。相关系数的值是多少?你如何判断它是正的还是负的?(提示:你可能需要查看之前的练习。)

    2. 检查残差图。你观察到了什么?简单最小二乘拟合是否适用于这些数据?LINE 条件中哪些满足,哪些不满足?

  1. 谋杀与贫困,随机化检验。 以下回归输出用于根据贫困人口百分比(annual_murders_per_mil)来预测一个由20个大都市区组成的随机样本中每百万人的年谋杀数(perc_pov)。下面展示了根据贫困人口百分比预测大都市地区每百万人年谋杀数的线性模型输出,以及来自 1000 个随机化数据集的斜率直方图(1000 次中, annual_murders_per_mil 被打乱并与 perc_pov进行回归)。红色竖线画在线性模型输出中得到的观测斜率值处。

    term 估计 的线性模型摘要 统计量 std.error
    p.value -29.90 7.79 -3.84 0.0012
    perc_pov 2.56 0.39 6.56 <0.0001
    1. 评估用贫困百分比预测年谋杀率的模型斜率是否不同于0的原假设和备择假设是什么?

    2. 使用描述原假设为真时斜率分布的直方图,求出p值,并在问题背景下(使用“谋杀率”和“贫困”等词语)对假设检验下结论。

    3. 基于随机化斜率直方图的结论与使用数学模型会得到的结论是否一致?请解释你的理由。

     

  2. 谋杀与贫困,数学检验。 下表显示了用生活在贫困中的百分比(annual_murders_per_mil)来预测一个由20个大都市区组成的随机样本中每百万人的年谋杀数(perc_pov)的线性模型在20个大都市地区随机样本中的输出结果。

    term 估计 的线性模型摘要 统计量 std.error
    p.value -29.90 7.79 -3.84 0.0012
    perc_pov 2.56 0.39 6.56 <0.0001
    1. 评估用贫困百分比预测年谋杀率的模型斜率是否不同于0的假设是什么?

    2. 在问题背景下陈述(a)部分假设检验的结论。这说明贫困百分比是否是年谋杀率的有用预测变量?

    3. 计算贫困百分比斜率的95%置信区间,并在问题背景下解释它。

    4. 你从假设检验和置信区间得到的结果是否一致?请解释你的理由。

  1. 谋杀与贫困,自助法百分位区间。 从20个大都市地区的随机样本中收集了每百万人年谋杀数(annual_murders_per_mil)和生活在贫困中的百分比(perc_pov)的数据。使用这些数据,我们想要估计预测模型的斜率。 annual_murders_per_mil 表 25.6:用 perc_pov. 我们从数据中抽取1,000个自助样本,并对每个自助样本拟合一个预测 annual_murders_per_mil 表 25.6:用 perc_pov 的线性模型。这些斜率的直方图如下所示。

    1. 使用百分位自助法及上面的直方图,求斜率参数的90%置信区间。

    2. 在问题的背景下解释该置信区间。

  2. 谋杀与贫困,标准误自助区间。 建立一个线性模型,根据生活在贫困中的百分比(annual_murders_per_mil)来预测一个由20个大都市区组成的随机样本中每百万人的年谋杀数(perc_pov)。下面展示了根据生活在贫困中的百分比预测大都市区每百万人年谋杀数的标准线性模型输出,以及来自数据的1000个不同自助样本的斜率统计量的自助分布。

    term 估计 的线性模型摘要 统计量 std.error
    p.value -29.90 7.79 -3.84 0.0012
    perc_pov 2.56 0.39 6.56 <0.0001

    1. 使用直方图,近似估计斜率统计量的标准误(即量化斜率统计量从样本到样本的变异性)。

    2. 求斜率参数的90%自助标准误(SE)置信区间。

    3. 在问题的背景下解释该置信区间。

  1. 谋杀与贫困,条件。 下面的散点图显示了一个由20个大都市区组成的随机样本中每百万人年谋杀数与生活在贫困中的百分比的关系。第二幅图以残差为y轴、以生活在贫困中的百分比为x轴。

    1. 对于这些数据, \(R^2\) 为70.56%。相关系数的值是多少?如何判断它是正的还是负的?

    2. 检查残差图。你观察到了什么?简单最小二乘拟合是否适合这些数据?LINE条件中哪些满足,哪些不满足?

  1. 我喜欢猫。 研究人员收集了 144 只成年家猫的心脏重量和体重数据。下表显示了用这些猫的体重(以千克为单位)预测心脏重量(以克为单位)的线性模型的输出结果。6

    term 估计 的线性模型摘要 统计量 std.error
    p.value -0.357 0.692 -0.515 0.6072
    Bwt 4.034 0.250 16.119 <0.0001
    1. 评估猫的体重与心脏重量之间是否存在正相关关系的假设是什么?

    2. 结合具体情境陈述 (a) 部分假设检验的结论。

    3. 计算体重斜率的 95% 置信区间,并结合具体情境进行解释。

    4. 你从假设检验和置信区间得到的结果是否一致?请解释你的理由。

  1. 啤酒与血液酒精含量。 许多人认为,在预测血液酒精含量(BAC)时,体重、饮酒习惯以及许多其他因素比仅仅考虑一个人喝了多少杯酒重要得多。在这里,我们考察来自俄亥俄州立大学 16 名学生志愿者的数据,每人喝了随机指定数量的罐装啤酒。这些学生男女各半,他们的体重和饮酒习惯各不相同。三十分钟后,一名警察测量了他们的血液酒精含量(BAC),单位为每分升血液中酒精的克数。散点图和回归表总结了这些发现。 7 (Malkevitch 和 Lesser 2008)

    term 估计 的线性模型摘要 统计量 std.error
    p.value -0.0127 0.0126 -1.00 0.332
    beers 0.0180 0.0024 7.48 <0.0001
    1. 描述啤酒罐数与 BAC 之间的关系。

    2. 写出回归直线的方程,并结合具体情境解释斜率和截距。

    3. 这些数据是否提供了令人信服的证据,表明喝更多罐啤酒与血液酒精含量升高相关?陈述原假设和备择假设,报告 p 值,并给出你的结论。

    4. 啤酒罐数与 BAC 的相关系数为 0.89。计算 \(R^2\) 并在上下文中对其进行解释。

    5. 假设我们在自己所在的城镇去一家酒吧,询问人们喝了几杯酒,并测量他们的血液酒精浓度(BAC)。饮酒杯数与BAC之间的关系是否会像俄亥俄州立大学研究中发现的那样强?为什么?

  1. 城市房主,条件。 下面的散点图显示了拥有住房的家庭百分比与居住在城市地区的人口百分比之间的关系。 (美国人口普查局 2010) 共有52个观测值,每个观测值对应美国的一个州。波多黎各和哥伦比亚特区也包括在内。第二幅图以残差为y轴,以居住在城市地区的人口百分比为x轴。

    1. 对于这些数据, \(R^2\) 为29.16%。相关系数的值是多少?如何判断它是正的还是负的?

    2. 检查残差图。你观察到了什么?简单最小二乘拟合是否适合这些数据?LINE条件中哪些满足,哪些不满足?

  1. 我爱猫,LINE条件。 研究人员收集了144只家养成年猫的心脏重量和体重数据。下图显示了由一个线性模型生成的预测值和残差的输出,该模型根据这些猫的体重(以千克为单位)预测心脏重量(以克为单位)。

    1. 检查残差图。注意,对于较小的预测值,残差的幅值小于较大预测值所对应的较大残差。残差幅值随预测值的变化表明违反了哪一条LINE技术条件?

    2. 如果(a)部分所述的LINE条件被违反,它是否可能导致关于模型本身(即最小二乘回归线)的错误结论、关于模型推断(即与最小二乘回归线相关的p值)的错误结论、两者都不是,还是两者都是?请解释你的理由。

  1. 啤酒与血液酒精含量,LINE条件。 下图显示了由一个线性模型生成的预测值和残差的输出,该模型根据俄亥俄州立大学十六名学生志愿者饮用的啤酒罐数预测血液酒精含量(BAC)。8 (Malkevitch 和 Lesser 2008)

    1. 检查残差图。注意,很难识别出任何令人信服的模式来支持或反驳违反LINE技术条件的说法。残差图的什么特点使得难以评估LINE技术条件?

    2. 残差图中是否有什么会让您在将线性模型用于对所有学生的推断时有所犹豫?该研究的实验设计中是否有什么会让您在将线性模型用于对所有学生的推断时有所犹豫?


  1. 这个问题的答案依赖于这样一个理念:统计分析在某种程度上是一门艺术。也就是说,在许多情况下,并不存在“正确”的答案。当您自己进行越来越多的分析时,您会逐渐认识到针对特定数据集所需的细致入微的理解。就大萧条而言,我们将提供两种相反的考虑。这些数据点中的每一个在任何最小二乘回归线上都会有很高的杠杆作用,而失业率如此之高的年份可能无法帮助我们理解失业率仅适度偏高的其他年份会发生什么。另一方面,大萧条年份是特殊情况,如果将它们从最终分析中排除,我们将丢弃重要信息。↩︎

  2. 我们查看对应于家庭收入变量的第二行。我们看到直线斜率的点估计为 -0.0431,该估计的标准误为 0.0108, \(t\)-检验统计量为 \(T = -3.98\)。p值恰好对应于我们感兴趣的双侧检验:0.0002。p值非常小,因此我们拒绝原假设,并得出结论:2011年入学的Elmhurst学院新生的家庭收入与经济资助呈负相关,且真实斜率参数确实小于0,正如我们在对 图 7.13.↩︎

  3. 该趋势看起来是线性的,数据分布在直线附近且没有明显的离群点,方差大致恒定。数据并非来自时间序列,也不存在其他明显违反独立性的情况。最小二乘回归可以应用于这些数据。↩︎

  4. bdims 本练习中使用的数据可在 openintro R 包中找到。↩︎

  5. births14 本练习中使用的数据可在 openintro R 包中找到。↩︎

  6. cats 本练习中使用的数据可在 MASS R 包中找到。↩︎

  7. bac 本练习中使用的数据可在 openintro R 包中找到。↩︎

  8. bac 本练习中使用的数据可在 openintro R 包中找到。↩︎