Clustering standard errors affects statistical inference in regression models
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
1 source for · 0 against
Evidence demonstrates that failing to account for clustering in sample designs when calculating standard errors leads to incorrect statistical inferences.
Most large national surveys, such as the National Survey of Families and Households (NSFH), involve clustered and stratified samples. These complex sample designs have consequences for data analysis techniques. Standard errors calculated using procedures that do not adjust for design effects often are too small and lead to incorrect inferences. We discuss design effects and estimate them for a set of variables selected from the 1988 NSFH. Included are examples of descriptive estimates and regression results with household income and marital happiness as dependent variables. Statistical software that adjusts standard errors in complex designs is discussed, as are issues related to weighting and the analysis of subsamples. As family researchers increase their use of data collected in large, complex personal-interview surveys, such as the NSFH, there is a need for greater awareness of the ways that the sampling design affects the analysis of the data and statistical inferences. Statistical techniques and the standard software (e.g., SPSS, SAS) used by most family researchers make the assumption that the data were collected by simple random sampling. Most large-scale personal-interview surveys, for reasons of efficiency and economy, use probability sampling designs that are not simple random samples. Stratification, clustering, and differential case weighting are common in the design of these multistage probability samples, all of which have consequences for statistical inference