--- title: "Working with Survey Weights" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Working with Survey Weights} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", message = FALSE, warning = FALSE ) ``` ```{r setup} library(mariposa) library(dplyr) data(survey_data) ``` ## Overview Survey weights correct for sampling biases so your results represent the target population, not just your sample. Common reasons for weighting include oversampling of small groups, differential non-response, and post-stratification to known demographics. In mariposa, adding weights is as simple as `weights = sampling_weight` in any function call. This guide covers the dedicated `w_*` functions, the effective sample size concept, and best practices. ## The Impact of Weights ```{r} # Without weights (describes the sample) unweighted <- survey_data %>% summarise( mean_age = mean(age, na.rm = TRUE), mean_income = mean(income, na.rm = TRUE) ) # With weights (represents the population) age_w <- w_mean(survey_data, age, weights = sampling_weight) income_w <- w_mean(survey_data, income, weights = sampling_weight) comparison <- data.frame( Variable = c("Age", "Income"), Unweighted = c(unweighted$mean_age, unweighted$mean_income), Weighted = c(age_w$results$weighted_mean, income_w$results$weighted_mean) ) print(comparison) ``` The difference between weighted and unweighted results reflects how biased your sample is. Larger differences mean weights are doing more important work. ## Weighted Statistics Functions mariposa provides 11 `w_*` functions for individual weighted statistics. ### Central Tendency ```{r} # Weighted mean w_mean(survey_data, income, weights = sampling_weight) # Weighted median w_median(survey_data, income, weights = sampling_weight) # Weighted mode w_modus(survey_data, education, weights = sampling_weight) ``` ### Dispersion ```{r} # Weighted standard deviation w_sd(survey_data, income, weights = sampling_weight) # Weighted variance w_var(survey_data, income, weights = sampling_weight) # Weighted interquartile range w_iqr(survey_data, income, weights = sampling_weight) ``` ### Distribution Shape ```{r} # Weighted skewness w_skew(survey_data, income, weights = sampling_weight) # Weighted kurtosis w_kurtosis(survey_data, income, weights = sampling_weight) ``` ### Precision and Range ```{r} # Weighted standard error w_se(survey_data, income, weights = sampling_weight) # Weighted quantiles w_quantile(survey_data, income, probs = c(0.25, 0.5, 0.75), weights = sampling_weight) # Weighted range w_range(survey_data, income, weights = sampling_weight) ``` ### Multiple Variables All `w_*` functions accept multiple variables: ```{r} w_mean(survey_data, age, income, life_satisfaction, weights = sampling_weight) ``` ## Effective Sample Size Weighting reduces statistical precision. The *effective sample size* ($n_{eff}$) tells you how much information your weighted sample actually carries: $$n_{eff} = \frac{\left(\sum w_i\right)^2}{\sum w_i^2}$$ ```{r} actual_n <- nrow(survey_data) age_w <- w_mean(survey_data, age, weights = sampling_weight) effective_n <- age_w$results$effective_n cat("Actual sample size:", actual_n, "\n") cat("Effective sample size:", round(effective_n), "\n") cat("Design effect:", round(actual_n / effective_n, 2), "\n") ``` A design effect of 1.0 means weights have no impact on precision. Values above 1 mean larger standard errors than an unweighted sample of the same size. ## Weights in Analysis Functions Every analysis function in mariposa supports the `weights` argument: ```{r} # Descriptive statistics survey_data %>% describe(income, life_satisfaction, weights = sampling_weight) ``` ```{r} # Frequency tables survey_data %>% frequency(education, weights = sampling_weight) ``` ```{r} # t-test survey_data %>% t_test(income, group = gender, weights = sampling_weight) ``` ```{r} # ANOVA survey_data %>% oneway_anova(life_satisfaction, group = education, weights = sampling_weight) ``` ```{r} # Chi-square survey_data %>% chi_square(education, employment, weights = sampling_weight) ``` ```{r} # Correlation survey_data %>% pearson_cor(age, income, weights = sampling_weight) ``` ```{r} # Regression survey_data %>% linear_regression(life_satisfaction ~ age + income, weights = sampling_weight) ``` ## Grouped Weighted Analysis All functions work seamlessly with `group_by()`: ```{r} survey_data %>% group_by(region) %>% describe(age, income, life_satisfaction, weights = sampling_weight) ``` ```{r} survey_data %>% group_by(region) %>% t_test(income, group = gender, weights = sampling_weight) ``` ## Diagnosing Weight Issues ### Checking Weight Distribution ```{r} weight_stats <- survey_data %>% summarise( min = min(sampling_weight, na.rm = TRUE), max = max(sampling_weight, na.rm = TRUE), mean = mean(sampling_weight, na.rm = TRUE), sd = sd(sampling_weight, na.rm = TRUE) ) print(weight_stats) ``` A rule of thumb: weights should range from about 0.2 to 5. Extreme values (below 0.1 or above 10) may indicate problems with the weighting scheme. ### Missing Weights ```{r} missing <- sum(is.na(survey_data$sampling_weight)) total <- nrow(survey_data) cat("Missing weights:", missing, "/", total, "(", round(missing / total * 100, 1), "%)\n") ``` ### Weight Trimming If weights are extreme, trim them to reduce variance at the cost of a small increase in bias: ```{r} trimmed <- survey_data %>% mutate( weight_trimmed = case_when( sampling_weight > 5 ~ 5, sampling_weight < 0.2 ~ 0.2, TRUE ~ sampling_weight ) ) original <- w_mean(survey_data, income, weights = sampling_weight) trimmed_result <- w_mean(trimmed, income, weights = weight_trimmed) cat("Original weighted mean:", round(original$results$weighted_mean, 1), "\n") cat("Trimmed weighted mean:", round(trimmed_result$results$weighted_mean, 1), "\n") ``` ## Complete Example ```{r} # 1. Inspect weight properties survey_data %>% summarise( n = n(), mean_weight = mean(sampling_weight, na.rm = TRUE), sd_weight = sd(sampling_weight, na.rm = TRUE), min_weight = min(sampling_weight, na.rm = TRUE), max_weight = max(sampling_weight, na.rm = TRUE) ) # 2. Compare weighted vs unweighted vars <- c("age", "income", "life_satisfaction") for (var in vars) { unw <- mean(survey_data[[var]], na.rm = TRUE) w_result <- w_mean(survey_data, !!sym(var), weights = sampling_weight) w <- w_result$results$weighted_mean cat(sprintf("%s: unweighted = %.2f, weighted = %.2f (diff = %.1f%%)\n", var, unw, w, (w - unw) / unw * 100)) } # 3. Weighted group comparison survey_data %>% group_by(region) %>% describe(income, weights = sampling_weight) # 4. Weighted hypothesis test survey_data %>% t_test(income, group = gender, weights = sampling_weight) ``` ## Best Practices 1. **Always use weights when they are provided.** They correct for sampling biases and make your results generalizable. 2. **Report both weighted and unweighted sample sizes.** The actual *n* tells readers about data collection; the effective *n* tells them about statistical precision. 3. **Check for extreme weights.** Large variation inflates standard errors and reduces statistical power. Consider trimming if necessary. 4. **Document the weight source.** Explain how weights were calculated and what they adjust for (e.g., age, gender, region post-stratification). 5. **Report the effective sample size.** This tells readers how much precision was lost due to weighting. ## Summary 1. Survey weights make results **representative of the population**, not just the sample 2. The 11 `w_*` functions provide individual weighted statistics 3. **All analysis functions** support weights via `weights =` 4. The **effective sample size** measures how much precision weighting costs 5. **Check and document** weight distributions, extreme values, and missing weights ## Next Steps - Get started with the package --- see `vignette("introduction")` - Explore your data --- see `vignette("descriptive-statistics")` - Test hypotheses --- see `vignette("hypothesis-testing")` - Import weighted data from SPSS --- see `vignette("data-io")`