--- title: "Comparing Groups and Testing Hypotheses" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Comparing Groups and Testing Hypotheses} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", message = FALSE, warning = FALSE ) ``` ```{r setup} library(mariposa) library(dplyr) data(survey_data) ``` ## Overview Statistical tests determine whether observed differences between groups are real or due to random chance. This guide covers all hypothesis tests in mariposa, organized by the type of data and research question. ### Choosing the Right Test | Your data | 2 groups | 3+ groups | Paired | |-----------|----------|-----------|--------| | **Continuous, normal** | `t_test()` | `oneway_anova()` | *paired t-test planned* | | **Continuous, non-normal** | `mann_whitney()` | `kruskal_wallis()` | `wilcoxon_test()` | | **Categorical** | `chi_square()` | `chi_square()` | `mcnemar_test()` | | **Small sample, categorical** | `fisher_test()` | --- | --- | | **Multiple factors** | --- | `factorial_anova()` | `friedman_test()` | | **With covariate** | --- | `ancova()` | --- | | **Proportion vs. expected** | `binomial_test()` | `chisq_gof()` | --- | ### Checking Normality First The parametric tests in the first row assume (approximately) normally distributed variables. `normality_test()` produces the SPSS EXAMINE "Tests of Normality" table — Kolmogorov-Smirnov with Lilliefors correction and Shapiro-Wilk: ```{r} survey_data %>% group_by(gender) %>% normality_test(age, income) ``` Run it grouped (as above) when you are about to compare groups: the assumption that matters is normality *within* each group. With large samples these tests flag even trivial deviations, so combine them with the skewness and kurtosis from `describe()` before switching to a non-parametric alternative. ## t-Tests ### Independent Samples Compare two groups on a continuous variable: ```{r} survey_data %>% t_test(life_satisfaction, group = gender, weights = sampling_weight) ``` The output includes both Student's t-test (equal variances assumed) and Welch's t-test (not assumed). When in doubt, use Welch --- it is more robust. For the detailed output with group descriptives, Levene's test, and confidence intervals: ```{r} survey_data %>% t_test(life_satisfaction, group = gender, weights = sampling_weight) %>% summary() ``` ### Multiple Variables at Once ```{r} survey_data %>% t_test(trust_government, trust_media, trust_science, group = gender, weights = sampling_weight) ``` ### One-Sample t-Test Test whether a mean differs from a specific value: ```{r} # Is average life satisfaction different from the scale midpoint (3)? survey_data %>% t_test(life_satisfaction, mu = 3, weights = sampling_weight) ``` ### Grouped Analysis Run separate tests per subgroup: ```{r} survey_data %>% group_by(region) %>% t_test(income, group = gender, weights = sampling_weight) ``` ## One-Way ANOVA Compare means across three or more groups: ```{r} result <- survey_data %>% oneway_anova(life_satisfaction, group = education, weights = sampling_weight) result ``` The effect size $\eta^2$ (eta-squared) indicates how much variance is explained by group membership: - **Small**: $\eta^2 \approx 0.01$ - **Medium**: $\eta^2 \approx 0.06$ - **Large**: $\eta^2 \approx 0.14$ ### Post-Hoc Tests A significant ANOVA tells you that groups differ, but not *which* groups. Use post-hoc tests: ```{r} # Tukey HSD: balanced comparison of all pairs tukey_test(result) ``` ```{r} # Scheffe: more conservative (fewer false positives) scheffe_test(result) ``` ### Assumption Check ANOVA assumes equal variances. Test with Levene's test: ```{r} levene_test(result) ``` If Levene's test is significant ($p < .05$), variances are unequal. Use the Welch correction included in the ANOVA output. ## Factorial ANOVA Test the effects of two or more factors and their interactions: ```{r} survey_data %>% factorial_anova(dv = income, between = c(gender, education), weights = sampling_weight) ``` The output uses Type III sums of squares and reports partial $\eta^2$ for each effect. Weighted analysis uses WLS estimation, matching SPSS UNIANOVA. For the full output with descriptive statistics per cell: ```{r} survey_data %>% factorial_anova(dv = life_satisfaction, between = c(gender, region), weights = sampling_weight) %>% summary() ``` ## ANCOVA Compare groups while controlling for a covariate: ```{r} survey_data %>% ancova(dv = income, between = education, covariate = age, weights = sampling_weight) ``` The output includes the covariate effect, the adjusted factor effect, and estimated marginal means (group means adjusted for the covariate). ## Non-Parametric Tests Use these when data is not normally distributed, ordinal, or based on small samples. ### Mann-Whitney U Test The non-parametric alternative to the independent t-test: ```{r} survey_data %>% mann_whitney(political_orientation, group = region, weights = sampling_weight) ``` ### Kruskal-Wallis H Test The non-parametric alternative to one-way ANOVA (3+ groups): ```{r} kw_result <- survey_data %>% kruskal_wallis(life_satisfaction, group = education) kw_result ``` When significant, use Dunn's post-hoc test with Bonferroni correction: ```{r} dunn_test(kw_result) ``` ### Wilcoxon Signed-Rank Test The non-parametric alternative to the paired t-test: ```{r} data(longitudinal_data_wide) longitudinal_data_wide %>% wilcoxon_test(score_T1, score_T2) ``` ### Friedman Test The non-parametric alternative to repeated-measures ANOVA (3+ measurements): ```{r} friedman_result <- longitudinal_data_wide %>% friedman_test(score_T1, score_T2, score_T3) friedman_result ``` When significant, use pairwise Wilcoxon post-hoc tests: ```{r} pairwise_wilcoxon(friedman_result) ``` ### Binomial Test Test whether an observed proportion differs from an expected value: ```{r} survey_data %>% binomial_test(gender) ``` ## Categorical Tests ### Chi-Square Test of Independence Test whether two categorical variables are related: ```{r} survey_data %>% chi_square(education, employment, weights = sampling_weight) ``` A significant result means the variables are not independent --- knowing one tells you something about the other. ### Effect Sizes for Categorical Data The helpers `phi()`, `cramers_v()`, and `goodman_gamma()` run the chi-square analysis internally and return just the requested effect size as a number (per group for grouped data). For the full test output, call `chi_square()` directly. ```{r} # Phi coefficient (2x2 tables) survey_data %>% phi(gender, employment, weights = sampling_weight) ``` ```{r} # Cramer's V (larger tables) survey_data %>% cramers_v(education, employment, weights = sampling_weight) ``` ### Fisher's Exact Test Use when expected cell frequencies are below 5: ```{r} small_sample <- survey_data %>% slice_sample(n = 30) small_sample %>% fisher_test(gender, region) ``` ### Chi-Square Goodness-of-Fit Test whether observed frequencies match expected proportions: ```{r} # Equal proportions (default) survey_data %>% chisq_gof(education) ``` ```{r} # Custom expected proportions survey_data %>% chisq_gof(education, expected = c(0.30, 0.25, 0.25, 0.20)) ``` ### McNemar's Test Compare paired proportions (e.g., before/after): ```{r, eval = FALSE} test_data <- survey_data %>% mutate( trust_gov_high = ifelse(trust_government > 3, 1, 0), trust_media_high = ifelse(trust_media > 3, 1, 0) ) test_data %>% mcnemar_test(var1 = trust_gov_high, var2 = trust_media_high) ``` ## Interpreting Results ### p-Values - $p < .05$: The difference is statistically significant - $p \geq .05$: No significant difference detected "Not significant" does **not** mean "no difference" --- it means we cannot rule out chance given the sample size. ### Effect Sizes With large samples, even tiny differences can be significant. Always check effect sizes: | Test | Effect size | Small | Medium | Large | |------|-------------|-------|--------|-------| | t-test | Cohen's *d* | 0.20 | 0.50 | 0.80 | | ANOVA | $\eta^2$ | 0.01 | 0.06 | 0.14 | | Chi-square | Cramer's *V* | 0.10 | 0.30 | 0.50 | | Correlation | *r* | 0.10 | 0.30 | 0.50 | ### Multiple Comparisons Running many tests inflates false positive rates. Post-hoc tests (`tukey_test()`, `dunn_test()`, `pairwise_wilcoxon()`) handle this automatically with corrections. ## Complete Example A typical hypothesis testing workflow: ```{r} # 1. Describe the groups survey_data %>% group_by(education) %>% describe(life_satisfaction, weights = sampling_weight) # 2. Test for overall differences anova_result <- survey_data %>% oneway_anova(life_satisfaction, group = education, weights = sampling_weight) anova_result # 3. Check assumptions levene_test(anova_result) # 4. Post-hoc: which groups differ? tukey_test(anova_result) ``` ## Practical Tips 1. **Check assumptions first.** Use `describe(show = "all")` to inspect skewness. For non-normal data, use non-parametric tests. 2. **Match the test to the data.** Normal continuous data: t-test / ANOVA. Non-normal or ordinal: Mann-Whitney / Kruskal-Wallis. Categorical: chi-square / Fisher. 3. **Always follow up significant omnibus tests.** Use `tukey_test()` for ANOVA, `dunn_test()` for Kruskal-Wallis, `pairwise_wilcoxon()` for Friedman. 4. **Report effect sizes alongside p-values.** A significant result with a negligible effect size may not be practically meaningful. 5. **Use weights when available.** They ensure results represent the population, not just the sample. ## Summary ### Parametric Tests - **`t_test()`** compares means between two groups - **`oneway_anova()`** extends to three or more groups, with `tukey_test()` / `scheffe_test()` post-hoc - **`factorial_anova()`** tests multiple factors and interactions - **`ancova()`** controls for a covariate ### Non-Parametric Tests - **`mann_whitney()`**, **`kruskal_wallis()`** (with `dunn_test()`), **`wilcoxon_test()`**, **`friedman_test()`** (with `pairwise_wilcoxon()`), **`binomial_test()`** ### Categorical Tests - **`chi_square()`**, **`fisher_test()`**, **`chisq_gof()`**, **`mcnemar_test()`** ### Effect Sizes - **`phi()`**, **`cramers_v()`**, **`goodman_gamma()`** ## Next Steps - Measure relationships between continuous variables --- see `vignette("correlation-analysis")` - Build predictive models --- see `vignette("regression-analysis")` - Construct reliable scales --- see `vignette("scale-analysis")`