---
title: "Descriptive Statistics and Frequencies"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Descriptive Statistics and Frequencies}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  message = FALSE,
  warning = FALSE
)
```

```{r setup}
library(mariposa)
library(dplyr)
data(survey_data)
```

## Overview

Before running any statistical tests, you should understand your data. Descriptive statistics summarize what is typical, what is unusual, and how responses are distributed.

mariposa provides four functions for data exploration:

| Function | Best for |
|----------|----------|
| `describe()` | Numeric variables --- means, medians, spread |
| `frequency()` | Categorical variables --- counts and percentages |
| `crosstab()` | Relationships between two categorical variables |
| `codebook()` | Full data dictionary with types, labels, and values |

All four support survey weights and grouped analysis.

## Summary Statistics with describe()

### Basic Usage

```{r}
survey_data %>%
  describe(age)
```

### Multiple Variables

```{r}
survey_data %>%
  describe(age, income, life_satisfaction)
```

### With Survey Weights

Add `weights` for population-representative statistics:

```{r}
survey_data %>%
  describe(age, income, life_satisfaction, weights = sampling_weight)
```

### Grouped Analysis

Compare statistics across subgroups:

```{r}
survey_data %>%
  group_by(region) %>%
  describe(income, life_satisfaction, weights = sampling_weight)
```

### Choosing Statistics

The `show` argument controls which statistics are displayed:

```{r}
# Just the essentials
survey_data %>%
  describe(age, income, show = c("mean", "sd", "range"))
```

```{r}
# Everything available
survey_data %>%
  describe(age, show = "all")
```

### Understanding the Output

- **n**: Number of valid (non-missing) cases
- **mean** ($\bar{x}$): The arithmetic average
- **sd**: Standard deviation --- how spread out the values are
- **median**: The middle value (50th percentile)
- **min / max**: The range of observed values
- **skewness**: Distribution symmetry. Values near 0 indicate symmetry; values beyond $\pm 1$ indicate notable skew
- **kurtosis**: Tail heaviness. Values near 0 indicate a normal-like shape

## Frequency Tables with frequency()

### Basic Frequency

```{r}
survey_data %>%
  frequency(education)
```

### Understanding the Output

- **Frequency**: Number of respondents in each category
- **Percent**: Percentage of all cases, including missing
- **Valid Percent**: Percentage of non-missing cases only
- **Cumulative Percent**: Running total of valid percentages

Use *Valid Percent* when missing values are truly missing (e.g., skip patterns). Use *Percent* when non-response is itself meaningful.

### Weighted Frequencies

```{r}
survey_data %>%
  frequency(education, weights = sampling_weight)
```

### Multiple Variables

```{r}
survey_data %>%
  frequency(education, employment, region, weights = sampling_weight)
```

### Grouped Frequencies

```{r}
survey_data %>%
  group_by(gender) %>%
  frequency(education, weights = sampling_weight)
```

## Cross-Tabulation with crosstab()

### Basic Crosstab

Examine the relationship between two categorical variables:

```{r}
survey_data %>%
  crosstab(education, employment)
```

### Understanding the Output

- **Count**: Number of cases in each cell
- **Row %**: Percentage within each row (sums to 100% across)
- **Column %**: Percentage within each column (sums to 100% down)
- **Cell %**: Percentage of the total sample

Row percentages answer: "Of people with this education level, what proportion has each employment status?"

Column percentages answer: "Of people with this employment status, what proportion has each education level?"

### Weighted Crosstabs

```{r}
survey_data %>%
  crosstab(education, employment, weights = sampling_weight)
```

### Grouped Crosstabs

```{r}
survey_data %>%
  group_by(region) %>%
  crosstab(education, employment, weights = sampling_weight)
```

## Multiple Response Sets with multiple_response()

"Check all that apply" questions arrive as a series of 0/1 indicator
variables. `multiple_response()` (SPSS `MULT RESPONSE`) analyzes them as
one set, with the two percentage bases SPSS users expect: *percent of
responses* (sums to 100%) and *percent of cases* (can sum above 100%,
because respondents tick several boxes):

```{r}
# Example set: institutions a respondent trusts highly (rating 4-5)
trust_set <- survey_data %>%
  mutate(
    gov     = as.integer(trust_government >= 4),
    media   = as.integer(trust_media >= 4),
    science = as.integer(trust_science >= 4)
  )

trust_set %>%
  multiple_response(gov, media, science, weights = sampling_weight) %>%
  summary()
```

Cross the set against a demographic with `by` — mentions and case-based
column percentages per group:

```{r}
trust_set %>%
  multiple_response(gov, media, science, by = gender,
                    weights = sampling_weight) %>%
  summary()
```

## Data Dictionary with codebook()

For a comprehensive overview of your entire dataset, use `codebook()`:

```{r, eval = FALSE}
codebook(survey_data)
```

This opens an interactive HTML view in the RStudio Viewer showing variable names, types, labels, value distributions, and missing data patterns. It is the fastest way to orient yourself in a new dataset.

## Complete Example

A typical descriptive analysis workflow:

```{r}
# 1. Explore the dataset
find_var(survey_data, "trust|satisfaction")

# 2. Summarize numeric variables
survey_data %>%
  describe(age, income, life_satisfaction,
           weights = sampling_weight)

# 3. Check categorical distributions
survey_data %>%
  frequency(education, employment,
            weights = sampling_weight)

# 4. Cross-tabulate key relationships
survey_data %>%
  crosstab(education, employment,
           weights = sampling_weight)

# 5. Compare across regions
survey_data %>%
  group_by(region) %>%
  describe(income, life_satisfaction,
           weights = sampling_weight)
```

## Practical Tips

1. **Start with descriptives before testing.** Understanding distributions, ranges, and missing data patterns prevents surprises in later analyses.

2. **Check for impossible values.** Negative ages, incomes above plausible limits, or out-of-range Likert responses indicate data quality issues.

3. **Look at skewness.** Values beyond $\pm 1$ suggest the distribution departs notably from normality. This affects the choice between parametric and non-parametric tests.

4. **Always use weights when available.** Unweighted statistics describe your sample; weighted statistics estimate the population.

5. **Consider your audience.** For technical readers, include SD, skewness, and kurtosis. For general audiences, focus on means, medians, and percentages.

## Summary

1. **`describe()`** summarizes numeric variables (mean, SD, median, range, skewness)
2. **`frequency()`** counts categories with percent, valid percent, and cumulative percent
3. **`crosstab()`** shows the joint distribution of two categorical variables
4. **`codebook()`** provides an interactive HTML data dictionary
5. All four functions support **survey weights** and **`group_by()`**

## Next Steps

- Test for significant differences --- see `vignette("hypothesis-testing")`
- Measure relationships between variables --- see `vignette("correlation-analysis")`
- Learn about survey weights --- see `vignette("survey-weights")`
