Before running any statistical tests, you should understand your data. Descriptive statistics summarize what is typical, what is unusual, and how responses are distributed.
mariposa provides four functions for data exploration:
| Function | Best for |
|---|---|
describe() |
Numeric variables — means, medians, spread |
frequency() |
Categorical variables — counts and percentages |
crosstab() |
Relationships between two categorical variables |
codebook() |
Full data dictionary with types, labels, and values |
All four support survey weights and grouped analysis.
survey_data %>%
describe(age, income, life_satisfaction)
#>
#> Descriptive Statistics
#> ----------------------
#> Variable Mean Median SD Range IQR Skewness N Missing
#> age 50.550 50 16.976 77 24 0.172 2500 0
#> income 3753.934 3500 1432.802 7200 1900 0.730 2186 314
#> life_satisfaction 3.628 4 1.153 4 2 -0.501 2421 79
#> ----------------------------------------Add weights for population-representative
statistics:
survey_data %>%
describe(age, income, life_satisfaction, weights = sampling_weight)
#>
#> Weighted Descriptive Statistics
#> -------------------------------
#> Variable Mean Median SD Range IQR Skewness Effective_N
#> age 50.514 50 17.084 77 25 0.159 2468.8
#> income 3743.099 3500 1423.966 7200 1900 0.725 2158.9
#> life_satisfaction 3.625 4 1.152 4 2 -0.499 2390.9
#> ----------------------------------------Compare statistics across subgroups:
survey_data %>%
group_by(region) %>%
describe(income, life_satisfaction, weights = sampling_weight)
#>
#> Weighted Descriptive Statistics
#> -------------------------------
#>
#> Group: region = East
#> --------------------
#> ----------------------------------------
#> Variable Mean Median SD Range IQR Skewness Effective_N
#> income 3760.687 3600 1388.321 7200 1700 0.721 421.9
#> life_satisfaction 3.623 4 1.203 4 2 -0.558 457.4
#> ----------------------------------------
#>
#> Group: region = West
#> --------------------
#> ----------------------------------------
#> Variable Mean Median SD Range IQR Skewness Effective_N
#> income 3738.586 3500 1433.325 7200 1900 0.727 1738.1
#> life_satisfaction 3.625 4 1.139 4 2 -0.481 1934.8
#> ----------------------------------------The show argument controls which statistics are
displayed:
# Just the essentials
survey_data %>%
describe(age, income, show = c("mean", "sd", "range"))
#>
#> Descriptive Statistics
#> ----------------------
#> Variable Mean SD Range N Missing
#> age 50.550 16.976 77 2500 0
#> income 3753.934 1432.802 7200 2186 314
#> ----------------------------------------# Everything available
survey_data %>%
describe(age, show = "all")
#>
#> Descriptive Statistics
#> ----------------------
#> Variable Mean Median SD SE Range IQR Skewness Kurtosis Variance Mode
#> age 50.55 50 16.976 0.34 77 24 0.172 -0.364 288.185 18
#> Q25 Q50 Q75 N Missing
#> 38 50 62 2500 0
#> ----------------------------------------Use Valid Percent when missing values are truly missing (e.g., skip patterns). Use Percent when non-response is itself meaningful.
survey_data %>%
frequency(education, employment, region, weights = sampling_weight)
#> Frequency: education [Weighted]
#> 4 categories, N valid = 2516, missing = 0
#> Frequency: employment [Weighted]
#> 5 categories, N valid = 2516, missing = 0
#> Frequency: region [Weighted]
#> 2 categories, N valid = 2516, missing = 0
#> Use summary() for detailed output.Examine the relationship between two categorical variables:
Row percentages answer: “Of people with this education level, what proportion has each employment status?”
Column percentages answer: “Of people with this employment status, what proportion has each education level?”
survey_data %>%
group_by(region) %>%
crosstab(education, employment, weights = sampling_weight)
#> Crosstab: education x employment [Weighted]
#> [region = East] 4 x 5 table, N = 509
#> [region = West] 4 x 5 table, N = 2007
#> Note: no significance test included - use chi_square() for a test of independence.
#> Use summary() for detailed output.“Check all that apply” questions arrive as a series of 0/1 indicator
variables. multiple_response() (SPSS
MULT RESPONSE) analyzes them as one set, with the two
percentage bases SPSS users expect: percent of responses (sums
to 100%) and percent of cases (can sum above 100%, because
respondents tick several boxes):
# Example set: institutions a respondent trusts highly (rating 4-5)
trust_set <- survey_data %>%
mutate(
gov = as.integer(trust_government >= 4),
media = as.integer(trust_media >= 4),
science = as.integer(trust_science >= 4)
)
trust_set %>%
multiple_response(gov, media, science, weights = sampling_weight) %>%
summary()
#>
#> Weighted Multiple Response Results
#> ----------------------------------
#> - Set: gov, media, science
#> - Counted value: 1
#> - Weights Variable: sampling_weight
#>
#> Frequencies
#> ---------------------------------------------
#> Option Responses n Responses % % of Cases
#> ---------------------------------------------
#> gov 585.8 22.9 23.3
#> media 476.5 18.7 18.9
#> science 1490.3 58.4 59.2
#> ---------------------------------------------
#> Valid cases: 2516 | Total responses: 2553 | Excluded (all missing): 0
#> % of Cases can sum above 100% (multiple mentions per case).Cross the set against a demographic with by — mentions
and case-based column percentages per group:
trust_set %>%
multiple_response(gov, media, science, by = gender,
weights = sampling_weight) %>%
summary()
#>
#> Weighted Multiple Response Results
#> ----------------------------------
#> - Set: gov, media, science
#> - Counted value: 1
#> - Weights Variable: sampling_weight
#> - By: gender
#>
#> Frequencies
#> ---------------------------------------------
#> Option Responses n Responses % % of Cases
#> ---------------------------------------------
#> gov 585.8 22.9 23.3
#> media 476.5 18.7 18.9
#> science 1490.3 58.4 59.2
#> ---------------------------------------------
#> Valid cases: 2516 | Total responses: 2553 | Excluded (all missing): 0
#> % of Cases can sum above 100% (multiple mentions per case).
#>
#> Crosstab: set BY gender (% of cases per column)
#> ---------------------------------
#> Option Male Female
#> ---------------------------------
#> gov 280 (23.4%) 306 (23.2%)
#> media 214 (17.9%) 262 (19.8%)
#> science 694 (58.1%) 797 (60.3%)
#> ---------------------------------
#> Cases per column - Male: 1195, Female: 1321For a comprehensive overview of your entire dataset, use
codebook():
This opens an interactive HTML view in the RStudio Viewer showing variable names, types, labels, value distributions, and missing data patterns. It is the fastest way to orient yourself in a new dataset.
A typical descriptive analysis workflow:
# 1. Explore the dataset
find_var(survey_data, "trust|satisfaction")
#> col name label
#> 1 10 life_satisfaction Life satisfaction (1=dissatisfied, 5=satisfied)
#> 2 11 trust_government Trust in government (1=none, 5=complete)
#> 3 12 trust_media Trust in media (1=none, 5=complete)
#> 4 13 trust_science Trust in science (1=none, 5=complete)
# 2. Summarize numeric variables
survey_data %>%
describe(age, income, life_satisfaction,
weights = sampling_weight)
#>
#> Weighted Descriptive Statistics
#> -------------------------------
#> Variable Mean Median SD Range IQR Skewness Effective_N
#> age 50.514 50 17.084 77 25 0.159 2468.8
#> income 3743.099 3500 1423.966 7200 1900 0.725 2158.9
#> life_satisfaction 3.625 4 1.152 4 2 -0.499 2390.9
#> ----------------------------------------
# 3. Check categorical distributions
survey_data %>%
frequency(education, employment,
weights = sampling_weight)
#> Frequency: education [Weighted]
#> 4 categories, N valid = 2516, missing = 0
#> Frequency: employment [Weighted]
#> 5 categories, N valid = 2516, missing = 0
#> Use summary() for detailed output.
# 4. Cross-tabulate key relationships
survey_data %>%
crosstab(education, employment,
weights = sampling_weight)
#> Crosstab: education x employment [Weighted]
#> 4 x 5 table, N = 2516
#> Note: no significance test included - use chi_square() for a test of independence.
#> Use summary() for detailed output.
# 5. Compare across regions
survey_data %>%
group_by(region) %>%
describe(income, life_satisfaction,
weights = sampling_weight)
#>
#> Weighted Descriptive Statistics
#> -------------------------------
#>
#> Group: region = East
#> --------------------
#> ----------------------------------------
#> Variable Mean Median SD Range IQR Skewness Effective_N
#> income 3760.687 3600 1388.321 7200 1700 0.721 421.9
#> life_satisfaction 3.623 4 1.203 4 2 -0.558 457.4
#> ----------------------------------------
#>
#> Group: region = West
#> --------------------
#> ----------------------------------------
#> Variable Mean Median SD Range IQR Skewness Effective_N
#> income 3738.586 3500 1433.325 7200 1900 0.727 1738.1
#> life_satisfaction 3.625 4 1.139 4 2 -0.481 1934.8
#> ----------------------------------------Start with descriptives before testing. Understanding distributions, ranges, and missing data patterns prevents surprises in later analyses.
Check for impossible values. Negative ages, incomes above plausible limits, or out-of-range Likert responses indicate data quality issues.
Look at skewness. Values beyond \(\pm 1\) suggest the distribution departs notably from normality. This affects the choice between parametric and non-parametric tests.
Always use weights when available. Unweighted statistics describe your sample; weighted statistics estimate the population.
Consider your audience. For technical readers, include SD, skewness, and kurtosis. For general audiences, focus on means, medians, and percentages.
describe() summarizes numeric
variables (mean, SD, median, range, skewness)frequency() counts categories with
percent, valid percent, and cumulative percentcrosstab() shows the joint
distribution of two categorical variablescodebook() provides an interactive
HTML data dictionarygroup_by()vignette("hypothesis-testing")vignette("correlation-analysis")vignette("survey-weights")