--- title: "Introduction to mariposa" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Introduction to mariposa} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", message = FALSE, warning = FALSE ) ``` ```{r setup} library(mariposa) library(dplyr) ``` ## What is mariposa? mariposa (*Marburg Initiative for Political and Social Analysis*) is a comprehensive R package for professional survey data analysis. It covers the entire workflow --- from importing SPSS, Stata, SAS, and Excel files through label management, recoding, and standardization to statistical analysis with survey weights and publication-ready output. Every statistical result is validated against SPSS v29, so researchers migrating from SPSS can trust their numbers. ### Key Features - **76 functions** across 15 categories - **Full data pipeline**: import → labels → transformation → analysis → export - **Survey weights** built into every function - **Tidyverse integration**: pipes (`%>%`), `group_by()`, tidyselect - **Two-level output**: compact `print()` and detailed `summary()` - **SPSS-validated**: 4,986+ tests ensure results match SPSS v29 ## The Example Dataset mariposa includes `survey_data`, a synthetic survey of 2,500 respondents with demographics, attitudes, and a sampling weight: ```{r} data(survey_data) glimpse(survey_data) ``` All examples in this guide use this dataset. ## Five-Minute Tour Here is a complete analysis workflow showing what mariposa can do: ### 1. Explore the Data ```{r} # Find variables related to "trust" find_var(survey_data, "trust") ``` ```{r} # Descriptive statistics with survey weights survey_data %>% describe(age, income, life_satisfaction, weights = sampling_weight) ``` ```{r} # Frequency table survey_data %>% frequency(education, weights = sampling_weight) ``` ### 2. Transform Variables ```{r} # Create age groups survey_data <- rec(survey_data, age, rules = "18:29=1 [Young]; 30:49=2 [Middle]; 50:99=3 [Older]", suffix = "_group", as_factor = TRUE) # Build a trust scale survey_data <- survey_data %>% mutate(m_trust = row_means(., trust_government, trust_media, trust_science, min_valid = 2)) ``` ### 3. Compare Groups ```{r} # t-test with survey weights survey_data %>% t_test(life_satisfaction, group = gender, weights = sampling_weight) ``` ```{r} # ANOVA across education levels result <- survey_data %>% oneway_anova(life_satisfaction, group = education, weights = sampling_weight) result ``` Every result has a detailed view with `summary()`: ```{r} summary(result, descriptives = FALSE) ``` ### 4. Post-Hoc Analysis ```{r} # Which education groups differ? tukey_test(result) ``` ### 5. Measure Relationships ```{r} survey_data %>% pearson_cor(age, income, life_satisfaction, weights = sampling_weight) ``` ### 6. Build Models ```{r} survey_data %>% linear_regression(life_satisfaction ~ age + income + m_trust, weights = sampling_weight) ``` ## Compact vs. Detailed Output Every analysis function in mariposa provides two output levels: - **`print()`** (default): A compact one-line summary with the key statistic - **`summary()`**: Full SPSS-style output with all details You can toggle individual sections in the detailed output: ```{r} result <- survey_data %>% t_test(life_satisfaction, group = gender, weights = sampling_weight) # Compact result # Detailed summary(result) # Detailed, skip effect sizes summary(result, effect_sizes = FALSE) ``` ## Grouped Analysis All functions support `dplyr::group_by()` for subgroup analysis: ```{r} survey_data %>% group_by(region) %>% describe(income, life_satisfaction, weights = sampling_weight) ``` ```{r} survey_data %>% group_by(region) %>% t_test(income, group = gender, weights = sampling_weight) ``` ## Quick Reference ### Data Import & Export | Function | Purpose | |----------|---------| | `read_spss()`, `read_por()` | Import SPSS files with tagged NA support | | `read_stata()` | Import Stata files | | `read_sas()`, `read_xpt()` | Import SAS files | | `read_xlsx()` | Import Excel files with label reconstruction | | `write_spss()` | Export to SPSS with label/missing roundtripping | | `write_stata()` | Export to Stata | | `write_xpt()` | Export to SAS transport format | | `write_xlsx()` | Export to Excel (data, codebook, frequencies) | ### Label Management | Function | Purpose | |----------|---------| | `var_label()` | Get/set variable labels | | `val_labels()` | Get/set value labels | | `find_var()` | Search variables by name or label | | `to_label()` | Labelled → factor | | `to_character()` | Labelled → character | | `to_numeric()` | Factor/labelled → numeric | | `to_labelled()` | Factor/character → labelled | | `set_na()` | Declare values as missing | | `unlabel()` | Strip all label metadata | | `copy_labels()` | Restore labels after dplyr operations | | `drop_labels()` | Remove unused value labels | ### Data Transformation | Function | Purpose | |----------|---------| | `rec()` | Recode with string syntax (ranges, reverse, median split) | | `to_dummy()` | One-hot encoding / dummy variables | | `std()` | Z-standardization (sd, 2sd, mad, gmd methods) | | `center()` | Mean-centering (grand-mean, group-mean) | | `row_means()` | Row-wise means with min_valid threshold | | `row_sums()` | Row-wise sums | | `row_count()` | Count specific values per row | | `pomps()` | Percent of Maximum Possible Scores (0--100) | ### Descriptive Statistics | Function | Purpose | |----------|---------| | `codebook()` | Interactive HTML data dictionary | | `describe()` | Numeric summaries (mean, sd, median, range, skewness) | | `frequency()` | Frequency tables with valid/cumulative percent | | `crosstab()` | Cross-tabulations with row/column/cell percentages | ### Hypothesis Testing | Function | Purpose | |----------|---------| | `t_test()` | Independent and one-sample t-tests | | `oneway_anova()` | One-way ANOVA | | `factorial_anova()` | Multi-factor ANOVA with Type III SS | | `ancova()` | ANCOVA with estimated marginal means | | `mann_whitney()` | Mann-Whitney U test | | `kruskal_wallis()` | Kruskal-Wallis H test | | `wilcoxon_test()` | Wilcoxon signed-rank test | | `friedman_test()` | Friedman test | | `binomial_test()` | Exact binomial test | | `chi_square()` | Chi-square test of independence | | `fisher_test()` | Fisher's exact test | | `chisq_gof()` | Chi-square goodness-of-fit | | `mcnemar_test()` | McNemar's test for paired proportions | ### Post-Hoc & Effect Sizes | Function | Purpose | |----------|---------| | `tukey_test()` | Tukey HSD pairwise comparisons | | `scheffe_test()` | Scheffe pairwise comparisons | | `levene_test()` | Test for homogeneity of variances | | `dunn_test()` | Dunn's post-hoc for Kruskal-Wallis | | `pairwise_wilcoxon()` | Pairwise Wilcoxon for Friedman | | `phi()` | Phi coefficient | | `cramers_v()` | Cramer's V | | `goodman_gamma()` | Goodman-Kruskal gamma | ### Scale Analysis | Function | Purpose | |----------|---------| | `reliability()` | Cronbach's Alpha with item statistics | | `efa()` | Exploratory Factor Analysis (PCA/ML, Varimax/Oblimin/Promax) | ### Regression | Function | Purpose | |----------|---------| | `linear_regression()` | Linear regression with SPSS-style output | | `logistic_regression()` | Logistic regression with odds ratios | ### Weighted Statistics | Function | Purpose | |----------|---------| | `w_mean()`, `w_median()`, `w_sd()`, `w_var()` | Central tendency and spread | | `w_se()`, `w_quantile()`, `w_iqr()`, `w_range()` | Precision and distribution | | `w_skew()`, `w_kurtosis()`, `w_modus()` | Shape and mode | ## Guides Explore the full documentation: - **Data Management** - `vignette("data-io")` --- Importing and exporting data - `vignette("labels-and-missing-values")` --- Working with labels and missing values - `vignette("data-transformation")` --- Recoding, standardization, and row operations - **Core Analysis** - `vignette("descriptive-statistics")` --- Summaries, frequencies, and cross-tabulations - `vignette("hypothesis-testing")` --- Comparing groups and testing hypotheses - `vignette("correlation-analysis")` --- Measuring relationships between variables - **Advanced Topics** - `vignette("scale-analysis")` --- Reliability, factor analysis, and scale construction - `vignette("regression-analysis")` --- Linear and logistic regression - `vignette("survey-weights")` --- Working with weighted data