--- title: "Calculating Health Inequality in Complex Surveys" author: "Sujon Mia - StatAid Research Lab" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Calculating Health Inequality in Complex Surveys} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} is_srvyr_ready <- requireNamespace("srvyr", quietly = TRUE) knitr::opts_chunk$set( collapse = TRUE, comment = "#>", eval = is_srvyr_ready ) ``` ## Introduction Analyzing health inequalities—such as the concentration of childhood malnutrition among the poorest wealth quintiles—is a critical component of global health research. However, calculating standardized metrics like the **Concentration Index** using data from complex, multi-stage stratified cluster surveys (such as DHS or MICS) is methodologically challenging. The `SurveyNCD` package, developed by the **StatAid Research Lab**, provides a streamlined, mathematically rigorous pipeline to automate these calculations. This vignette demonstrates how to process raw anthropometric data and calculate survey-weighted health inequality metrics in just a few lines of code. ## Setup and Package Loading First, load `SurveyNCD` along with the standard `tidyverse` and `srvyr` packages for data manipulation and survey design handling. ```{r setup-packages, message=FALSE, warning=FALSE} library(SurveyNCD) library(dplyr) library(srvyr) ``` ## Phase 1: Cleaning Anthropometric Data Raw survey datasets often store Z-scores as integers (e.g., `-254` instead of `-2.54`) and use specific numerical flags (like `9999`) to indicate missing or biologically implausible data. Failing to account for these flags will severely distort your means and indices. The `who_anthro_score()` function automatically cleans these values and categorizes them according to strict WHO child growth standards. ```{r anthro-clean} # Simulating raw survey data with a missing flag (9999) raw_survey_data <- tibble( cluster_id = c(1, 1, 2, 2), strata = c(1, 1, 2, 2), sample_weight = c(1.2, 0.8, 1.1, 0.9), wealth_score = c(-1.5, -0.2, 0.5, 1.8), raw_hw70 = c(-310, -254, 50, 9999) # Raw Height-for-Age ) # Apply the SurveyNCD cleaning function cleaned_data <- raw_survey_data %>% mutate( stunting_category = who_anthro_score(raw_hw70, indicator = "stunting"), is_stunted = case_when( stunting_category %in% c("Severe stunting", "Moderate stunting") ~ 1, stunting_category == "Normal stunting" ~ 0, TRUE ~ NA_real_ ) ) cleaned_data %>% select(raw_hw70, stunting_category, is_stunted) ``` ## Phase 2: Specifying the Complex Survey Design Before calculating inequality, we must declare the complex survey design. ```{r survey-design} svy_design <- cleaned_data %>% as_survey_design( ids = cluster_id, strata = strata, weights = sample_weight ) ``` ## Phase 3: Calculating the Concentration Index The `survey_concentration_index()` function handles the internal complexities of calculating weighted fractional ranks and covariances across the survey design. ```{r calc-ci} stunting_inequality <- svy_design %>% survey_concentration_index( outcome = is_stunted, wealth = wealth_score ) print(stunting_inequality) ``` ### Interpretation The output provides both the weighted prevalence of the outcome and the precise `Concentration_Index`. A negative index confirms a pro-poor inequality in child stunting.