--- title: "clean_the_nest(): Standardising Surveillance Data" subtitle: "Preparing cases, hospitalisations, and vaccination records for linkage" author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service" date: "`r Sys.Date()`" output: html_document: toc: true toc_depth: 3 toc_float: true theme: flatly pdf_document: toc: true toc_depth: 3 number_sections: true latex_engine: xelatex vignette: > %\VignetteIndexEntry{clean_the_nest(): Standardising Surveillance Data} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE) library(mudnester) ``` ## What `clean_the_nest()` does `clean_the_nest()` is the foundational layer of the `mudnester` pipeline. It: - Renames your columns to the mudnester internal schema (standardised names like `lettername1`, `dob`, `onset_date`) so every downstream function knows what to expect - Validates date formats and raises a clear error if a date column is not `Date` class - Strips and lowercases name fields for linkage consistency - Derives blocking variables (`block1`, `block2`, `block3`) used by `starling::murmuration()` - Standardises Medicare numbers into 9-, 10-, and 11-digit variants - Derives age, length of stay, ICU outcome, and death outcome automatically when the required columns are present - Pivots long vaccination data to wide (one row per person) via `lie_nest_flat = TRUE` Nothing downstream is trustworthy until this layer is structurally sound. Think of it as laying the mud — subsequent functions can only build on what is solid here. --- ## The three data types `data_type` is required. The three values map to different source systems in a Queensland public health context. | `data_type` | Typical source | Key mandatory dates | |---|---|---| | `"cases"` | NNDSS, NoCS, EDIS linelists | `onset_date` | | `"hospital"` | iPM, HBCIS, admitted patient collections | `admission_date` | | `"vaccination"` | Australian Immunisation Register (AIR), local registers | `vax_date` | Age and age categories are derived for `"cases"` and `"hospital"` when `dob` and `onset_date` / `admission_date` are supplied. They are not derived for `"vaccination"` (the AIR does not reliably carry age at time of vaccination). --- ## Worked examples ### Cases (notifiable disease linelist) ```{r cases-example} set.seed(1) cases_raw <- data.frame( identity = paste0("PT", 1:20), first_name = sample(c("James","Sarah","Michael"), 20, TRUE), surname = sample(c("Smith","Jones","Williams"), 20, TRUE), date_of_birth = as.Date("1970-01-01") + sample(-5000:5000, 20), date_of_onset = as.Date("2024-03-01") + sample(0:120, 20), disease_name = sample(c("COVID-19","Influenza A","RSV"), 20, TRUE), gender = sample(c("M","F"), 20, TRUE), postcode = sample(c("4556","4557","4560"), 20, TRUE), medicare_no = paste0(sample(2000:9999, 20), sample(10000:99999, 20)), indigenous_status = sample(c("Non-Indigenous","Aboriginal","Unknown"), 20, TRUE, prob = c(0.85, 0.10, 0.05)), stringsAsFactors = FALSE ) df_cases <- clean_the_nest( cases_raw, data_type = "cases", drop_eggs = TRUE, id_var = "identity", diagnosis = "disease_name", lettername1 = "first_name", lettername2 = "surname", dob = "date_of_birth", medicare = "medicare_no", gender = "gender", postcode = "postcode", fn = "indigenous_status", onset_date = "date_of_onset" ) head(df_cases[, c("lettername1","lettername2","dob","age","diagnosis","block1")], 4) ``` Notice that `lettername1` is lowercased, punctuation-stripped, and truncated to the first name token only. `block1` is `gender + postcode + birth_year` — used as the primary blocking variable in `starling::murmuration()`. --- ### Hospital admissions When `admission_date` and `discharge_date` are both supplied, length of stay (`los`, in days) is derived automatically. `admission_outcome` is a factor indicating whether the person was actually admitted. ```{r hospital-example} set.seed(2) hosp_raw <- data.frame( patient_id = paste0("UR", 1:15), firstname = sample(c("James","Sarah","Michael"), 15, TRUE), last_name = sample(c("Smith","Jones"), 15, TRUE), birth_date = as.Date("1945-01-01") + sample(0:10000, 15), date_of_admission = as.Date("2024-06-01") + sample(0:180, 15), date_of_discharge = as.Date("2024-06-15") + sample(0:180, 15), medicare_number = paste0(sample(2000:9999, 15), sample(10000:99999, 15)), sex = sample(c("M","F"), 15, TRUE), zip_codes = sample(c("4556","4557"), 15, TRUE), icd_codes = sample(c("J06.9","U07.1","J44.1"), 15, TRUE), stringsAsFactors = FALSE ) df_hosp <- clean_the_nest( hosp_raw, data_type = "hospital", drop_eggs = TRUE, id_var = "patient_id", lettername1 = "firstname", lettername2 = "last_name", dob = "birth_date", medicare = "medicare_number", gender = "sex", postcode = "zip_codes", icd_code = "icd_codes", admission_date = "date_of_admission", discharge_date = "date_of_discharge" ) df_hosp[, c("los","admission_outcome","icd_code")] |> head(4) ``` --- ### Vaccination data — long to wide The Australian Immunisation Register exports one row per vaccination event. `lie_nest_flat = TRUE` pivots this to one row per person, with `vax_date_1`, `vax_type_1`, `vax_date_2`, `vax_type_2`, etc. ```{r vax-example} set.seed(3) vax_raw <- data.frame( patient_id = rep(paste0("VAX", 1:10), each = 2), firstname = rep(c("Alice","Bob","Carol","Dan","Eve", "Frank","Grace","Henry","Iris","Jack"), each = 2), last_name = rep(c("Smith","Jones","Williams","Taylor","Brown", "White","Black","Green","Blue","Red"), each = 2), birth_date = rep(as.Date("1980-01-01") + sample(-2000:2000, 10), each = 2), gender = rep(sample(c("M","F"), 10, TRUE), each = 2), postcode = rep(sample(c("4556","4557"), 10, TRUE), each = 2), medicare_number = rep(paste0(sample(2000:9999, 10), sample(10000:99999, 10)), each = 2), vaccine_delivered = rep(c("COVID-19 XBB.1.5","COVID-19 XBB.1.5"), 10), service_date = c(rbind( as.Date("2024-01-15") + sample(0:30, 10), as.Date("2024-05-01") + sample(0:30, 10) )), stringsAsFactors = FALSE ) df_vax <- clean_the_nest( vax_raw, data_type = "vaccination", lie_nest_flat = TRUE, id_var = "patient_id", lettername1 = "firstname", lettername2 = "last_name", dob = "birth_date", medicare = "medicare_number", gender = "gender", postcode = "postcode", vax_type = "vaccine_delivered", vax_date = "service_date" ) df_vax[, c("id_var","vax_date_1","vax_type_1","vax_date_2","vax_type_2")] |> head(4) ``` --- ### Birth-cohort studies In birth-cohort vaccine effectiveness studies (e.g. nirsevimab or Abrysvo effectiveness in infants), the date of birth *is* the cohort entry date. Pass the same column name to both `dob` and `cohort_entry_date` — no duplication needed. ```{r cohort-example, eval=FALSE} df_cohort <- clean_the_nest( birth_cohort_data, data_type = "cases", id_var = "baby_id", lettername1 = "first_name", lettername2 = "last_name", dob = "babys_date_of_birth", cohort_entry_date = "babys_date_of_birth", # same column — aliased internally cohort_exit_date = "end_of_followup_date", gender = "sex" ) ``` --- ## Australian Medicare numbers `clean_the_nest()` understands the full structure of Australian Medicare numbers and produces semantically named output columns for each component, rather than opaque digit-count names. ### Structure A complete Medicare number has up to 11 characters: | Component | Digits | Column | Description | |---|---|---|---| | Account identifier | 1–8 | `medicare08` | Unique household/account number. First digit is always 2–6. | | Checksum | 9 | `medicare09` | Mathematically derived from digits 1–8; used to validate the number. `clean_the_nest()` validates this automatically. | | Card issue number | 10 | `medicare10` | Increments each time a new card is issued (lost, expired, family member added). Never 0. | | IRN | 11 | `medicare_irn` | Individual Reference Number — the digit printed left of a person's name. Primary cardholder = 1, partner = 2, children = 3, 4, etc. Up to 9 people share a card. | Two additional columns are always produced: `medicare_clean` (spaces removed, as supplied) and `medicare_valid` (logical — `TRUE` if the checksum is mathematically correct). ### Which column to use for linkage? | Linkage purpose | Use | |---|---| | Person-level linkage (cases ↔ AIR) | `medicare09` (stable across card reissues and family members) | | Card-specific linkage (e.g. AIR dose records) | `medicare10` | | Household-level blocking | `medicare08` | | Distinguishing individuals on the same card | `medicare_irn` | ```{r medicare} mc_data <- data.frame( patient_id = c("PT001", "PT002", "PT003"), # PT001: valid 10-digit number (no IRN appended) # PT002: valid 11-digit number (IRN = 1) # PT003: invalid checksum (digit 9 is 7, should be 3) mcare = c("2428778132", "24287781321", "2428778172"), stringsAsFactors = FALSE ) # suppressWarnings() here because PT003 has an invalid checksum — # that's exactly what we want to demonstrate. df_mc <- suppressWarnings(suppressMessages( clean_the_nest(mc_data, data_type = "cases", id_var = "patient_id", medicare = "mcare") )) df_mc[, c("id_var", "medicare08", "medicare09", "medicare10", "medicare_irn", "medicare_valid")] ``` Notice that PT003 has `medicare_valid = FALSE` — `clean_the_nest()` warns you about invalid numbers before linkage, because an invalid Medicare number will never match in `starling::murmuration()`. In real data, investigate these rows: a common cause is the IRN being stored as part of the 10-digit number rather than separately. --- ## The `drop_eggs` argument `drop_eggs = TRUE` retains only the columns needed for linkage and downstream analysis. This produces a lean, manageable dataset for early-stage work. Use `keep_vars` to retain additional columns you need alongside the defaults. ```{r drop-eggs} # Without drop_eggs — all original columns plus derived ones df_full <- clean_the_nest( cases_raw, data_type = "cases", id_var = "identity", lettername1 = "first_name", lettername2 = "surname", dob = "date_of_birth", onset_date = "date_of_onset" ) ncol(df_full) # With drop_eggs — only the linkage and analysis essentials df_lean <- clean_the_nest( cases_raw, data_type = "cases", drop_eggs = TRUE, id_var = "identity", lettername1 = "first_name", lettername2 = "surname", dob = "date_of_birth", onset_date = "date_of_onset", keep_vars = "disease_name" # retain this one extra ) ncol(df_lean) names(df_lean) ``` --- ## What comes next Once your data is cleaned: - **`preening()`** — assign age bands using one of ~50 named schemes (see `vignette("preening")`) - **`plumage()`** — detect chronic comorbidities from ICD-10-AM coding (see `vignette("plumage")`). For hospital datasets, this is typically the next step after `clean_the_nest()` — the standardised `icd_code` column produced here is the direct input to `plumage(icd_column = "icd_code")`. - **`roost()`** — aggregate to a time unit for an epi curve (see `vignette("roost")`) - **`starling::murmuration()`** — probabilistic record linkage across the cleaned datasets - **`molting()`** — de-identify before sharing (see `vignette("molting")`)