--- title: "preening(): Age Categorisation Against ~50 Named Schemes" subtitle: "From ABS 5-year bands to ATAGI program-specific groupings" author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service" date: "`r Sys.Date()`" output: html_document: toc: true toc_depth: 3 toc_float: true theme: flatly pdf_document: toc: true toc_depth: 3 number_sections: true latex_engine: xelatex vignette: > %\VignetteIndexEntry{preening(): Age Categorisation Against ~50 Named Schemes} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE) library(mudnester) ``` ## The problem `preening()` solves Age categorisation is one of the most routine and most error-prone steps in surveillance analysis. The same dataset might need ABS 5-year bands for a national comparison, ATAGI program bands for a vaccine effectiveness study, and FluCAN bands for a sentinel surveillance report — all in the same week. `preening()` provides a single function backed by a catalogue of ~50 named, citable schemes so that band choice is explicit, reproducible, and traceable to a published source. The name comes from the way a bird re-sorts its feathers into whichever functional arrangement suits the moment, without changing anything about the bird itself. The same raw age values are re-sorted into whichever standard grouping the analysis calls for. --- ## Three ways to specify a scheme ### 1. Exact name The clearest approach — name the scheme directly. ```{r exact} set.seed(1) df <- data.frame(age = c(0.2, 3, 14, 25, 50, 67, 80, 92)) preening(df, age_col = "age", scheme = "atagi_covid19_2025")$age_group ``` ### 2. Filter to a single match Supply `family`, `focus`, and/or `max_bands` — if exactly one scheme matches, it is applied automatically and a message names it so the choice is never silent. ```{r filter-single} # vaccination + paediatric + max 3 bands → exactly one match preening(df, age_col = "age", family = "vaccination", focus = "paediatric", max_bands = 3)$age_group ``` ### 3. Let `list_age_schemes()` guide you Browse the catalogue before committing to a scheme. ```{r list-schemes} list_age_schemes(family = "surveillance") ``` When a filter matches multiple schemes, `preening()` stops and lists them so you can pick one explicitly. This is intentional — `preening()` never guesses among ties. ```{r multi-match-error, error=TRUE} # Multiple paediatric schemes exist — preening() asks you to choose preening(df, age_col = "age", focus = "paediatric") ``` --- ## The scheme families Schemes are organised into six families. Use `family =` to restrict your search. | Family | `family =` value | Count | Examples | |---|---|---|---| | National statistical standards | `"national_stats"` | 10 | `abs_5yr`, `abs_broad_lifecourse` | | International statistical standards | `"international_stats"` | 7 | `who_life_course`, `eurostat_5yr` | | Vaccination/immunisation guidance | `"vaccination"` | 10 | `atagi_covid19_2025`, `flucan_sentinel` | | Surveillance-system conventions | `"surveillance"` | 8 | `nndss_standard`, `racf_aged_care` | | Clinical/developmental staging | `"clinical_developmental"` | 8 | `geriatric_fine`, `paediatric_developmental` | | Disease/research-specific | `"disease_specific"` | 7 | `rsv_research`, `covid19_severity_strata` | --- ## Focus tags — cross-cutting filters Every scheme carries one or more focus tags that cut across families. ```{r focus-tags} list_age_schemes(focus = "aged_care") ``` ```{r focus-broad} # Quick summary schemes for small datasets or executive reports list_age_schemes(focus = "broad", max_bands = 5) ``` --- ## Neonatal schemes: `age_unit = "days"` Schemes in the neonatal family use days rather than years. Pass `age_unit = "days"` and ensure the age column is in days. ```{r neonatal} neonates <- data.frame(age_days = c(0, 0.5, 2, 5, 15, 30)) preening(neonates, age_col = "age_days", scheme = "neonatal_early", age_unit = "days")$age_group ``` --- ## Custom schemes When no standard scheme fits, supply your own breaks and labels. ```{r custom} preening( df, age_col = "age", scheme = "custom", age_breaks = c(0, 18, 40, 65, Inf), age_labels = c("0-17", "18-39", "40-64", "65+") )$age_group ``` --- ## Catch-all bands and the full-lifespan rule Every scheme in the `mudnester` library spans 0 to `Inf`, so `preening()` never returns `NA` purely because a record fell outside a scheme's "intended" range. Bands marked with `⁺` in the documentation (e.g. the `0-<60` floor band in `rsv_older_adult`) are catch-alls — a meaningful count in one of these bands is a signal to review the scheme choice, not a finding to report. ```{r catchall} # rsv_older_adult is scoped to 60+. A child record still gets a band. data.frame(age = c(3, 65, 80)) |> preening(age_col = "age", scheme = "rsv_older_adult") ``` --- ## After preening: what comes next `preening()` is typically called before `roost()` to enable age-stratified counts: ```{r roost-integration, eval=FALSE} df |> preening(age_col = "age", scheme = "flucan_sentinel") |> roost(date_col = "onset_date", time_unit = "month", group_cols = "age_group") ``` See `vignette("roost")` for aggregation options, and `vignette("age-schemes")` for the full scheme catalogue with source citations.