--- title: "flyway(): Aggregating Several Linked Event Dates in One Table" subtitle: "From onset to admission to outcome, aligned by time and stratum" author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service" date: "`r Sys.Date()`" output: html_document: toc: true toc_depth: 3 toc_float: true theme: flatly pdf_document: toc: true toc_depth: 3 number_sections: true latex_engine: xelatex vignette: > %\VignetteIndexEntry{flyway(): Aggregating Several Linked Event Dates in One Table} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE) library(mudnester) ``` ## What `flyway()` does A `starling::murmuration()`-style linkage (e.g. case-to-hospital, or "c2h") produces one linked record per person, carrying several milestone dates for the same case: onset, hospital admission, ICU admission, a complication, death. `roost()` aggregates a *single* date column at a time — exactly what you want for a simple epi curve, but not directly what you want when the goal is a rate that compares two of those milestones, like an observed case-hospitalisation rate. `flyway()` runs `roost()` once per milestone date, then joins the results into one table, aligned by the same time unit and stratum, with one `n_` count column per event. Named after a flyway — the migratory corridor linking a bird population's successive stopover sites — because that's exactly the shape of the data: the same cohort, traced through each waypoint of a single journey. It's implemented as a thin wrapper around `roost()`, not a change to `roost()` itself. Every event's own calendar-grid zero-filling, NA-date handling, and validation is exactly `roost()`'s own — nothing about that carefully-tested logic is duplicated or re-implemented. --- ## Synthetic linked data ```{r data} set.seed(7) n <- 500 linked <- data.frame( onset_date = as.Date("2024-01-01") + sample(0:89, n, replace = TRUE), age = sample(0:90, n, replace = TRUE), stringsAsFactors = FALSE ) linked <- preening(linked, age_col = "age", scheme = "flucan_sentinel") # Simulate a care pathway: ~8% of cases are admitted, ~15% of those reach # ICU, ~5% of those die -- each downstream event's date lags the previous admitted <- sample(seq_len(n), size = round(n * 0.08)) linked$admission_date <- as.Date(NA) linked$admission_date[admitted] <- linked$onset_date[admitted] + sample(1:5, length(admitted), TRUE) icu <- sample(admitted, size = round(length(admitted) * 0.15)) linked$icu_date <- as.Date(NA) linked$icu_date[icu] <- linked$admission_date[icu] + sample(1:3, length(icu), TRUE) died <- sample(icu, size = max(1, round(length(icu) * 0.30))) linked$fatality_date <- as.Date(NA) linked$fatality_date[died] <- linked$icu_date[died] + sample(1:7, length(died), TRUE) head(linked[, c("onset_date", "age_group", "admission_date", "icu_date", "fatality_date")]) ``` --- ## One call, several event columns ```{r flyway} linked_monthly <- flyway( linked, events = c( cases = "onset_date", hospitalisations = "admission_date", icu = "icu_date", deaths = "fatality_date" ), time_unit = "month", group_cols = "age_group" ) linked_monthly ``` Every row is one `age_group` × `month` stratum, with `n_cases`, `n_hospitalisations`, `n_icu`, and `n_deaths` sitting side by side — because every one of those counts was aggregated with the *same* `time_unit` and `group_cols`, they're directly comparable. --- ## Observed rates, directly ```{r rates} linked_monthly$chr_obs <- linked_monthly$n_hospitalisations / linked_monthly$n_cases linked_monthly$cfr_obs <- linked_monthly$n_deaths / linked_monthly$n_cases linked_monthly[, c("age_group", "month", "n_cases", "n_hospitalisations", "n_deaths", "chr_obs", "cfr_obs")] ``` These are *naive* (uncorrected) rates — exactly the quantity `corncrake()`'s `method = "severity_anchor"` expects as its starting point (see `vignette("corncrake")`), since correcting for under-ascertainment means comparing this observed ratio against a reference rate from a well-ascertained source. --- ## Why a wrapper, not a change to `roost()` Two events rarely share the same calendar range — deaths lag onset by definition, so a `roost()` call on `fatality_date` alone would only start its calendar grid wherever the *first* death happens to fall. `flyway()` handles this by zero-filling *after* the join: a stratum/month with cases but no deaths yet gets `n_deaths = 0`, not `NA`, consistent with `roost()`'s own zero-filling philosophy — just extended to work correctly across several joined calendar grids rather than one. ```{r na-check} # No NAs in any count column, even though the underlying events have very # different calendar coverage colSums(is.na(linked_monthly[, c("n_cases", "n_hospitalisations", "n_icu", "n_deaths")])) ``` --- ## Feeding straight into `corncrake()` Because `flyway()` output inherits `roost_tbl`, `corncrake()`'s `time_col` auto-detection works on it exactly as it would on plain `roost()` output — just point `count_col` and `severity_count_col` at the two event columns you want to compare: ```{r corncrake-integration} corncrake( linked_monthly, count_col = "n_cases", method = "severity_anchor", group_by = "age_group", severity_count_col = "n_deaths", reference_rate = 0.01, reference_rate_lower = 0.005, reference_rate_upper = 0.02, reference_source = "Illustrative reference IFR" )[, c("age_group", "month", "n_cases", "n_deaths", "ascertainment_factor", "corrected_count")] ``` See `vignette("corncrake")` for the full detail on `method = "severity_anchor"`, including why the confidence bounds invert relative to `reference_rate`'s own bounds. --- ## Unnamed events If `events` isn't named, `flyway()` falls back to `event_1`, `event_2`, ... with a message — usable, but `n_hospitalisations` is a lot more legible three months from now than `n_event_2`: ```{r unnamed, error=TRUE} flyway(linked, events = c("onset_date", "admission_date")) ```