flyway() doesA starling::murmuration()-style linkage
(e.g. case-to-hospital, or “c2h”) produces one linked record per person,
carrying several milestone dates for the same case: onset, hospital
admission, ICU admission, a complication, death. roost()
aggregates a single date column at a time — exactly what you
want for a simple epi curve, but not directly what you want when the
goal is a rate that compares two of those milestones, like an observed
case-hospitalisation rate.
flyway() runs roost() once per milestone
date, then joins the results into one table, aligned by the same time
unit and stratum, with one n_<name> count column per
event. Named after a flyway — the migratory corridor linking a bird
population’s successive stopover sites — because that’s exactly the
shape of the data: the same cohort, traced through each waypoint of a
single journey.
It’s implemented as a thin wrapper around roost(), not a
change to roost() itself. Every event’s own calendar-grid
zero-filling, NA-date handling, and validation is exactly
roost()’s own — nothing about that carefully-tested logic
is duplicated or re-implemented.
set.seed(7)
n <- 500
linked <- data.frame(
onset_date = as.Date("2024-01-01") + sample(0:89, n, replace = TRUE),
age = sample(0:90, n, replace = TRUE),
stringsAsFactors = FALSE
)
linked <- preening(linked, age_col = "age", scheme = "flucan_sentinel")
# Simulate a care pathway: ~8% of cases are admitted, ~15% of those reach
# ICU, ~5% of those die -- each downstream event's date lags the previous
admitted <- sample(seq_len(n), size = round(n * 0.08))
linked$admission_date <- as.Date(NA)
linked$admission_date[admitted] <- linked$onset_date[admitted] + sample(1:5, length(admitted), TRUE)
icu <- sample(admitted, size = round(length(admitted) * 0.15))
linked$icu_date <- as.Date(NA)
linked$icu_date[icu] <- linked$admission_date[icu] + sample(1:3, length(icu), TRUE)
died <- sample(icu, size = max(1, round(length(icu) * 0.30)))
linked$fatality_date <- as.Date(NA)
linked$fatality_date[died] <- linked$icu_date[died] + sample(1:7, length(died), TRUE)
head(linked[, c("onset_date", "age_group", "admission_date", "icu_date", "fatality_date")])
#> onset_date age_group admission_date icu_date fatality_date
#> 1 2024-02-11 50-64 <NA> <NA> <NA>
#> 2 2024-03-23 50-64 <NA> <NA> <NA>
#> 3 2024-01-31 50-64 <NA> <NA> <NA>
#> 4 2024-03-06 5-15 <NA> <NA> <NA>
#> 5 2024-01-15 16-49 <NA> <NA> <NA>
#> 6 2024-03-30 0-4 <NA> <NA> <NA>
linked_monthly <- flyway(
linked,
events = c(
cases = "onset_date",
hospitalisations = "admission_date",
icu = "icu_date",
deaths = "fatality_date"
),
time_unit = "month",
group_cols = "age_group"
)
linked_monthly
#> # A tibble: 15 × 6
#> age_group month n_cases n_hospitalisations n_icu n_deaths
#> <ord> <date> <int> <int> <int> <int>
#> 1 0-4 2024-01-01 11 1 0 0
#> 2 0-4 2024-02-01 13 2 0 0
#> 3 0-4 2024-03-01 4 0 0 0
#> 4 5-15 2024-01-01 21 2 0 0
#> 5 5-15 2024-02-01 16 0 0 0
#> 6 5-15 2024-03-01 25 3 1 0
#> 7 16-49 2024-01-01 65 5 0 0
#> 8 16-49 2024-02-01 59 3 0 0
#> 9 16-49 2024-03-01 67 3 0 0
#> 10 50-64 2024-01-01 32 3 0 0
#> 11 50-64 2024-02-01 22 6 1 0
#> 12 50-64 2024-03-01 26 1 0 0
#> 13 65+ 2024-01-01 49 3 0 0
#> 14 65+ 2024-02-01 45 5 3 1
#> 15 65+ 2024-03-01 45 3 1 1
#>
#> -- roost_meta --------------------------------------
#> time_unit : month
#> date_range : 2024-01-01 to 2024-03-31
#> group_cols : age_group
#> hemisphere : southern
#> n_rows_in : 500 (1452 NA dates dropped)
#>
#> -- flyway events ----------------------------------
#> cases <- onset_date (n_cases)
#> hospitalisations <- admission_date (n_hospitalisations)
#> icu <- icu_date (n_icu)
#> deaths <- fatality_date (n_deaths)
Every row is one age_group × month stratum,
with n_cases, n_hospitalisations,
n_icu, and n_deaths sitting side by side —
because every one of those counts was aggregated with the same
time_unit and group_cols, they’re directly
comparable.
linked_monthly$chr_obs <- linked_monthly$n_hospitalisations / linked_monthly$n_cases
linked_monthly$cfr_obs <- linked_monthly$n_deaths / linked_monthly$n_cases
linked_monthly[, c("age_group", "month", "n_cases", "n_hospitalisations", "n_deaths",
"chr_obs", "cfr_obs")]
#> # A tibble: 15 × 7
#> age_group month n_cases n_hospitalisations n_deaths chr_obs cfr_obs
#> <ord> <date> <int> <int> <int> <dbl> <dbl>
#> 1 0-4 2024-01-01 11 1 0 0.0909 0
#> 2 0-4 2024-02-01 13 2 0 0.154 0
#> 3 0-4 2024-03-01 4 0 0 0 0
#> 4 5-15 2024-01-01 21 2 0 0.0952 0
#> 5 5-15 2024-02-01 16 0 0 0 0
#> 6 5-15 2024-03-01 25 3 0 0.12 0
#> 7 16-49 2024-01-01 65 5 0 0.0769 0
#> 8 16-49 2024-02-01 59 3 0 0.0508 0
#> 9 16-49 2024-03-01 67 3 0 0.0448 0
#> 10 50-64 2024-01-01 32 3 0 0.0938 0
#> 11 50-64 2024-02-01 22 6 0 0.273 0
#> 12 50-64 2024-03-01 26 1 0 0.0385 0
#> 13 65+ 2024-01-01 49 3 0 0.0612 0
#> 14 65+ 2024-02-01 45 5 1 0.111 0.0222
#> 15 65+ 2024-03-01 45 3 1 0.0667 0.0222
#>
#> -- roost_meta --------------------------------------
#> time_unit : month
#> date_range : 2024-01-01 to 2024-03-31
#> group_cols : age_group
#> hemisphere : southern
#> n_rows_in : 500 (1452 NA dates dropped)
#>
#> -- flyway events ----------------------------------
#> cases <- onset_date (n_cases)
#> hospitalisations <- admission_date (n_hospitalisations)
#> icu <- icu_date (n_icu)
#> deaths <- fatality_date (n_deaths)
These are naive (uncorrected) rates — exactly the quantity
corncrake()’s method = "severity_anchor"
expects as its starting point (see vignette("corncrake")),
since correcting for under-ascertainment means comparing this observed
ratio against a reference rate from a well-ascertained source.
roost()Two events rarely share the same calendar range — deaths lag onset by
definition, so a roost() call on fatality_date
alone would only start its calendar grid wherever the first
death happens to fall. flyway() handles this by
zero-filling after the join: a stratum/month with cases but no
deaths yet gets n_deaths = 0, not NA,
consistent with roost()’s own zero-filling philosophy —
just extended to work correctly across several joined calendar grids
rather than one.
# No NAs in any count column, even though the underlying events have very
# different calendar coverage
colSums(is.na(linked_monthly[, c("n_cases", "n_hospitalisations", "n_icu", "n_deaths")]))
#> n_cases n_hospitalisations n_icu n_deaths
#> 0 0 0 0
corncrake()Because flyway() output inherits roost_tbl,
corncrake()’s time_col auto-detection works on
it exactly as it would on plain roost() output — just point
count_col and severity_count_col at the two
event columns you want to compare:
corncrake(
linked_monthly,
count_col = "n_cases",
method = "severity_anchor",
group_by = "age_group",
severity_count_col = "n_deaths",
reference_rate = 0.01,
reference_rate_lower = 0.005,
reference_rate_upper = 0.02,
reference_source = "Illustrative reference IFR"
)[, c("age_group", "month", "n_cases", "n_deaths", "ascertainment_factor", "corrected_count")]
#> # A tibble: 15 × 6
#> age_group month n_cases n_deaths ascertainment_factor corrected_count
#> <ord> <date> <int> <int> <dbl> <dbl>
#> 1 0-4 2024-01-01 11 0 0 0
#> 2 0-4 2024-02-01 13 0 0 0
#> 3 0-4 2024-03-01 4 0 0 0
#> 4 5-15 2024-01-01 21 0 0 0
#> 5 5-15 2024-02-01 16 0 0 0
#> 6 5-15 2024-03-01 25 0 0 0
#> 7 16-49 2024-01-01 65 0 0 0
#> 8 16-49 2024-02-01 59 0 0 0
#> 9 16-49 2024-03-01 67 0 0 0
#> 10 50-64 2024-01-01 32 0 0 0
#> 11 50-64 2024-02-01 22 0 0 0
#> 12 50-64 2024-03-01 26 0 0 0
#> 13 65+ 2024-01-01 49 0 0 0
#> 14 65+ 2024-02-01 45 1 2.22 100
#> 15 65+ 2024-03-01 45 1 2.22 100
#>
#> -- roost_meta --------------------------------------
#> time_unit : month
#> date_range : 2024-01-01 to 2024-03-31
#> group_cols : age_group
#> hemisphere : southern
#> n_rows_in : 500 (1452 NA dates dropped)
#>
#> -- flyway events ----------------------------------
#> cases <- onset_date (n_cases)
#> hospitalisations <- admission_date (n_hospitalisations)
#> icu <- icu_date (n_icu)
#> deaths <- fatality_date (n_deaths)
See vignette("corncrake") for the full detail on
method = "severity_anchor", including why the confidence
bounds invert relative to reference_rate’s own bounds.
If events isn’t named, flyway() falls back
to event_1, event_2, … with a message —
usable, but n_hospitalisations is a lot more legible three
months from now than n_event_2:
flyway(linked, events = c("onset_date", "admission_date"))
#> # A tibble: 3 × 3
#> month n_event_1 n_event_2
#> <date> <int> <int>
#> 1 2024-01-01 178 14
#> 2 2024-02-01 155 16
#> 3 2024-03-01 167 10
#>
#> -- roost_meta --------------------------------------
#> time_unit : month
#> date_range : 2024-01-01 to 2024-03-30
#> hemisphere : southern
#> n_rows_in : 500 (460 NA dates dropped)
#>
#> -- flyway events ----------------------------------
#> event_1 <- onset_date (n_event_1)
#> event_2 <- admission_date (n_event_2)