What flyway() does

A starling::murmuration()-style linkage (e.g. case-to-hospital, or “c2h”) produces one linked record per person, carrying several milestone dates for the same case: onset, hospital admission, ICU admission, a complication, death. roost() aggregates a single date column at a time — exactly what you want for a simple epi curve, but not directly what you want when the goal is a rate that compares two of those milestones, like an observed case-hospitalisation rate.

flyway() runs roost() once per milestone date, then joins the results into one table, aligned by the same time unit and stratum, with one n_<name> count column per event. Named after a flyway — the migratory corridor linking a bird population’s successive stopover sites — because that’s exactly the shape of the data: the same cohort, traced through each waypoint of a single journey.

It’s implemented as a thin wrapper around roost(), not a change to roost() itself. Every event’s own calendar-grid zero-filling, NA-date handling, and validation is exactly roost()’s own — nothing about that carefully-tested logic is duplicated or re-implemented.


Synthetic linked data

set.seed(7)
n <- 500
linked <- data.frame(
  onset_date       = as.Date("2024-01-01") + sample(0:89, n, replace = TRUE),
  age               = sample(0:90, n, replace = TRUE),
  stringsAsFactors  = FALSE
)
linked <- preening(linked, age_col = "age", scheme = "flucan_sentinel")

# Simulate a care pathway: ~8% of cases are admitted, ~15% of those reach
# ICU, ~5% of those die -- each downstream event's date lags the previous
admitted <- sample(seq_len(n), size = round(n * 0.08))
linked$admission_date <- as.Date(NA)
linked$admission_date[admitted] <- linked$onset_date[admitted] + sample(1:5, length(admitted), TRUE)

icu <- sample(admitted, size = round(length(admitted) * 0.15))
linked$icu_date <- as.Date(NA)
linked$icu_date[icu] <- linked$admission_date[icu] + sample(1:3, length(icu), TRUE)

died <- sample(icu, size = max(1, round(length(icu) * 0.30)))
linked$fatality_date <- as.Date(NA)
linked$fatality_date[died] <- linked$icu_date[died] + sample(1:7, length(died), TRUE)

head(linked[, c("onset_date", "age_group", "admission_date", "icu_date", "fatality_date")])
#>   onset_date age_group admission_date icu_date fatality_date
#> 1 2024-02-11     50-64           <NA>     <NA>          <NA>
#> 2 2024-03-23     50-64           <NA>     <NA>          <NA>
#> 3 2024-01-31     50-64           <NA>     <NA>          <NA>
#> 4 2024-03-06      5-15           <NA>     <NA>          <NA>
#> 5 2024-01-15     16-49           <NA>     <NA>          <NA>
#> 6 2024-03-30       0-4           <NA>     <NA>          <NA>

One call, several event columns

linked_monthly <- flyway(
  linked,
  events = c(
    cases            = "onset_date",
    hospitalisations = "admission_date",
    icu              = "icu_date",
    deaths           = "fatality_date"
  ),
  time_unit  = "month",
  group_cols = "age_group"
)

linked_monthly
#> # A tibble: 15 × 6
#>    age_group month      n_cases n_hospitalisations n_icu n_deaths
#>    <ord>     <date>       <int>              <int> <int>    <int>
#>  1 0-4       2024-01-01      11                  1     0        0
#>  2 0-4       2024-02-01      13                  2     0        0
#>  3 0-4       2024-03-01       4                  0     0        0
#>  4 5-15      2024-01-01      21                  2     0        0
#>  5 5-15      2024-02-01      16                  0     0        0
#>  6 5-15      2024-03-01      25                  3     1        0
#>  7 16-49     2024-01-01      65                  5     0        0
#>  8 16-49     2024-02-01      59                  3     0        0
#>  9 16-49     2024-03-01      67                  3     0        0
#> 10 50-64     2024-01-01      32                  3     0        0
#> 11 50-64     2024-02-01      22                  6     1        0
#> 12 50-64     2024-03-01      26                  1     0        0
#> 13 65+       2024-01-01      49                  3     0        0
#> 14 65+       2024-02-01      45                  5     3        1
#> 15 65+       2024-03-01      45                  3     1        1
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-03-31 
#>  group_cols : age_group 
#>  hemisphere : southern 
#>  n_rows_in  : 500  (1452 NA dates dropped)
#> 
#> -- flyway events ----------------------------------
#>  cases <- onset_date  (n_cases)
#>  hospitalisations <- admission_date  (n_hospitalisations)
#>  icu <- icu_date  (n_icu)
#>  deaths <- fatality_date  (n_deaths)

Every row is one age_group × month stratum, with n_cases, n_hospitalisations, n_icu, and n_deaths sitting side by side — because every one of those counts was aggregated with the same time_unit and group_cols, they’re directly comparable.


Observed rates, directly

linked_monthly$chr_obs <- linked_monthly$n_hospitalisations / linked_monthly$n_cases
linked_monthly$cfr_obs <- linked_monthly$n_deaths / linked_monthly$n_cases

linked_monthly[, c("age_group", "month", "n_cases", "n_hospitalisations", "n_deaths",
                    "chr_obs", "cfr_obs")]
#> # A tibble: 15 × 7
#>    age_group month      n_cases n_hospitalisations n_deaths chr_obs cfr_obs
#>    <ord>     <date>       <int>              <int>    <int>   <dbl>   <dbl>
#>  1 0-4       2024-01-01      11                  1        0  0.0909  0     
#>  2 0-4       2024-02-01      13                  2        0  0.154   0     
#>  3 0-4       2024-03-01       4                  0        0  0       0     
#>  4 5-15      2024-01-01      21                  2        0  0.0952  0     
#>  5 5-15      2024-02-01      16                  0        0  0       0     
#>  6 5-15      2024-03-01      25                  3        0  0.12    0     
#>  7 16-49     2024-01-01      65                  5        0  0.0769  0     
#>  8 16-49     2024-02-01      59                  3        0  0.0508  0     
#>  9 16-49     2024-03-01      67                  3        0  0.0448  0     
#> 10 50-64     2024-01-01      32                  3        0  0.0938  0     
#> 11 50-64     2024-02-01      22                  6        0  0.273   0     
#> 12 50-64     2024-03-01      26                  1        0  0.0385  0     
#> 13 65+       2024-01-01      49                  3        0  0.0612  0     
#> 14 65+       2024-02-01      45                  5        1  0.111   0.0222
#> 15 65+       2024-03-01      45                  3        1  0.0667  0.0222
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-03-31 
#>  group_cols : age_group 
#>  hemisphere : southern 
#>  n_rows_in  : 500  (1452 NA dates dropped)
#> 
#> -- flyway events ----------------------------------
#>  cases <- onset_date  (n_cases)
#>  hospitalisations <- admission_date  (n_hospitalisations)
#>  icu <- icu_date  (n_icu)
#>  deaths <- fatality_date  (n_deaths)

These are naive (uncorrected) rates — exactly the quantity corncrake()’s method = "severity_anchor" expects as its starting point (see vignette("corncrake")), since correcting for under-ascertainment means comparing this observed ratio against a reference rate from a well-ascertained source.


Why a wrapper, not a change to roost()

Two events rarely share the same calendar range — deaths lag onset by definition, so a roost() call on fatality_date alone would only start its calendar grid wherever the first death happens to fall. flyway() handles this by zero-filling after the join: a stratum/month with cases but no deaths yet gets n_deaths = 0, not NA, consistent with roost()’s own zero-filling philosophy — just extended to work correctly across several joined calendar grids rather than one.

# No NAs in any count column, even though the underlying events have very
# different calendar coverage
colSums(is.na(linked_monthly[, c("n_cases", "n_hospitalisations", "n_icu", "n_deaths")]))
#>            n_cases n_hospitalisations              n_icu           n_deaths 
#>                  0                  0                  0                  0

Feeding straight into corncrake()

Because flyway() output inherits roost_tbl, corncrake()’s time_col auto-detection works on it exactly as it would on plain roost() output — just point count_col and severity_count_col at the two event columns you want to compare:

corncrake(
  linked_monthly,
  count_col           = "n_cases",
  method               = "severity_anchor",
  group_by             = "age_group",
  severity_count_col   = "n_deaths",
  reference_rate       = 0.01,
  reference_rate_lower = 0.005,
  reference_rate_upper = 0.02,
  reference_source     = "Illustrative reference IFR"
)[, c("age_group", "month", "n_cases", "n_deaths", "ascertainment_factor", "corrected_count")]
#> # A tibble: 15 × 6
#>    age_group month      n_cases n_deaths ascertainment_factor corrected_count
#>    <ord>     <date>       <int>    <int>                <dbl>           <dbl>
#>  1 0-4       2024-01-01      11        0                 0                  0
#>  2 0-4       2024-02-01      13        0                 0                  0
#>  3 0-4       2024-03-01       4        0                 0                  0
#>  4 5-15      2024-01-01      21        0                 0                  0
#>  5 5-15      2024-02-01      16        0                 0                  0
#>  6 5-15      2024-03-01      25        0                 0                  0
#>  7 16-49     2024-01-01      65        0                 0                  0
#>  8 16-49     2024-02-01      59        0                 0                  0
#>  9 16-49     2024-03-01      67        0                 0                  0
#> 10 50-64     2024-01-01      32        0                 0                  0
#> 11 50-64     2024-02-01      22        0                 0                  0
#> 12 50-64     2024-03-01      26        0                 0                  0
#> 13 65+       2024-01-01      49        0                 0                  0
#> 14 65+       2024-02-01      45        1                 2.22             100
#> 15 65+       2024-03-01      45        1                 2.22             100
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-03-31 
#>  group_cols : age_group 
#>  hemisphere : southern 
#>  n_rows_in  : 500  (1452 NA dates dropped)
#> 
#> -- flyway events ----------------------------------
#>  cases <- onset_date  (n_cases)
#>  hospitalisations <- admission_date  (n_hospitalisations)
#>  icu <- icu_date  (n_icu)
#>  deaths <- fatality_date  (n_deaths)

See vignette("corncrake") for the full detail on method = "severity_anchor", including why the confidence bounds invert relative to reference_rate’s own bounds.


Unnamed events

If events isn’t named, flyway() falls back to event_1, event_2, … with a message — usable, but n_hospitalisations is a lot more legible three months from now than n_event_2:

flyway(linked, events = c("onset_date", "admission_date"))
#> # A tibble: 3 × 3
#>   month      n_event_1 n_event_2
#>   <date>         <int>     <int>
#> 1 2024-01-01       178        14
#> 2 2024-02-01       155        16
#> 3 2024-03-01       167        10
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-03-30 
#>  hemisphere : southern 
#>  n_rows_in  : 500  (460 NA dates dropped)
#> 
#> -- flyway events ----------------------------------
#>  event_1 <- onset_date  (n_event_1)
#>  event_2 <- admission_date  (n_event_2)