--- title: "Cookbook: Count Outcome, End to End" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Cookbook: Count Outcome, End to End} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` One complete, runnable script for a count outcome — number of events per subject (visits, relapses, defects). Same four steps as every cookbook (design → assign → record → infer; see `vignette("cookbook-continuous")`). The natural estimand is a log rate ratio; the inference classes are the `InferenceCount*` family, covering Poisson, negative binomial, hurdle and zero-inflated variants for over-dispersed or zero-heavy counts. ## Setup EDI is not on CRAN yet, so `install.packages("EDI")` fails — install from R-universe (fallback: GitHub, `subdir = "R/EDI"`). Not evaluated here. ```{r install, eval = FALSE, purl = FALSE} install.packages("EDI", repos = c("https://kapelner.r-universe.dev", "https://cloud.r-project.org")) # or: remotes::install_github("kapelner/EDI", subdir = "R/EDI") ``` ```{r setup} library(EDI) set.seed(20260916) n = 80 X = data.frame( baseline_rate = round(rgamma(n, 4, 1), 1), urban = rbinom(n, 1, 0.5) ) true_log_rr = -0.4 # treatment reduces the event rate by ~33% ``` ## Fixed design, Poisson regression ```{r fixed} des = DesignFixedBernoulli$new(n = n, response_type = "count", verbose = FALSE) des$add_all_subjects_to_experiment(X) des$assign_w_to_all_subjects() w = des$get_w() mu = exp(0.5 + true_log_rr * w + 0.15 * X$baseline_rate + 0.3 * X$urban) y = rpois(n, mu) des$add_all_subject_responses(y) inf = InferenceCountPoisson$new(des, verbose = FALSE) inf$num_cores = 1L inf$compute_estimate() # log rate ratio for treatment inf$compute_asymp_confidence_interval(alpha = 0.05) inf$compute_asymp_two_sided_pval() ``` ```{r fixed-resampling} inf$set_seed(1) inf$compute_rand_two_sided_pval(r = 200, show_progress = FALSE) inf$set_seed(1) inf$compute_bootstrap_confidence_interval(alpha = 0.05, B = 200, show_progress = FALSE) ``` ## Over-dispersion: negative binomial on the same design Real counts are usually over-dispersed relative to Poisson. Swap the class; the design object is unchanged. ```{r negbin} inf_nb = InferenceCountNegBin$new(des, verbose = FALSE) inf_nb$num_cores = 1L inf_nb$compute_estimate() inf_nb$compute_asymp_confidence_interval(alpha = 0.05) ``` ## Everything at once ```{r suite} suite = InferenceSuite$new(des) res = suite$run_all_inference(screen = TRUE, plots = FALSE, num_cores = 1L, methods = c("wald", "score", "lik_ratio"), max_secs_per_class = 15) ``` ## Sequential design The one-by-one API is identical across response types — only the generated response changes. Randomization inference is the tool for sequential designs whose assignments depend on earlier subjects. ```{r seq} des_seq = DesignSeqOneByOneKK14$new(n = n, response_type = "count", verbose = FALSE) for (i in seq_len(n)) { w_i = des_seq$add_one_subject_to_experiment_and_assign(X[i, , drop = FALSE]) mu_i = exp(0.5 + true_log_rr * w_i + 0.15 * X$baseline_rate[i] + 0.3 * X$urban[i]) des_seq$add_one_subject_response(i, rpois(1, mu_i)) } inf_seq = InferenceCountPoisson$new(des_seq, verbose = FALSE) inf_seq$num_cores = 1L inf_seq$compute_estimate() inf_seq$set_seed(1) inf_seq$compute_rand_two_sided_pval(r = 200, show_progress = FALSE) ``` ## Where to go next - Zero-heavy data: `InferenceCountHurdlePoisson`, `InferenceCountHurdleNegBin`, `InferenceCountZeroInflatedPoisson`, `InferenceCountZeroInflatedNegBin`. - Exposure/offset (rates per unit time) is a planned addition — see `ROADMAP.md`. - `vignette("validation-evidence")`: each kernel is checked against `stats::glm`, `MASS::glm.nb`, or `pscl`.