Getting started with fdb

Fit a hybrid-control trial

fdb fits penalized Cox models with a treatment coefficient theta and an external-versus-concurrent control drift coefficient delta. Both are on the log hazard-ratio scale. Penalizing drift toward zero introduces borrowing.

library(fdb)
set.seed(41)
dat <- simulate_hybrid_cox(nI1 = 40, nI0 = 40, nE = 80,
                           theta0 = log(0.8), delta0 = log(1.1), p = 0)$data
head(dat)
#>        time status T Z
#> 1 45.062050      1 1 0
#> 2  1.650595      1 1 0
#> 3 19.001591      1 1 0
#> 4 54.766561      1 1 0
#> 5  7.245401      1 1 0
#> 6  2.087950      1 1 0

The no-covariate model uses p = 0. Direct fitting expects time, status, T (treatment indicator), Z (external indicator), and optional covariates named X1, X2, and so on.

fit <- fit_one_penalized_method(dat, method = "P1", lambda = 0.2)
fit[c("method", "theta_hat", "delta_hat", "se_theta", "se_sand")]
#> $method
#> [1] "P1_SEScaledL1"
#> 
#> $theta_hat
#> [1] 0.196989
#> 
#> $delta_hat
#> [1] 0.2790806
#> 
#> $se_theta
#> [1] 0.1920929
#> 
#> $se_sand
#> [1] 0.233306

This tuning value illustrates the API; it has not been calibrated. fit_all_methods(dat) compares internal-only, pooled, Li, and P1–P4 fits. P2 uses the integrated gate and P3 uses smoothed MCP. Record eps, rho_mcp, penalty parameters and optimization bounds with results.

Interpret uncertainty

se_theta from the default profiled fit treats estimated drift as a fixed offset and can underestimate uncertainty. se_sand is a local plug-in sandwich approximation holding adaptive first-stage quantities fixed. It does not guarantee nominal conditional or unconditional coverage. Clipped curvature changes variance calculations, not fitted coefficients. Check coverage, bias, RMSE and missing-result counts in the design under study.

Calibrate a design

The larger examples below are not evaluated when building this vignette. Choose drift grids and an error-rate criterion before studying power.

cal <- calibrate_lambda_grid_two_stage(
  method = "P1", lambda_grid_coarse = c(0.02, 0.05, 0.1, 0.2, 0.5),
  scenario_base = scenario_S1, drift_set_cal = log(c(1, 1.05, 1.1)),
  nsim_cal = 10000, nsim_confirm = 10000,
  alpha = 0.025, alpha_cal = 0.04, seed = 81)
cal$lambda_star
cal$confirm$details

Confirmation defaults to the calibration grid. Set drift_set_confirm explicitly to confirm on a different grid. An NA selection means no tested candidate qualified; it is not permission to use an uncalibrated default. Inspect valid counts and distinguish failed fits from rates above the threshold.

Finite-grid calibration has Monte Carlo error. A threshold of 0.04 permits more rejection than a nominal 0.025 test. Neither a passing point estimate nor the pointwise normal-approximation Monte Carlo upper bound establishes simultaneous control across a continuum of drift values. The upper bound is especially approximate for few or zero rejections; use adequate replication.

Retain replicate estimates

run_simulation() always returns $raw. Curve summaries retain estimates only when requested:

curve <- run_drift_curve(
  theta0 = 0, drift_set = log(c(1, 1.1)), scenario_base = scenario_S1,
  lambdas = lambdas_default, nsim = 10000, seed = 71, keep_raw = TRUE)
curve$summary
head(curve$raw)
saveRDS(curve, "curve_with_replicates.rds")  # Choose your own output path.

Default lambdas illustrate the workflow and are not calibrated for every design. n_valid and n_missing identify results used in summary rates. mcse_rej_rate and mcse_coverage describe precision conditional on fixed tuning; they do not include calibration uncertainty. Pair estimates by method, drift_index and sim. Do not pool drift scenarios for a normality check.

ESS is the variance-equivalent gain NS * (variance_internal / variance_method - 1), with NS = nI1 + nI0 in the wrappers. Negative values indicate greater variance than internal-only analysis. ESS ignores bias and need not equal a literal number of borrowed controls. Paired bootstrap intervals can quantify its Monte Carlo uncertainty.

Parallel execution

Install the package before using workers. Parallel execution is opt-in and defaults to two workers. Set ncores explicitly for larger local runs; examples and package checks must use at most two. Workers use the package library selected by the parent session. Fixed seeds reproduce a fixed configuration, but changing worker counts or switching between serial and parallel execution can change simulated samples.

Full studies should run outside package checks. Save the package version, sessionInfo(), tuning, grids, seeds and core count with the results.