Differentially private synthesis (Track B)

Two tracks, one interface

flexsynth has two engines. Track A (the default) maximises statistical utility and carries no formal privacy guarantee — you judge the residual risk with diagnose() and disclosure_risk(). Track B trades some utility for a mathematically provable guarantee: differential privacy (DP). You opt in by passing a dp_control() to the privacy argument of synth(); nothing else about the call changes.

dp <- dp_control(epsilon = 1, delta = 1e-6, mechanism = "gaussian")
dp
#> <dp_control> differentially private synthesis (Track B)
#>   epsilon   : 1 
#>   delta     : 1e-06 
#>   unit      : person 
#>   mechanism : gaussian 
#>   dependence: tree 
#>   bins      : 12 
#>   domain    : dp (0.1 of budget for edges) 
#>   max rows/person: 1 (auto) 

Differential privacy bounds how much any single individual can change the output distribution. The guarantee is stated as a pair ((\epsilon), (\delta)): smaller values mean stronger privacy and more noise. The default privacy unit here is person-level — the guarantee protects an individual across all of their rows, not just a single record.

A worked example

We use a small synthetic cardiac dataset (never real patient data): one row per patient, with an age, sex, a smoking flag and a systolic blood pressure that depends on both.

n <- 1500
real <- data.frame(
  id     = seq_len(n),
  age    = round(rnorm(n, 62, 11)),
  sex    = factor(sample(c("F", "M"), n, TRUE, prob = c(0.45, 0.55))),
  smoker = sample(c(FALSE, TRUE), n, TRUE, prob = c(0.7, 0.3))
)
real$sbp <- round(0.6 * real$age + ifelse(real$smoker, 8, 0) + rnorm(n, 90, 10))
head(real)
id age sex smoker sbp
1 55 F TRUE 138
2 64 M FALSE 132
3 53 M FALSE 135
4 80 M FALSE 130
5 66 F FALSE 114
6 53 F TRUE 139

Where the bin edges come from

The DP engine works over a discrete grid: numeric variables are cut into bins. Where those bin edges come from matters, because reading them straight from the data (its min and max) can betray an individual’s presence — an outlier’s exact value is the maximum. The domain argument controls this, and the default keeps the accounting honest without any work on your part:

Categorical columns carry their domain in their type, so pass them as factor (or logical) — their levels are public metadata. A bare character column has its category set discovered under differential privacy by DP set-union under the default domain = "dp": each present category’s count is noised and kept only if it clears a threshold that hides any category a single person could have created, and rare categories fold into an "(other)" catch-all. The cost comes out of the same domain_frac slice as numeric bin-edge estimation. Because a threshold cannot hide a lone category’s presence at delta = 0, this needs delta > 0 (a pure-(\epsilon) release, or domain = "public", still requires public factor levels).

df_txt <- data.frame(id = seq_len(400),
                     age = round(rnorm(400, 60, 8)),
                     site = sample(c("north", "south", "east", "west"), 400, TRUE),
                     stringsAsFactors = FALSE)          # site is bare character
dp_txt <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
                     bounds = list(age = c(18, 100)))
synth(df_txt, ~ id, privacy = dp_txt, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.442 zCDP) 
#>   model     : tree over 2 variables  
#>   marginals : 3 (measured under composed budget)
#>   noise     : Gaussian sd 1.843 per cell  
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)
#>   categories: DP set-union discovered site (4 kept) (per-op eps 0.6, delta 1e-07)

For a clean, fully data-independent release with no budget spent on the domain, pass public ranges via bounds:

dp <- dp_control(
  epsilon    = 2,
  delta      = 1e-6,
  mechanism  = "gaussian",
  bounds     = list(age = c(18, 100), sbp = c(60, 240))
)
res <- synth(real, structure = ~ id, privacy = dp, seed = 1)
res
#> <synth_result>
#>   track        : B (differentially private; epsilon = 2, delta = 1e-06, person level)
#>   datasets (m) : 1 
#>   rows each    : 1514  (input: 1500 )
#>   synthesised  : age, sex, smoker, sbp 
#>   method       : dp-gaussian 
#> 
#> Get the data with as.data.frame(x)  

The result is an ordinary synth_result; as.data.frame() gives the synthetic table, whose id is a fresh surrogate key.

syn <- as.data.frame(res)
head(syn)
id age sex smoker sbp
1 61 M FALSE 116
2 61 F FALSE 128
3 95 F TRUE 152
4 70 F FALSE 160
5 57 M TRUE 166
6 48 M FALSE 119

The privacy accounting

Every DP result carries an accounting record describing exactly what was spent and how the noise was calibrated. This is the object to keep with a governed release.

res$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 2 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.0676 zCDP) 
#>   model     : tree over 4 variables  
#>   marginals : 10 (measured under composed budget)
#>   noise     : Gaussian sd 8.602 per cell  
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

The engine measures a set of low-order marginals under a single, composed budget: one-way marginals for every variable and (with the default dependence = "tree") the pairwise marginals needed to learn a Chow-Liu dependency tree. The number of marginals, the per-cell noise scale, and — for the Gaussian mechanism — the zero-concentrated-DP (\rho) are all reported.

Or let the engine estimate them privately

If you have no public ranges, the default domain = "dp" estimates each numeric variable’s working range under DP — a clamp-free exponential-mechanism quantile at each end — and charges it to the budget (a small slice, domain_frac, default 10%). No warning, and the reported ((\epsilon), (\delta)) is exact:

dp_auto <- dp_control(epsilon = 2, delta = 1e-6, mechanism = "gaussian")
res_auto <- synth(real, structure = ~ id, privacy = dp_auto, seed = 1)
res_auto$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 2 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.0676 zCDP) 
#>   model     : tree over 4 variables  
#>   marginals : 10 (measured under composed budget)
#>   noise     : Gaussian sd 9.067 per cell  
#>   row cap   : 1 per person  
#>   bin edges : dp - estimated under DP for age, sbp (0.1 of budget) 

The accounting names the variables whose edges were estimated and the fraction of budget it took. The remaining budget pays for the marginals, exactly as before.

Choices that matter

Mechanism. "laplace" gives pure (\epsilon)-DP (set delta = 0); "gaussian" gives approximate ((\epsilon), (\delta))-DP with tighter composition when many marginals are measured (it needs delta > 0).

Dependence. "tree" (default) keeps first- and second-order structure by fitting a Chow-Liu tree over the noisy pairwise marginals — no extra budget, because the same measurements are reused for structure and parameters. "independent" measures only one-way marginals: fewer queries, so less noise per query, but cross-variable correlations are lost. The real age/sbp correlation here is about r round(cor(real$age, real$sbp), 2). Under DP the tree recovers a substantial part of it while the independent model necessarily drives it to zero:

b <- list(age = c(18, 100), sbp = c(60, 240))
dp_tree  <- dp_control(epsilon = 6, mechanism = "laplace", dependence = "tree",
                       bins = 8, bounds = b)
dp_indep <- dp_control(epsilon = 6, mechanism = "laplace", dependence = "independent",
                       bins = 8, bounds = b)
s_tree  <- as.data.frame(synth(real, ~ id, privacy = dp_tree,  seed = 1))
s_indep <- as.data.frame(synth(real, ~ id, privacy = dp_indep, seed = 1))
c(real  = cor(real$age, real$sbp),
  tree  = cor(s_tree$age,  s_tree$sbp),
  indep = cor(s_indep$age, s_indep$sbp))
#>        real        tree       indep 
#> 0.541373105 0.268146206 0.006322181 

Two things are visible in that vector: the tree keeps a clear positive association (never the full strength — DP noise attenuates it), and the independent model has no correlation at all by construction. The gap between the tree and the real value is the honest price of the guarantee; a larger epsilon or a coarser bins narrows it.

Budget-efficient structure learning. With dependence = "tree" the default fitter measures every pairwise marginal at full fidelity, then keeps only the (d-1) that end up in the Chow-Liu tree — so budget was spent measuring pairs the model discards. structure_frac splits the work instead: a small fraction of the marginal budget buys a deliberately rough all-pairs scan used only to pick the tree, and the rest is concentrated on re-measuring just the chosen edges. Because structure selection tolerates noise far better than the conditional parameters do, this sharpens the surviving edges — increasingly so as the number of variables grows (with only a handful it is roughly break-even, and it is inert below three). Both passes are reported and compose into the same exact budget — the cheap scan and the concentrated re-measurement are two sequential releases whose pure-(\epsilon) adds (or, for Gaussian, whose zCDP (\rho) adds):

dp_eff <- dp_control(epsilon = 6, mechanism = "laplace", dependence = "tree",
                     bins = 8, bounds = b, structure_frac = 0.25)
synth(real, ~ id, privacy = dp_eff, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , pure eps )-DP, person level
#>   mechanism : laplace  
#>   model     : tree over 4 variables  
#>   marginals : 13 (6 pairwise scans + 7 parameter marginals, composed budget)
#>   noise     : Laplace scale 1.556 per cell (parameters) 
#>   structure : budget-efficient, 0.25 of budget selects the tree; scan Laplace scale 4 per cell
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

The accounting line splits into the pairwise scans and the parameter marginals, and the per-cell noise is reported separately for each — the scan is noisier (it only has to rank dependencies), the parameters sharper.

Adaptive marginal selection (AIM-style). structure_frac still measures a fixed menu of marginals (all pairs) and then discards most; it also cannot escape the tree — every dependency is squeezed onto (d-1) edges. select = "adaptive" takes the next step: after the one-way marginals it grows the model one marginal at a time, each round using the exponential mechanism to privately pick the marginal the model-so-far fits worst, then measuring that one. Selection spends budget too (it reads the true data to score candidates), so it and the measurements compose into the same exact ((\epsilon, \delta)); select_frac sets the split, and the round count is fixed in advance so the accounting stays data-independent.

The real payoff is treewidth. At treewidth = 1 the selected marginals form a spanning tree (the same model class as dependence = "tree", just chosen adaptively). At treewidth = 2 they form triangles, so the model can hold a genuine three-way interaction — the kind where a variable depends on a pair of others while looking independent of each one alone, which no tree (and so no amount of structure_frac) can represent:

dp_aim <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
                     select = "adaptive", treewidth = 2,
                     bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : adaptive junction tree (treewidth 2) over 4 variables  
#>   marginals : 6 (4 one-way + 2 adaptively selected cliques, composed budget)
#>   noise     : Gaussian sd 2.723 per cell  
#>   selection : adaptive (AIM-style), 0.25 of budget over 2 exponential-mechanism round(s); per-round eps 0.3673 
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

The accounting names the model an adaptive junction tree of the chosen treewidth, splits the histogram count into one-ways plus adaptively selected cliques, and reports the selection slice (how much budget, over how many exponential-mechanism rounds) separately from the measurement noise. treewidth = 3 goes one level further, to four-way cliques that hold interactions no three-way marginal can see (a 3-bit parity is the textbook case); the cost is cell sparsity, and the control warns when a clique’s cell count grows large enough for the per-cell noise to dominate. Adaptive selection is flat-table only for now — it is refused on a longitudinal or linked release.

Budget annealing (anneal = TRUE). By default the adaptive selector runs a fixed d - treewidth rounds with a uniform per-round budget. anneal = TRUE makes the schedule data-adaptive, in the spirit of AIM: the one-way marginals take a fair fixed share, then each clique round starts at a small quantum (large noise) and the budget doubles whenever a round’s measured signal fails to beat its noise floor. Once the mandatory spanning cliques are in place — so every variable is still covered and the sampler needs no PGM inference — any surplus budget is spent on extra rounds that re-measure the worst-fit clique, combined by inverse-variance weighting so the privacy cost adds exactly. The final round absorbs the exact remainder, so the release is still precisely ((\epsilon, \delta)), now over a variable number of rounds that the data chooses:

dp_anneal <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
                        select = "adaptive", treewidth = 2, anneal = TRUE,
                        bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_anneal, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : adaptive junction tree (treewidth 2) over 4 variables  
#>   marginals : 8 (4 one-way + 2 spanning + 2 refinement clique(s), composed budget)
#>   noise     : Gaussian sd 2.723 - 3.851 per cell (annealed range)
#>   selection : annealed (AIM-style), 0.25 of budget; 4 round(s) = 2 spanning + 2 refinement; sigma halved 0x 
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

The accounting now reports the realised schedule — total rounds split into spanning and refinement, how many times the budget doubled, and the annealed noise range. Because the model stays a spanning junction tree (no PGM inference), refinement can only sharpen the cliques it already measured, not add loopy marginals; annealing helps most when there are few variables and the fixed schedule would otherwise leave budget unspent.

PrivBayes networks (degree = k). A Chow-Liu tree is a degree-1 Bayesian network: every variable conditions on one parent. dp_control(dependence = "tree", degree = k) opts into GreedyBayes — a degree-k network in which each variable may condition on up to k of the already-generated variables, its parent set chosen greedily with the exponential mechanism (scored by the parents-to-node association, so internally-correlated parents are not preferred over the parents a node truly depends on). Unlike the adaptive junction tree — whose parents must lie inside one existing clique — a degree-k network can give a node any k predecessors, so it captures a variable that depends on two otherwise-unrelated parents (a v-structure) that a tree cannot represent at any budget:

dp_bayes <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
                       dependence = "tree", degree = 2,
                       bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_bayes, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : degree 2 Bayesian network (GreedyBayes) over 4 variables  
#>   marginals : 7 (4 one-way + 3 family joints, composed budget)
#>   noise     : Gaussian sd 2.941 per cell  
#>   selection : GreedyBayes, 0.25 of budget over 3 exponential-mechanism round(s); per-round eps 0.2999 
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

It measures the d one-way marginals plus one (parents, node) family joint per non-root node (2d - 1 marginals) and spends a select_frac slice on the d - 1 greedy picks; every slice composes into the same exact ((\epsilon, \delta)). The network is forward-sampled ancestrally, so like the tree it needs no PGM inference. degree is capped to d - 1, and — being a distinct structure search — it is an alternative to structure_frac and to select = "adaptive" (setting either alongside degree > 1 is an error). Raising degree makes each family a (degree + 1)-way histogram, so the same cell-count caution as a high treewidth applies.

Private-PGM inference (estimator = "pgm"). Every model above is PGM-free: it uses each measured marginal locally — the tree takes the root’s one-way and each edge’s raw noisy 2-way as (P(\text{child}\mid\text{parent})), the adaptive junction tree each clique’s own array — so the other one-way marginals, and the fact that overlapping noisy marginals disagree (each measured independently under noise), are thrown away. dp_control(estimator = "pgm") adds the reconciliation they omit, following McKenna et al.’s Private-PGM / MST: the whole measured set is reconciled into the single graphical-model distribution that best fits all of it at once (least squares), by belief propagation on a junction tree of the measured cliques plus entropic mirror descent, and the model is sampled from that. Because a tree (or the adaptive junction tree) has bounded treewidth, the inference is exact and cheap.

dp_pgm <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
                     dependence = "tree", estimator = "pgm",
                     bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_pgm, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : tree over 4 variables  
#>   marginals : 10 (measured under composed budget)
#>   noise     : Gaussian sd 3.044 per cell  
#>   estimator : Private-PGM reconciliation of the measured marginals
#>               (belief propagation + mirror descent; post-processing, no extra budget)
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

Crucially, reconciliation is pure post-processing of the already-privatised marginals, so it spends no extra budget — the ((\epsilon, \delta)) is identical to the same release with the default estimator = "local"; only the fitted model changes. It denoises (overlapping marginals are made mutually consistent) and lets the otherwise-discarded one-way marginals constrain the model, usually sharpening both the marginals and the conditionals at the same budget. It is available for the flat dependence = "tree" release and for select = "adaptive", and is an alternative to structure_frac, to degree > 1, and to anneal = TRUE.

Full AIM (select = "aim"). The adaptive selector above is deliberately constrained: every new clique attaches to an existing one and covers a fresh variable, so the measured cliques always form a junction tree and the model is forward-sampled with no inference. The price is that it can never measure a marginal between two variables that are both already in the model — so it cannot represent a loop of pairwise dependence (a cycle (A!-!B!-!C!-!A)), which no tree-shaped junction structure holds at any budget. select = "aim" removes that constraint on selection: the exponential mechanism may pick any pair, loops included. A loopy measured set has no forward sampler, so Private-PGM inference stops being optional and becomes the estimator — the whole set (the one-ways plus the selected pairs) is reconciled over a triangulated junction tree and sampled from that.

dp_aim <- dp_control(epsilon = 8, delta = 1e-6, mechanism = "gaussian",
                     select = "aim", treewidth = 2,
                     bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 8 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.91 zCDP) 
#>   model     : Full AIM graphical model (treewidth 2, Private-PGM) over 4 variables  
#>   marginals : 10 (4 one-way + 6 adaptively selected pairs, loopy, composed budget)
#>   noise     : Gaussian sd 2.707 per cell  
#>   selection : Full AIM (loopy marginals), 0.25 of budget over 6 exponential-mechanism round(s); per-round eps 0.2753 
#>   estimator : Private-PGM reconciliation over the triangulated model
#>               (belief propagation + mirror descent; post-processing, no extra budget)
#>   scoring   : model-projection (candidates scored against the reconciled model; post-processing, default)
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

It runs a data-independent (\min!\big(\binom{d}{2},, \text{treewidth}\times(d-1)\big)) selection rounds, each rejecting any new pair whose triangulated clique would exceed treewidth + 1 (so the model’s treewidth — and hence the inference cost — stays bounded); at treewidth = 1 no loop can close, so it reduces to an adaptively-selected tree. Selection and measurement compose to the same exact ((\epsilon, \delta)), the reconciliation is budget-neutral, and it is flat-table only. Use it when you suspect cyclic or higher-order dependence that a single tree cannot hold; keep treewidth small, since each clique is a bins^(treewidth+1) histogram.

Adding anneal = TRUE swaps the fixed round count for the same data-adaptive (\sigma)-halving schedule the adaptive selector offers: a baseline of treewidth-capped new loopy pairs is selected first, then any surplus budget re-measures the worst-fit already-measured pair (inverse-variance combined), with the per-round budget doubling whenever a measurement fails its noise floor. The annealed set is still triangulated and reconciled with Private-PGM, and the total spend is exactly the same ((\epsilon, \delta)) over a variable number of rounds — the privacy record reports the realised schedule.

dp_aim_anneal <- dp_control(epsilon = 8, delta = 1e-6, mechanism = "gaussian",
                            select = "aim", treewidth = 2, anneal = TRUE,
                            bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim_anneal, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 8 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.91 zCDP) 
#>   model     : Full AIM graphical model (treewidth 2, Private-PGM) over 4 variables  
#>   marginals : 15 (4 one-way + 5 loopy + 6 refinement pair(s), composed budget)
#>   noise     : Gaussian sd 2.707 - 3.828 per cell (annealed range)
#>   selection : annealed Full AIM (loopy marginals), 0.25 of budget; 11 round(s) = 5 loopy + 6 refinement; sigma halved 1x 
#>   estimator : Private-PGM reconciliation over the triangulated model
#>               (belief propagation + mirror descent; post-processing, no extra budget)
#>   scoring   : model-projection (candidates scored against the reconciled model; post-processing, default)
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

Model-projection scoring (the default). By default Full AIM uses AIM’s actual quality function: each candidate pair is scored against the current reconciled model’s own marginal over that pair, so a loopy pair the model already explains — through the marginals measured so far — no longer looks surprising, and the budget is steered to the genuinely worst-fit interaction. That reference is read from the already-privatised marginals (reconciled each round and projected onto the candidate, which for a not-yet-measured pair crosses cliques of the junction tree), so it is pure post-processing: the exponential mechanism’s sensitivity and the exact ((\epsilon, \delta)) are unchanged — only which marginals get selected changes. It composes with anneal = TRUE and costs a reconciliation per selection round. The cheaper one-way-product reference is still available as scoring = "independence".

# The default (scoring = "model") is shown above; this opts back to the cheaper
# one-way-product reference.
dp_aim_indep <- dp_control(epsilon = 8, delta = 1e-6, mechanism = "gaussian",
                           select = "aim", treewidth = 2, scoring = "independence",
                           bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim_indep, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 8 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.91 zCDP) 
#>   model     : Full AIM graphical model (treewidth 2, Private-PGM) over 4 variables  
#>   marginals : 10 (4 one-way + 6 adaptively selected pairs, loopy, composed budget)
#>   noise     : Gaussian sd 2.707 per cell  
#>   selection : Full AIM (loopy marginals), 0.25 of budget over 6 exponential-mechanism round(s); per-round eps 0.2753 
#>   estimator : Private-PGM reconciliation over the triangulated model
#>               (belief propagation + mirror descent; post-processing, no extra budget)
#>   scoring   : independence (one-way product; model-projection disabled)
#>   row cap   : 1 per person  
#>   bin edges : dp (public bounds; no budget spent)

Contribution bound. At unit = "person" each person may contribute at most max_rows_per_person rows (default 1, appropriate for one-row-per-person tables). If a person legitimately has several rows, set this from public domain knowledge; rows beyond the cap are subsampled away before any budget is spent, so the guarantee holds by construction.

More budget, more fidelity. Larger epsilon means less noise. The synthetic marginals converge to the real ones as (\epsilon) grows:

for (e in c(0.5, 2, 8)) {
  d <- dp_control(epsilon = e, mechanism = "laplace",
                  dependence = "independent",
                  bounds = list(age = c(18, 100), sbp = c(60, 240)))
  s <- as.data.frame(synth(real, ~ id, privacy = d, seed = 1))
  cat(sprintf("epsilon = %-3g  mean(sbp): real %.1f  syn %.1f\n",
              e, mean(real$sbp), mean(s$sbp)))
}
#> epsilon = 0.5  mean(sbp): real 129.7  syn 129.3
#> epsilon = 2    mean(sbp): real 129.7  syn 130.1
#> epsilon = 8    mean(sbp): real 129.7  syn 130.4

Longitudinal releases: a DP Markov model

If the structure declares a nesting index (~ id / visit), Track B preserves within-unit temporal structure with a first-order Markov model — the private analogue of Track A’s initial-state + lag-1 transition model. Three things are measured under one composed budget: a length histogram (how many rows a person contributes), the initial-state marginals (the t = 1 row, one- and two-way as before), and a transition matrix (P(v_t \mid v_{t-1})) for each variable. Generation draws a length, an initial row, then steps each variable’s transition matrix, so autocorrelation across visits is retained.

Because a person now contributes a whole trajectory, you must set max_rows_per_person to the public maximum number of visits — this bounds each person’s effect on the transition histograms (a length-(\le c) trajectory has at most (c - 1) consecutive pairs). The budget still composes to exactly ((\epsilon), (\delta)); the accounting reports the split.

long <- do.call(rbind, lapply(1:400, function(i) {
  nv <- sample(2:4, 1); s <- numeric(nv); s[1] <- rnorm(1, 130, 12)
  for (t in seq_len(nv)[-1]) s[t] <- 0.85 * s[t - 1] + 0.15 * 130 + rnorm(1, 0, 5)
  data.frame(id = i, visit = seq_len(nv), sbp = round(s),
             sex = sample(c("F", "M"), 1))
}))
long$sex <- factor(long$sex)

dp_long <- dp_control(epsilon = 8, mechanism = "laplace",
                      max_rows_per_person = 4, bounds = list(sbp = c(60, 240)))
res_long <- synth(long, structure = ~ id / visit, privacy = dp_long, seed = 1)
res_long$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 8 , pure eps )-DP, person level
#>   mechanism : laplace  
#>   model     : tree over 2 variables (DP Markov: initial-state + transitions) 
#>   histograms: 6 (1 length + 3 initial + 2 transition, composed budget)
#>   noise     : Laplace scale 1.25 per cell  
#>   row cap   : 4 per person  
#>   bin edges : dp (public bounds; no budget spent)

The synthetic visit index is regenerated as the within-person position, and a lag-1 correlation like the real one survives the noise (DP attenuates it):

lag1 <- function(d) {
  d <- d[order(d$id, d$visit), ]
  prev <- ave(d$sbp, d$id, FUN = function(x) c(NA, head(x, -1)))
  ok <- !is.na(prev); cor(prev[ok], d$sbp[ok])
}
syn_long <- as.data.frame(res_long)
c(real = lag1(long), synthetic = lag1(syn_long))
#>      real synthetic 
#> 0.8922741 0.3424675 

Baseline columns held exactly constant

By default every non-index column is treated as time-varying, so a genuinely subject-invariant covariate (here sex) is stepped through its own transition matrix and can drift between a synthetic person’s visits. If you know a column is constant within a person — public schema knowledge, not something read from the data — name it in baseline. It is then modelled once in the initial-state model (keeping its distribution and its correlation with the first visit) and broadcast unchanged to every row, so it is exactly constant within each synthetic unit. Because a baseline column needs no transition histogram, declaring it also removes that histogram from the release, so the remaining measurements are sharper at the same ((\epsilon), (\delta)):

res_base <- synth(long, structure = ~ id / visit,
                  privacy = dp_control(epsilon = 8, mechanism = "laplace",
                                       max_rows_per_person = 4,
                                       bounds = list(sbp = c(60, 240)),
                                       baseline = "sex"),
                  seed = 1)

# every synthetic person now has a single sex across their visits
syn_base <- as.data.frame(res_base)
max(tapply(syn_base$sex, syn_base$id, function(s) length(unique(s))))
#> [1] 1
res_base$privacy   # one fewer transition histogram than the release above
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 8 , pure eps )-DP, person level
#>   mechanism : laplace  
#>   model     : tree over 2 variables (DP Markov: initial-state + transitions) 
#>   histograms: 5 (1 length + 3 initial + 1 transition, composed budget)
#>   baseline  : held constant within unit: sex 
#>   noise     : Laplace scale 0.875 per cell  
#>   row cap   : 4 per person  
#>   bin edges : dp (public bounds; no budget spent)

Higher-order and cross-variable transitions

The default transition model is first-order and per-variable: each variable’s next value depends only on its own previous value. Two knobs deepen it. transition_order = k conditions on a variable’s own last k values (momentum), and transition_cross = m additionally conditions on the lag-1 values of its m most strongly associated companions — so cross-variable coupling is re-measured at every step, not just carried at t = 1 by the initial-state tree. Companions are picked from the pairwise marginals the tree already measures, so cross-conditioning spends no extra budget (a transition tuple still lands in exactly one cell); a higher order actually lowers the transition sensitivity to cap - order, since a person then contributes fewer, deeper tuples. The order must therefore be at most max_rows_per_person - 1, and transition_cross > 0 needs the tree model.

res_ho <- synth(long, structure = ~ id / visit,
                privacy = dp_control(epsilon = 8, mechanism = "laplace",
                                     dependence = "tree", max_rows_per_person = 4,
                                     bounds = list(sbp = c(60, 240)),
                                     baseline = "sex",
                                     transition_order = 2, transition_cross = 1),
                seed = 1)
res_ho$privacy   # note: transitions order 2 + 1 cross-parent, same (eps, delta)
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 8 , pure eps )-DP, person level
#>   mechanism : laplace  
#>   model     : tree over 2 variables (DP Markov: initial-state + transitions) 
#>   histograms: 5 (1 length + 3 initial + 1 transition, composed budget)
#>   baseline  : held constant within unit: sex 
#>   transitions: order 2 + 1 cross-parent(s) (sensitivity cap - order)
#>                cross-parents: sbp ~ sex 
#>   noise     : Laplace scale 0.75 per cell  
#>   row cap   : 4 per person  
#>   bin edges : dp (public bounds; no budget spent)

Early rows (before position order + 1, which have no full history yet) are generated by marginalising the very same measured tensor, which is post-processing and costs nothing. The trade-off is cell sparsity: a wider conditioning grid spreads the same noisy counts more thinly, so reach for depth only when the extra structure is worth it.

Linked multi-table releases

synth_linked() accepts a dp_control() too, giving a DP release across a whole key hierarchy at once. The privacy unit is the root entity (e.g. a patient): adding or removing one individual — its root row and every descendant row that cascades from it — changes the release within the budget. Contribution is bounded hierarchically, so max_rows_per_person here means the maximum children kept per parent for each child table: a single integer for every child table, or a named list keyed by table name. The root cap is always 1.

set.seed(1)
patients <- data.frame(
  id  = 1:300,
  age = round(rnorm(300, 60, 11)),
  sex = factor(sample(c("F", "M"), 300, TRUE)))
adm <- do.call(rbind, lapply(patients$id, function(pid) {
  n <- rpois(1, 1.4); if (n == 0) return(NULL)
  data.frame(id = pid, admission_id = seq_len(n), los = 1L + rpois(n, 4))
}))

dp_link <- dp_control(
  epsilon = 4, mechanism = "laplace",
  max_rows_per_person = list(admissions = 6),        # <= 6 admissions per patient
  domain = "public",
  bounds = list(age = c(18, 100), los = c(0, 60)))

res_link <- synth_linked(
  tables     = list(patients = patients, admissions = adm),
  structures = list(patients = ~ id, admissions = ~ id / admission_id),
  keys       = list(patients = "id", admissions = c("id", "admission_id")),
  privacy    = dp_link, seed = 1)
res_link$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 4 , pure eps )-DP, person level
#>   mechanism : laplace  
#>   model     : tree over 3 variables (linked: 2 tables) 
#>   histograms: 5 (per-table variable marginals + child count models, composed budget)
#>   noise     : Laplace scale 2.5 per cell  
#>     - patients       root, cap 1
#>     - admissions     child of patients, <=6/parent, path cap 6 (2 rows dropped)
#>   bin edges : public (public bounds; no budget spent)

Each table’s own variable marginals and a children-per-parent count histogram are measured under one exactly-composed budget; the summed L1 (Laplace) and summed squared L2 (Gaussian zCDP) sensitivities add over every histogram, with each table’s contribution scaled by its per-entity path cap (the product of branching caps from the root). Synthetic children copy their synthetic parent’s surrogate key, so referential integrity holds by construction — check_linkage() confirms it:

check_linkage(res_link)
#> Linkage OK: keys unique, no orphan child rows.

By default each child table’s variables are modelled by their own within-table marginals, so the synthetic child links to a synthetic parent but is statistically independent of it.

Conditioning children on the synthetic parent

dp_control(cross_table = TRUE) closes that gap. For each child table with a modellable immediate parent, the engine also measures parent-by-child joint marginals — counted at the child grain (one observation per child row, at the parent’s value carried down the foreign key). Their person-sensitivity is the child’s path cap, exactly like a child one-way marginal, so they fold into the same composed budget. The parent’s variables then enter the child’s Chow-Liu structure as fixed context nodes, and at generation the synthetic parent’s already-drawn value conditions the child draw:

dp_cross <- dp_control(
  epsilon = 4, mechanism = "laplace", cross_table = TRUE,
  max_rows_per_person = list(admissions = 6),
  domain = "public",
  bounds = list(age = c(18, 100), los = c(0, 60)))

res_cross <- synth_linked(
  tables     = list(patients = patients, admissions = adm),
  structures = list(patients = ~ id, admissions = ~ id / admission_id),
  keys       = list(patients = "id", admissions = c("id", "admission_id")),
  privacy    = dp_cross, seed = 1)
res_cross$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 4 , pure eps )-DP, person level
#>   mechanism : laplace  
#>   model     : tree over 3 variables (linked: 2 tables) 
#>   histograms: 7 (per-table variable marginals + child count models, composed budget)
#>   noise     : Laplace scale 5.5 per cell  
#>     - patients       root, cap 1
#>     - admissions     child of patients, <=6/parent, path cap 6, cond. on parent (2 rows dropped)
#>   bin edges : public (public bounds; no budget spent)

The accounting now lists admissions as cond. on parent and the extra joints raise the histogram count and the per-cell noise — the price of the added structure, still inside the same \((\epsilon, \delta)\). Each child variable’s single strongest predictor may be a parent variable or another child variable (under dependence = "independent" it is the best parent variable only), and a three-level hierarchy chains it: a grandchild conditions on its already-conditioned parent. Only the immediate parent conditions a child; deeper ancestors reach it through the parent’s synthesised values, as in Track A.

Only the immediate parent conditions a child; deeper ancestors reach it through the parent’s synthesised values, as in Track A. Constraints are refused, as elsewhere under DP.

Modelling a child table’s rows over time

When a child table’s rows are a time series within each parent unit — repeated visits ordered by the child’s own key index — dp_control(longitudinal = TRUE) (or a vector of child-table names) models them as a within-unit DP Markov trajectory instead of exchangeable records. The children-per-parent count model doubles as the trajectory-length model; an initial-state model is measured over each unit’s first (earliest) child row; and a first-order transition matrix \(P(v_t \mid v_{t-1})\) is measured per variable over consecutive within-unit rows:

set.seed(1)
visits <- do.call(rbind, lapply(patients$id, function(pid) {
  k  <- 2L + rpois(1, 1.2)
  st <- character(k); st[1] <- sample(c("stable", "worse"), 1)
  for (i in 2:k) st[i] <- if (runif(1) < 0.85) st[i - 1]
                          else setdiff(c("stable", "worse"), st[i - 1])
  data.frame(id = pid, visit_num = seq_len(k),
             status = factor(st, levels = c("stable", "worse")))
}))

dp_longi <- dp_control(
  epsilon = 6, delta = 1e-6, mechanism = "gaussian",
  max_rows_per_person = c(visits = 6), longitudinal = "visits",
  domain = "public", bounds = list(age = c(18, 100)))

res_longi <- synth_linked(
  tables     = list(patients = patients, visits = visits),
  structures = list(patients = ~ id, visits = ~ id / visit_num),
  keys       = list(patients = "id", visits = c("id", "visit_num")),
  privacy    = dp_longi, seed = 1)
res_longi$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : tree over 3 variables (linked: 2 tables) 
#>   histograms: 6 (per-table variable marginals + child count models, composed budget)
#>   noise     : Gaussian sd 5.273 per cell  
#>     - patients       root, cap 1
#>     - visits         child of patients, <=6/parent, path cap 6, DP Markov over rows (5 rows dropped)
#>   bin edges : public (public bounds; no budget spent)

The accounting lists visits as DP Markov over rows. The transition histograms have person-sensitivity path_cap[parent] * (branching_cap - 1) while the initial-state marginals sit at the parent path cap — heterogeneous, but folded into the same exact \((\epsilon, \delta)\). Over-cap units are prefix-truncated in temporal order so their consecutive pairs stay intact. This keeps a patient’s status trending across visits, the autocorrelation an exchangeable child would drop.

Combining cross-table conditioning with a longitudinal child

Setting cross_table = TRUE together with a longitudinal model on the same child combines the two: the child’s initial-state model is cross-conditioned on the synthetic parent — the first row of each unit draws from a parent-conditioned Chow-Liu tree — and the within-unit transition chain then carries that parent dependence across the trajectory. The parent shapes where a trajectory starts; the transitions stay parent-free, so the only extra cost is the nC * nP parent-by-child initial-state joints, at the same first-row sensitivity as the initial marginals:

dp_both <- dp_control(
  epsilon = 6, delta = 1e-6, mechanism = "gaussian",
  max_rows_per_person = c(visits = 6), longitudinal = "visits",
  cross_table = TRUE, domain = "public", bounds = list(age = c(18, 100)))

res_both <- synth_linked(
  tables     = list(patients = patients, visits = visits),
  structures = list(patients = ~ id, visits = ~ id / visit_num),
  keys       = list(patients = "id", visits = c("id", "visit_num")),
  privacy    = dp_both, seed = 1)
res_both$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : tree over 3 variables (linked: 2 tables) 
#>   histograms: 8 (per-table variable marginals + child count models, composed budget)
#>   noise     : Gaussian sd 5.446 per cell  
#>     - patients       root, cap 1
#>     - visits         child of patients, <=6/parent, path cap 6, DP Markov over rows (initial state cond. on parent) (5 rows dropped)
#>   bin edges : public (public bounds; no budget spent)

The accounting now lists visits as DP Markov over rows (initial state cond. on parent). A patient’s whole status trajectory can then inherit a dependence on their baseline attributes, not just autocorrelate with itself.

Baseline columns and deeper transitions on a linked child

The two within-unit transition controls of the flat DP Markov engine also apply, per table, to a longitudinally-modelled linked child. dp_control(baseline = c(...)) names subject-invariant columns of a child (matched against that child’s own columns) and holds them exactly constant within a unit — dropping their transition histograms — while dp_control(transition_order = k, transition_cross = m) deepens each time-varying column’s transition to its own last k values plus the lag-1 values of its m most associated companions. A higher order lowers the child’s transition sensitivity to path_cap[parent] * (branching_cap - order), so the order must be at most one less than the child’s branching cap:

dp_order2 <- dp_control(
  epsilon = 6, delta = 1e-6, mechanism = "gaussian",
  max_rows_per_person = c(visits = 6), longitudinal = "visits",
  transition_order = 2, domain = "public", bounds = list(age = c(18, 100)))

res_order2 <- synth_linked(
  tables     = list(patients = patients, visits = visits),
  structures = list(patients = ~ id, visits = ~ id / visit_num),
  keys       = list(patients = "id", visits = c("id", "visit_num")),
  privacy    = dp_order2, seed = 1)
res_order2$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : tree over 3 variables (linked: 2 tables) 
#>   histograms: 6 (per-table variable marginals + child count models, composed budget)
#>   noise     : Gaussian sd 4.411 per cell  
#>     - patients       root, cap 1
#>     - visits         child of patients, <=6/parent, path cap 6, DP Markov over rows (5 rows dropped)
#>         transitions: order 2 + 0 cross-parent(s) (sensitivity cap - order)
#>   bin edges : public (public bounds; no budget spent)

The visits line now reports transitions: order 2; with a baseline column the accounting would add a baseline held: ... line and charge one fewer transition histogram. Both settings compose with the combined cross-conditioned initial state above.

Anchoring the parent across the whole trajectory

The combined cross-conditioned model above cross-conditions only a child’s initial state; the parent’s influence then rides the own-lag chain and fades. dp_control(transition_parent = p) instead re-injects the synthetic parent’s subject-invariant attributes into the child’s transition at every step — each time-varying column conditions its next value on the p immediate-parent attributes most strongly associated with it, so parent → child dependence stays anchored across the trajectory. The parents are chosen budget-neutrally from the parent-by-child joints the cross-conditioned initial state already measures, so it adds no histogram and no sensitivity — which is exactly why it requires cross_table = TRUE:

dp_tp <- dp_control(
  epsilon = 6, delta = 1e-6, mechanism = "gaussian",
  max_rows_per_person = c(visits = 6), longitudinal = "visits",
  cross_table = TRUE, transition_parent = 1,
  domain = "public", bounds = list(age = c(18, 100)))

res_tp <- synth_linked(
  tables     = list(patients = patients, visits = visits),
  structures = list(patients = ~ id, visits = ~ id / visit_num),
  keys       = list(patients = "id", visits = c("id", "visit_num")),
  privacy    = dp_tp, seed = 1)
res_tp$privacy
#> <dp_accounting> Track B differential privacy
#>   guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#>   mechanism : gaussian (rho = 0.54 zCDP) 
#>   model     : tree over 3 variables (linked: 2 tables) 
#>   histograms: 8 (per-table variable marginals + child count models, composed budget)
#>   noise     : Gaussian sd 5.446 per cell  
#>     - patients       root, cap 1
#>     - visits         child of patients, <=6/parent, path cap 6, DP Markov over rows (initial state cond. on parent) (5 rows dropped)
#>         transitions: order 1 + 0 cross-parent(s) + 1 parent-attr(s) (sensitivity cap - order)
#>         parent-attrs: status ~ age 
#>   bin edges : public (public bounds; no budget spent)

The visits line reports initial state cond. on parent plus a `transitions: …

What Track B does not do (yet)

Within those limits the guarantee is exact and reported. For anything labelled differentially private, keep the accounting record with the release. The discretisation adds no unaccounted leakage in either rigorous mode — supply public bounds (domain = "public") to spend nothing on the domain, or let the default domain = "dp" estimate the edges privately and charge them to the budget.