flexsynth has two engines. Track A (the default) maximises statistical
utility and carries no formal privacy guarantee — you judge the residual risk
with diagnose() and disclosure_risk(). Track B trades some utility for a
mathematically provable guarantee: differential privacy (DP). You opt in by
passing a dp_control() to the privacy argument of synth(); nothing else
about the call changes.
dp <- dp_control(epsilon = 1, delta = 1e-6, mechanism = "gaussian")
dp
#> <dp_control> differentially private synthesis (Track B)
#> epsilon : 1
#> delta : 1e-06
#> unit : person
#> mechanism : gaussian
#> dependence: tree
#> bins : 12
#> domain : dp (0.1 of budget for edges)
#> max rows/person: 1 (auto)
Differential privacy bounds how much any single individual can change the output distribution. The guarantee is stated as a pair ((\epsilon), (\delta)): smaller values mean stronger privacy and more noise. The default privacy unit here is person-level — the guarantee protects an individual across all of their rows, not just a single record.
We use a small synthetic cardiac dataset (never real patient data): one row per patient, with an age, sex, a smoking flag and a systolic blood pressure that depends on both.
n <- 1500
real <- data.frame(
id = seq_len(n),
age = round(rnorm(n, 62, 11)),
sex = factor(sample(c("F", "M"), n, TRUE, prob = c(0.45, 0.55))),
smoker = sample(c(FALSE, TRUE), n, TRUE, prob = c(0.7, 0.3))
)
real$sbp <- round(0.6 * real$age + ifelse(real$smoker, 8, 0) + rnorm(n, 90, 10))
head(real)
| id | age | sex | smoker | sbp |
|---|---|---|---|---|
| 1 | 55 | F | TRUE | 138 |
| 2 | 64 | M | FALSE | 132 |
| 3 | 53 | M | FALSE | 135 |
| 4 | 80 | M | FALSE | 130 |
| 5 | 66 | F | FALSE | 114 |
| 6 | 53 | F | TRUE | 139 |
The DP engine works over a discrete grid: numeric variables are cut into
bins. Where those bin edges come from matters, because reading them straight from
the data (its min and max) can betray an individual’s presence — an outlier’s
exact value is the maximum. The domain argument controls this, and the
default keeps the accounting honest without any work on your part:
domain = "dp" (default): numeric variables you name in bounds use those
public edges for free; any others have their working range estimated under
differential privacy and the cost is folded into the reported
((\epsilon), (\delta)). Nothing leaks unaccounted.domain = "public": every numeric variable must have a public range in
bounds (an error otherwise); no budget is spent on the domain. Use this when
you have codebook or physiological limits.domain = "data": the old behaviour — edges read from the data with a warning,
excluded from the accounting. Benchmarking only.Categorical columns carry their domain in their type, so pass them as factor
(or logical) — their levels are public metadata. A bare character column has
its category set discovered under differential privacy by DP set-union under the
default domain = "dp": each present category’s count is noised and kept only if it
clears a threshold that hides any category a single person could have created, and
rare categories fold into an "(other)" catch-all. The cost comes out of the same
domain_frac slice as numeric bin-edge estimation. Because a threshold cannot hide a
lone category’s presence at delta = 0, this needs delta > 0 (a pure-(\epsilon)
release, or domain = "public", still requires public factor levels).
df_txt <- data.frame(id = seq_len(400),
age = round(rnorm(400, 60, 8)),
site = sample(c("north", "south", "east", "west"), 400, TRUE),
stringsAsFactors = FALSE) # site is bare character
dp_txt <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
bounds = list(age = c(18, 100)))
synth(df_txt, ~ id, privacy = dp_txt, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.442 zCDP)
#> model : tree over 2 variables
#> marginals : 3 (measured under composed budget)
#> noise : Gaussian sd 1.843 per cell
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
#> categories: DP set-union discovered site (4 kept) (per-op eps 0.6, delta 1e-07)
For a clean, fully data-independent release with no budget spent on the
domain, pass public ranges via bounds:
dp <- dp_control(
epsilon = 2,
delta = 1e-6,
mechanism = "gaussian",
bounds = list(age = c(18, 100), sbp = c(60, 240))
)
res <- synth(real, structure = ~ id, privacy = dp, seed = 1)
res
#> <synth_result>
#> track : B (differentially private; epsilon = 2, delta = 1e-06, person level)
#> datasets (m) : 1
#> rows each : 1514 (input: 1500 )
#> synthesised : age, sex, smoker, sbp
#> method : dp-gaussian
#>
#> Get the data with as.data.frame(x)
The result is an ordinary synth_result; as.data.frame() gives the synthetic
table, whose id is a fresh surrogate key.
syn <- as.data.frame(res)
head(syn)
| id | age | sex | smoker | sbp |
|---|---|---|---|---|
| 1 | 61 | M | FALSE | 116 |
| 2 | 61 | F | FALSE | 128 |
| 3 | 95 | F | TRUE | 152 |
| 4 | 70 | F | FALSE | 160 |
| 5 | 57 | M | TRUE | 166 |
| 6 | 48 | M | FALSE | 119 |
Every DP result carries an accounting record describing exactly what was spent and how the noise was calibrated. This is the object to keep with a governed release.
res$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 2 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.0676 zCDP)
#> model : tree over 4 variables
#> marginals : 10 (measured under composed budget)
#> noise : Gaussian sd 8.602 per cell
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
The engine measures a set of low-order marginals under a single, composed budget:
one-way marginals for every variable and (with the default
dependence = "tree") the pairwise marginals needed to learn a Chow-Liu
dependency tree. The number of marginals, the per-cell noise scale, and — for the
Gaussian mechanism — the zero-concentrated-DP (\rho) are all reported.
If you have no public ranges, the default domain = "dp" estimates each numeric
variable’s working range under DP — a clamp-free exponential-mechanism quantile
at each end — and charges it to the budget (a small slice, domain_frac, default
10%). No warning, and the reported ((\epsilon), (\delta)) is exact:
dp_auto <- dp_control(epsilon = 2, delta = 1e-6, mechanism = "gaussian")
res_auto <- synth(real, structure = ~ id, privacy = dp_auto, seed = 1)
res_auto$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 2 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.0676 zCDP)
#> model : tree over 4 variables
#> marginals : 10 (measured under composed budget)
#> noise : Gaussian sd 9.067 per cell
#> row cap : 1 per person
#> bin edges : dp - estimated under DP for age, sbp (0.1 of budget)
The accounting names the variables whose edges were estimated and the fraction of budget it took. The remaining budget pays for the marginals, exactly as before.
Mechanism. "laplace" gives pure (\epsilon)-DP (set delta = 0);
"gaussian" gives approximate ((\epsilon), (\delta))-DP with tighter
composition when many marginals are measured (it needs delta > 0).
Dependence. "tree" (default) keeps first- and second-order structure by
fitting a Chow-Liu tree over the noisy pairwise marginals — no extra budget,
because the same measurements are reused for structure and parameters.
"independent" measures only one-way marginals: fewer queries, so less noise per
query, but cross-variable correlations are lost. The real age/sbp
correlation here is about r round(cor(real$age, real$sbp), 2). Under DP the tree
recovers a substantial part of it while the independent model necessarily drives
it to zero:
b <- list(age = c(18, 100), sbp = c(60, 240))
dp_tree <- dp_control(epsilon = 6, mechanism = "laplace", dependence = "tree",
bins = 8, bounds = b)
dp_indep <- dp_control(epsilon = 6, mechanism = "laplace", dependence = "independent",
bins = 8, bounds = b)
s_tree <- as.data.frame(synth(real, ~ id, privacy = dp_tree, seed = 1))
s_indep <- as.data.frame(synth(real, ~ id, privacy = dp_indep, seed = 1))
c(real = cor(real$age, real$sbp),
tree = cor(s_tree$age, s_tree$sbp),
indep = cor(s_indep$age, s_indep$sbp))
#> real tree indep
#> 0.541373105 0.268146206 0.006322181
Two things are visible in that vector: the tree keeps a clear positive
association (never the full strength — DP noise attenuates it), and the
independent model has no correlation at all by construction. The gap between the
tree and the real value is the honest price of the guarantee; a larger epsilon
or a coarser bins narrows it.
Budget-efficient structure learning. With dependence = "tree" the default
fitter measures every pairwise marginal at full fidelity, then keeps only the
(d-1) that end up in the Chow-Liu tree — so budget was spent measuring pairs the
model discards. structure_frac splits the work instead: a small fraction of the
marginal budget buys a deliberately rough all-pairs scan used only to pick the
tree, and the rest is concentrated on re-measuring just the chosen edges. Because
structure selection tolerates noise far better than the conditional parameters
do, this sharpens the surviving edges — increasingly so as the number of variables
grows (with only a handful it is roughly break-even, and it is inert below three).
Both passes are reported and compose into the same exact budget — the cheap scan
and the concentrated re-measurement are two sequential releases whose
pure-(\epsilon) adds (or, for Gaussian, whose zCDP (\rho) adds):
dp_eff <- dp_control(epsilon = 6, mechanism = "laplace", dependence = "tree",
bins = 8, bounds = b, structure_frac = 0.25)
synth(real, ~ id, privacy = dp_eff, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , pure eps )-DP, person level
#> mechanism : laplace
#> model : tree over 4 variables
#> marginals : 13 (6 pairwise scans + 7 parameter marginals, composed budget)
#> noise : Laplace scale 1.556 per cell (parameters)
#> structure : budget-efficient, 0.25 of budget selects the tree; scan Laplace scale 4 per cell
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
The accounting line splits into the pairwise scans and the parameter marginals, and the per-cell noise is reported separately for each — the scan is noisier (it only has to rank dependencies), the parameters sharper.
Adaptive marginal selection (AIM-style). structure_frac still measures a
fixed menu of marginals (all pairs) and then discards most; it also cannot
escape the tree — every dependency is squeezed onto (d-1) edges. select = "adaptive" takes the next step: after the one-way marginals it grows the model
one marginal at a time, each round using the exponential mechanism to
privately pick the marginal the model-so-far fits worst, then measuring that one.
Selection spends budget too (it reads the true data to score candidates), so it
and the measurements compose into the same exact ((\epsilon, \delta));
select_frac sets the split, and the round count is fixed in advance so the
accounting stays data-independent.
The real payoff is treewidth. At treewidth = 1 the selected marginals form a
spanning tree (the same model class as dependence = "tree", just chosen
adaptively). At treewidth = 2 they form triangles, so the model can hold a
genuine three-way interaction — the kind where a variable depends on a pair of
others while looking independent of each one alone, which no tree (and so no
amount of structure_frac) can represent:
dp_aim <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
select = "adaptive", treewidth = 2,
bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : adaptive junction tree (treewidth 2) over 4 variables
#> marginals : 6 (4 one-way + 2 adaptively selected cliques, composed budget)
#> noise : Gaussian sd 2.723 per cell
#> selection : adaptive (AIM-style), 0.25 of budget over 2 exponential-mechanism round(s); per-round eps 0.3673
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
The accounting names the model an adaptive junction tree of the chosen
treewidth, splits the histogram count into one-ways plus adaptively selected
cliques, and reports the selection slice (how much budget, over how many
exponential-mechanism rounds) separately from the measurement noise. treewidth = 3 goes one level further, to four-way cliques that hold interactions no three-way
marginal can see (a 3-bit parity is the textbook case); the cost is cell
sparsity, and the control warns when a clique’s cell count grows large enough for
the per-cell noise to dominate. Adaptive selection is flat-table only for now
— it is refused on a longitudinal or linked release.
Budget annealing (anneal = TRUE). By default the adaptive selector runs a
fixed d - treewidth rounds with a uniform per-round budget. anneal = TRUE
makes the schedule data-adaptive, in the spirit of AIM: the one-way marginals
take a fair fixed share, then each clique round starts at a small quantum (large
noise) and the budget doubles whenever a round’s measured signal fails to beat
its noise floor. Once the mandatory spanning cliques are in place — so every
variable is still covered and the sampler needs no PGM inference — any surplus
budget is spent on extra rounds that re-measure the worst-fit clique, combined by
inverse-variance weighting so the privacy cost adds exactly. The final round
absorbs the exact remainder, so the release is still precisely ((\epsilon,
\delta)), now over a variable number of rounds that the data chooses:
dp_anneal <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
select = "adaptive", treewidth = 2, anneal = TRUE,
bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_anneal, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : adaptive junction tree (treewidth 2) over 4 variables
#> marginals : 8 (4 one-way + 2 spanning + 2 refinement clique(s), composed budget)
#> noise : Gaussian sd 2.723 - 3.851 per cell (annealed range)
#> selection : annealed (AIM-style), 0.25 of budget; 4 round(s) = 2 spanning + 2 refinement; sigma halved 0x
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
The accounting now reports the realised schedule — total rounds split into spanning and refinement, how many times the budget doubled, and the annealed noise range. Because the model stays a spanning junction tree (no PGM inference), refinement can only sharpen the cliques it already measured, not add loopy marginals; annealing helps most when there are few variables and the fixed schedule would otherwise leave budget unspent.
PrivBayes networks (degree = k). A Chow-Liu tree is a degree-1 Bayesian
network: every variable conditions on one parent. dp_control(dependence = "tree", degree = k) opts into GreedyBayes — a degree-k network in which each
variable may condition on up to k of the already-generated variables, its parent
set chosen greedily with the exponential mechanism (scored by the parents-to-node
association, so internally-correlated parents are not preferred over the parents a
node truly depends on). Unlike the adaptive junction tree — whose parents must lie
inside one existing clique — a degree-k network can give a node any k
predecessors, so it captures a variable that depends on two otherwise-unrelated
parents (a v-structure) that a tree cannot represent at any budget:
dp_bayes <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
dependence = "tree", degree = 2,
bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_bayes, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : degree 2 Bayesian network (GreedyBayes) over 4 variables
#> marginals : 7 (4 one-way + 3 family joints, composed budget)
#> noise : Gaussian sd 2.941 per cell
#> selection : GreedyBayes, 0.25 of budget over 3 exponential-mechanism round(s); per-round eps 0.2999
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
It measures the d one-way marginals plus one (parents, node) family joint per
non-root node (2d - 1 marginals) and spends a select_frac slice on the d - 1
greedy picks; every slice composes into the same exact ((\epsilon, \delta)). The
network is forward-sampled ancestrally, so like the tree it needs no PGM inference.
degree is capped to d - 1, and — being a distinct structure search — it is an
alternative to structure_frac and to select = "adaptive" (setting either
alongside degree > 1 is an error). Raising degree makes each family a
(degree + 1)-way histogram, so the same cell-count caution as a high treewidth
applies.
Private-PGM inference (estimator = "pgm"). Every model above is PGM-free:
it uses each measured marginal locally — the tree takes the root’s one-way and
each edge’s raw noisy 2-way as (P(\text{child}\mid\text{parent})), the adaptive
junction tree each clique’s own array — so the other one-way marginals, and the
fact that overlapping noisy marginals disagree (each measured independently under
noise), are thrown away. dp_control(estimator = "pgm") adds the reconciliation
they omit, following McKenna et al.’s Private-PGM / MST: the whole measured set is
reconciled into the single graphical-model distribution that best fits all of it at
once (least squares), by belief propagation on a junction tree of the measured
cliques plus entropic mirror descent, and the model is sampled from that. Because a
tree (or the adaptive junction tree) has bounded treewidth, the inference is exact
and cheap.
dp_pgm <- dp_control(epsilon = 6, delta = 1e-6, mechanism = "gaussian",
dependence = "tree", estimator = "pgm",
bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_pgm, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : tree over 4 variables
#> marginals : 10 (measured under composed budget)
#> noise : Gaussian sd 3.044 per cell
#> estimator : Private-PGM reconciliation of the measured marginals
#> (belief propagation + mirror descent; post-processing, no extra budget)
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
Crucially, reconciliation is pure post-processing of the already-privatised
marginals, so it spends no extra budget — the ((\epsilon, \delta)) is
identical to the same release with the default estimator = "local"; only the
fitted model changes. It denoises (overlapping marginals are made mutually
consistent) and lets the otherwise-discarded one-way marginals constrain the model,
usually sharpening both the marginals and the conditionals at the same budget. It is
available for the flat dependence = "tree" release and for select = "adaptive",
and is an alternative to structure_frac, to degree > 1, and to anneal = TRUE.
Full AIM (select = "aim"). The adaptive selector above is deliberately
constrained: every new clique attaches to an existing one and covers a fresh
variable, so the measured cliques always form a junction tree and the model is
forward-sampled with no inference. The price is that it can never measure a marginal
between two variables that are both already in the model — so it cannot represent
a loop of pairwise dependence (a cycle (A!-!B!-!C!-!A)), which no
tree-shaped junction structure holds at any budget. select = "aim" removes that
constraint on selection: the exponential mechanism may pick any pair, loops
included. A loopy measured set has no forward sampler, so Private-PGM inference stops
being optional and becomes the estimator — the whole set (the one-ways plus the
selected pairs) is reconciled over a triangulated junction tree and sampled from
that.
dp_aim <- dp_control(epsilon = 8, delta = 1e-6, mechanism = "gaussian",
select = "aim", treewidth = 2,
bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 8 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.91 zCDP)
#> model : Full AIM graphical model (treewidth 2, Private-PGM) over 4 variables
#> marginals : 10 (4 one-way + 6 adaptively selected pairs, loopy, composed budget)
#> noise : Gaussian sd 2.707 per cell
#> selection : Full AIM (loopy marginals), 0.25 of budget over 6 exponential-mechanism round(s); per-round eps 0.2753
#> estimator : Private-PGM reconciliation over the triangulated model
#> (belief propagation + mirror descent; post-processing, no extra budget)
#> scoring : model-projection (candidates scored against the reconciled model; post-processing, default)
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
It runs a data-independent (\min!\big(\binom{d}{2},, \text{treewidth}\times(d-1)\big))
selection rounds, each rejecting any new pair whose triangulated clique would exceed
treewidth + 1 (so the model’s treewidth — and hence the inference cost — stays
bounded); at treewidth = 1 no loop can close, so it reduces to an
adaptively-selected tree. Selection and measurement compose to the same exact
((\epsilon, \delta)), the reconciliation is budget-neutral, and it is flat-table
only. Use it when you suspect cyclic or higher-order dependence that a single tree
cannot hold; keep treewidth small, since each clique is a bins^(treewidth+1)
histogram.
Adding anneal = TRUE swaps the fixed round count for the same data-adaptive
(\sigma)-halving schedule the adaptive selector offers: a baseline of
treewidth-capped new loopy pairs is selected first, then any surplus budget
re-measures the worst-fit already-measured pair (inverse-variance combined), with the
per-round budget doubling whenever a measurement fails its noise floor. The annealed
set is still triangulated and reconciled with Private-PGM, and the total spend is
exactly the same ((\epsilon, \delta)) over a variable number of rounds — the
privacy record reports the realised schedule.
dp_aim_anneal <- dp_control(epsilon = 8, delta = 1e-6, mechanism = "gaussian",
select = "aim", treewidth = 2, anneal = TRUE,
bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim_anneal, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 8 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.91 zCDP)
#> model : Full AIM graphical model (treewidth 2, Private-PGM) over 4 variables
#> marginals : 15 (4 one-way + 5 loopy + 6 refinement pair(s), composed budget)
#> noise : Gaussian sd 2.707 - 3.828 per cell (annealed range)
#> selection : annealed Full AIM (loopy marginals), 0.25 of budget; 11 round(s) = 5 loopy + 6 refinement; sigma halved 1x
#> estimator : Private-PGM reconciliation over the triangulated model
#> (belief propagation + mirror descent; post-processing, no extra budget)
#> scoring : model-projection (candidates scored against the reconciled model; post-processing, default)
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
Model-projection scoring (the default). By default Full AIM uses AIM’s actual
quality function: each candidate pair is scored against the current reconciled
model’s own marginal over that pair, so a loopy pair the model already explains —
through the marginals measured so far — no longer looks surprising, and the budget is
steered to the genuinely worst-fit interaction. That reference is read from the
already-privatised marginals (reconciled each round and projected onto the candidate,
which for a not-yet-measured pair crosses cliques of the junction tree), so it is pure
post-processing: the exponential mechanism’s sensitivity and the exact
((\epsilon, \delta)) are unchanged — only which marginals get selected changes. It
composes with anneal = TRUE and costs a reconciliation per selection round. The
cheaper one-way-product reference is still available as scoring = "independence".
# The default (scoring = "model") is shown above; this opts back to the cheaper
# one-way-product reference.
dp_aim_indep <- dp_control(epsilon = 8, delta = 1e-6, mechanism = "gaussian",
select = "aim", treewidth = 2, scoring = "independence",
bounds = list(age = c(18, 100), sbp = c(60, 240)))
synth(real, ~ id, privacy = dp_aim_indep, seed = 1)$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 8 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.91 zCDP)
#> model : Full AIM graphical model (treewidth 2, Private-PGM) over 4 variables
#> marginals : 10 (4 one-way + 6 adaptively selected pairs, loopy, composed budget)
#> noise : Gaussian sd 2.707 per cell
#> selection : Full AIM (loopy marginals), 0.25 of budget over 6 exponential-mechanism round(s); per-round eps 0.2753
#> estimator : Private-PGM reconciliation over the triangulated model
#> (belief propagation + mirror descent; post-processing, no extra budget)
#> scoring : independence (one-way product; model-projection disabled)
#> row cap : 1 per person
#> bin edges : dp (public bounds; no budget spent)
Contribution bound. At unit = "person" each person may contribute at most
max_rows_per_person rows (default 1, appropriate for one-row-per-person
tables). If a person legitimately has several rows, set this from public domain
knowledge; rows beyond the cap are subsampled away before any budget is spent, so
the guarantee holds by construction.
More budget, more fidelity. Larger epsilon means less noise. The synthetic
marginals converge to the real ones as (\epsilon) grows:
for (e in c(0.5, 2, 8)) {
d <- dp_control(epsilon = e, mechanism = "laplace",
dependence = "independent",
bounds = list(age = c(18, 100), sbp = c(60, 240)))
s <- as.data.frame(synth(real, ~ id, privacy = d, seed = 1))
cat(sprintf("epsilon = %-3g mean(sbp): real %.1f syn %.1f\n",
e, mean(real$sbp), mean(s$sbp)))
}
#> epsilon = 0.5 mean(sbp): real 129.7 syn 129.3
#> epsilon = 2 mean(sbp): real 129.7 syn 130.1
#> epsilon = 8 mean(sbp): real 129.7 syn 130.4
If the structure declares a nesting index (~ id / visit), Track B preserves
within-unit temporal structure with a first-order Markov model — the private
analogue of Track A’s initial-state + lag-1 transition model. Three things are
measured under one composed budget: a length histogram (how many rows a
person contributes), the initial-state marginals (the t = 1 row, one- and
two-way as before), and a transition matrix (P(v_t \mid v_{t-1})) for each
variable. Generation draws a length, an initial row, then steps each variable’s
transition matrix, so autocorrelation across visits is retained.
Because a person now contributes a whole trajectory, you must set
max_rows_per_person to the public maximum number of visits — this bounds each
person’s effect on the transition histograms (a length-(\le c) trajectory has
at most (c - 1) consecutive pairs). The budget still composes to exactly
((\epsilon), (\delta)); the accounting reports the split.
long <- do.call(rbind, lapply(1:400, function(i) {
nv <- sample(2:4, 1); s <- numeric(nv); s[1] <- rnorm(1, 130, 12)
for (t in seq_len(nv)[-1]) s[t] <- 0.85 * s[t - 1] + 0.15 * 130 + rnorm(1, 0, 5)
data.frame(id = i, visit = seq_len(nv), sbp = round(s),
sex = sample(c("F", "M"), 1))
}))
long$sex <- factor(long$sex)
dp_long <- dp_control(epsilon = 8, mechanism = "laplace",
max_rows_per_person = 4, bounds = list(sbp = c(60, 240)))
res_long <- synth(long, structure = ~ id / visit, privacy = dp_long, seed = 1)
res_long$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 8 , pure eps )-DP, person level
#> mechanism : laplace
#> model : tree over 2 variables (DP Markov: initial-state + transitions)
#> histograms: 6 (1 length + 3 initial + 2 transition, composed budget)
#> noise : Laplace scale 1.25 per cell
#> row cap : 4 per person
#> bin edges : dp (public bounds; no budget spent)
The synthetic visit index is regenerated as the within-person position, and a
lag-1 correlation like the real one survives the noise (DP attenuates it):
lag1 <- function(d) {
d <- d[order(d$id, d$visit), ]
prev <- ave(d$sbp, d$id, FUN = function(x) c(NA, head(x, -1)))
ok <- !is.na(prev); cor(prev[ok], d$sbp[ok])
}
syn_long <- as.data.frame(res_long)
c(real = lag1(long), synthetic = lag1(syn_long))
#> real synthetic
#> 0.8922741 0.3424675
By default every non-index column is treated as time-varying, so a genuinely
subject-invariant covariate (here sex) is stepped through its own transition
matrix and can drift between a synthetic person’s visits. If you know a column is
constant within a person — public schema knowledge, not something read from the
data — name it in baseline. It is then modelled once in the initial-state
model (keeping its distribution and its correlation with the first visit) and
broadcast unchanged to every row, so it is exactly constant within each synthetic
unit. Because a baseline column needs no transition histogram, declaring it also
removes that histogram from the release, so the remaining measurements are
sharper at the same ((\epsilon), (\delta)):
res_base <- synth(long, structure = ~ id / visit,
privacy = dp_control(epsilon = 8, mechanism = "laplace",
max_rows_per_person = 4,
bounds = list(sbp = c(60, 240)),
baseline = "sex"),
seed = 1)
# every synthetic person now has a single sex across their visits
syn_base <- as.data.frame(res_base)
max(tapply(syn_base$sex, syn_base$id, function(s) length(unique(s))))
#> [1] 1
res_base$privacy # one fewer transition histogram than the release above
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 8 , pure eps )-DP, person level
#> mechanism : laplace
#> model : tree over 2 variables (DP Markov: initial-state + transitions)
#> histograms: 5 (1 length + 3 initial + 1 transition, composed budget)
#> baseline : held constant within unit: sex
#> noise : Laplace scale 0.875 per cell
#> row cap : 4 per person
#> bin edges : dp (public bounds; no budget spent)
The default transition model is first-order and per-variable: each variable’s next
value depends only on its own previous value. Two knobs deepen it. transition_order = k conditions on a variable’s own last k values (momentum), and transition_cross = m additionally conditions on the lag-1 values of its m most strongly associated
companions — so cross-variable coupling is re-measured at every step, not just carried
at t = 1 by the initial-state tree. Companions are picked from the pairwise marginals
the tree already measures, so cross-conditioning spends no extra budget (a
transition tuple still lands in exactly one cell); a higher order actually lowers the
transition sensitivity to cap - order, since a person then contributes fewer, deeper
tuples. The order must therefore be at most max_rows_per_person - 1, and
transition_cross > 0 needs the tree model.
res_ho <- synth(long, structure = ~ id / visit,
privacy = dp_control(epsilon = 8, mechanism = "laplace",
dependence = "tree", max_rows_per_person = 4,
bounds = list(sbp = c(60, 240)),
baseline = "sex",
transition_order = 2, transition_cross = 1),
seed = 1)
res_ho$privacy # note: transitions order 2 + 1 cross-parent, same (eps, delta)
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 8 , pure eps )-DP, person level
#> mechanism : laplace
#> model : tree over 2 variables (DP Markov: initial-state + transitions)
#> histograms: 5 (1 length + 3 initial + 1 transition, composed budget)
#> baseline : held constant within unit: sex
#> transitions: order 2 + 1 cross-parent(s) (sensitivity cap - order)
#> cross-parents: sbp ~ sex
#> noise : Laplace scale 0.75 per cell
#> row cap : 4 per person
#> bin edges : dp (public bounds; no budget spent)
Early rows (before position order + 1, which have no full history yet) are generated
by marginalising the very same measured tensor, which is post-processing and costs
nothing. The trade-off is cell sparsity: a wider conditioning grid spreads the same
noisy counts more thinly, so reach for depth only when the extra structure is worth it.
synth_linked() accepts a dp_control() too, giving a DP release across a whole
key hierarchy at once. The privacy unit is the root entity (e.g. a patient):
adding or removing one individual — its root row and every descendant row that
cascades from it — changes the release within the budget. Contribution is bounded
hierarchically, so max_rows_per_person here means the maximum children kept per
parent for each child table: a single integer for every child table, or a named
list keyed by table name. The root cap is always 1.
set.seed(1)
patients <- data.frame(
id = 1:300,
age = round(rnorm(300, 60, 11)),
sex = factor(sample(c("F", "M"), 300, TRUE)))
adm <- do.call(rbind, lapply(patients$id, function(pid) {
n <- rpois(1, 1.4); if (n == 0) return(NULL)
data.frame(id = pid, admission_id = seq_len(n), los = 1L + rpois(n, 4))
}))
dp_link <- dp_control(
epsilon = 4, mechanism = "laplace",
max_rows_per_person = list(admissions = 6), # <= 6 admissions per patient
domain = "public",
bounds = list(age = c(18, 100), los = c(0, 60)))
res_link <- synth_linked(
tables = list(patients = patients, admissions = adm),
structures = list(patients = ~ id, admissions = ~ id / admission_id),
keys = list(patients = "id", admissions = c("id", "admission_id")),
privacy = dp_link, seed = 1)
res_link$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 4 , pure eps )-DP, person level
#> mechanism : laplace
#> model : tree over 3 variables (linked: 2 tables)
#> histograms: 5 (per-table variable marginals + child count models, composed budget)
#> noise : Laplace scale 2.5 per cell
#> - patients root, cap 1
#> - admissions child of patients, <=6/parent, path cap 6 (2 rows dropped)
#> bin edges : public (public bounds; no budget spent)
Each table’s own variable marginals and a children-per-parent count histogram
are measured under one exactly-composed budget; the summed L1 (Laplace) and summed
squared L2 (Gaussian zCDP) sensitivities add over every histogram, with each
table’s contribution scaled by its per-entity path cap (the product of branching
caps from the root). Synthetic children copy their synthetic parent’s surrogate
key, so referential integrity holds by construction — check_linkage() confirms
it:
check_linkage(res_link)
By default each child table’s variables are modelled by their own within-table marginals, so the synthetic child links to a synthetic parent but is statistically independent of it.
dp_control(cross_table = TRUE) closes that gap. For each child table with a
modellable immediate parent, the engine also measures parent-by-child joint
marginals — counted at the child grain (one observation per child row, at the
parent’s value carried down the foreign key). Their person-sensitivity is the
child’s path cap, exactly like a child one-way marginal, so they fold into the
same composed budget. The parent’s variables then enter the child’s Chow-Liu
structure as fixed context nodes, and at generation the synthetic parent’s
already-drawn value conditions the child draw:
dp_cross <- dp_control(
epsilon = 4, mechanism = "laplace", cross_table = TRUE,
max_rows_per_person = list(admissions = 6),
domain = "public",
bounds = list(age = c(18, 100), los = c(0, 60)))
res_cross <- synth_linked(
tables = list(patients = patients, admissions = adm),
structures = list(patients = ~ id, admissions = ~ id / admission_id),
keys = list(patients = "id", admissions = c("id", "admission_id")),
privacy = dp_cross, seed = 1)
res_cross$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 4 , pure eps )-DP, person level
#> mechanism : laplace
#> model : tree over 3 variables (linked: 2 tables)
#> histograms: 7 (per-table variable marginals + child count models, composed budget)
#> noise : Laplace scale 5.5 per cell
#> - patients root, cap 1
#> - admissions child of patients, <=6/parent, path cap 6, cond. on parent (2 rows dropped)
#> bin edges : public (public bounds; no budget spent)
The accounting now lists admissions as cond. on parent and the extra joints
raise the histogram count and the per-cell noise — the price of the added
structure, still inside the same \((\epsilon, \delta)\). Each child variable’s
single strongest predictor may be a parent variable or another child variable
(under dependence = "independent" it is the best parent variable only), and a
three-level hierarchy chains it: a grandchild conditions on its already-conditioned
parent. Only the immediate parent conditions a child; deeper ancestors reach it
through the parent’s synthesised values, as in Track A.
Only the immediate parent conditions a child; deeper ancestors reach it through the parent’s synthesised values, as in Track A. Constraints are refused, as elsewhere under DP.
When a child table’s rows are a time series within each parent unit — repeated
visits ordered by the child’s own key index — dp_control(longitudinal = TRUE)
(or a vector of child-table names) models them as a within-unit DP Markov
trajectory instead of exchangeable records. The children-per-parent count model
doubles as the trajectory-length model; an initial-state model is measured over
each unit’s first (earliest) child row; and a first-order transition matrix
\(P(v_t \mid v_{t-1})\) is measured per variable over consecutive within-unit rows:
set.seed(1)
visits <- do.call(rbind, lapply(patients$id, function(pid) {
k <- 2L + rpois(1, 1.2)
st <- character(k); st[1] <- sample(c("stable", "worse"), 1)
for (i in 2:k) st[i] <- if (runif(1) < 0.85) st[i - 1]
else setdiff(c("stable", "worse"), st[i - 1])
data.frame(id = pid, visit_num = seq_len(k),
status = factor(st, levels = c("stable", "worse")))
}))
dp_longi <- dp_control(
epsilon = 6, delta = 1e-6, mechanism = "gaussian",
max_rows_per_person = c(visits = 6), longitudinal = "visits",
domain = "public", bounds = list(age = c(18, 100)))
res_longi <- synth_linked(
tables = list(patients = patients, visits = visits),
structures = list(patients = ~ id, visits = ~ id / visit_num),
keys = list(patients = "id", visits = c("id", "visit_num")),
privacy = dp_longi, seed = 1)
res_longi$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : tree over 3 variables (linked: 2 tables)
#> histograms: 6 (per-table variable marginals + child count models, composed budget)
#> noise : Gaussian sd 5.273 per cell
#> - patients root, cap 1
#> - visits child of patients, <=6/parent, path cap 6, DP Markov over rows (5 rows dropped)
#> bin edges : public (public bounds; no budget spent)
The accounting lists visits as DP Markov over rows. The transition histograms
have person-sensitivity path_cap[parent] * (branching_cap - 1) while the
initial-state marginals sit at the parent path cap — heterogeneous, but folded into
the same exact \((\epsilon, \delta)\). Over-cap units are prefix-truncated in temporal
order so their consecutive pairs stay intact. This keeps a patient’s status
trending across visits, the autocorrelation an exchangeable child would drop.
Setting cross_table = TRUE together with a longitudinal model on the same
child combines the two: the child’s initial-state model is cross-conditioned on
the synthetic parent — the first row of each unit draws from a parent-conditioned
Chow-Liu tree — and the within-unit transition chain then carries that parent
dependence across the trajectory. The parent shapes where a trajectory starts;
the transitions stay parent-free, so the only extra cost is the nC * nP
parent-by-child initial-state joints, at the same first-row sensitivity as the
initial marginals:
dp_both <- dp_control(
epsilon = 6, delta = 1e-6, mechanism = "gaussian",
max_rows_per_person = c(visits = 6), longitudinal = "visits",
cross_table = TRUE, domain = "public", bounds = list(age = c(18, 100)))
res_both <- synth_linked(
tables = list(patients = patients, visits = visits),
structures = list(patients = ~ id, visits = ~ id / visit_num),
keys = list(patients = "id", visits = c("id", "visit_num")),
privacy = dp_both, seed = 1)
res_both$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : tree over 3 variables (linked: 2 tables)
#> histograms: 8 (per-table variable marginals + child count models, composed budget)
#> noise : Gaussian sd 5.446 per cell
#> - patients root, cap 1
#> - visits child of patients, <=6/parent, path cap 6, DP Markov over rows (initial state cond. on parent) (5 rows dropped)
#> bin edges : public (public bounds; no budget spent)
The accounting now lists visits as DP Markov over rows (initial state cond. on parent). A patient’s whole status trajectory can then inherit a dependence on
their baseline attributes, not just autocorrelate with itself.
The two within-unit transition controls of the flat DP Markov engine also apply,
per table, to a longitudinally-modelled linked child. dp_control(baseline = c(...)) names subject-invariant columns of a child (matched against that child’s
own columns) and holds them exactly constant within a unit — dropping their
transition histograms — while dp_control(transition_order = k, transition_cross = m) deepens each time-varying column’s transition to its own last k values plus
the lag-1 values of its m most associated companions. A higher order lowers the
child’s transition sensitivity to path_cap[parent] * (branching_cap - order), so
the order must be at most one less than the child’s branching cap:
dp_order2 <- dp_control(
epsilon = 6, delta = 1e-6, mechanism = "gaussian",
max_rows_per_person = c(visits = 6), longitudinal = "visits",
transition_order = 2, domain = "public", bounds = list(age = c(18, 100)))
res_order2 <- synth_linked(
tables = list(patients = patients, visits = visits),
structures = list(patients = ~ id, visits = ~ id / visit_num),
keys = list(patients = "id", visits = c("id", "visit_num")),
privacy = dp_order2, seed = 1)
res_order2$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : tree over 3 variables (linked: 2 tables)
#> histograms: 6 (per-table variable marginals + child count models, composed budget)
#> noise : Gaussian sd 4.411 per cell
#> - patients root, cap 1
#> - visits child of patients, <=6/parent, path cap 6, DP Markov over rows (5 rows dropped)
#> transitions: order 2 + 0 cross-parent(s) (sensitivity cap - order)
#> bin edges : public (public bounds; no budget spent)
The visits line now reports transitions: order 2; with a baseline column the
accounting would add a baseline held: ... line and charge one fewer transition
histogram. Both settings compose with the combined cross-conditioned initial state
above.
The combined cross-conditioned model above cross-conditions only a child’s
initial state; the parent’s influence then rides the own-lag chain and fades.
dp_control(transition_parent = p) instead re-injects the synthetic parent’s
subject-invariant attributes into the child’s transition at every step — each
time-varying column conditions its next value on the p immediate-parent
attributes most strongly associated with it, so parent → child dependence stays
anchored across the trajectory. The parents are chosen budget-neutrally from
the parent-by-child joints the cross-conditioned initial state already measures, so
it adds no histogram and no sensitivity — which is exactly why it requires
cross_table = TRUE:
dp_tp <- dp_control(
epsilon = 6, delta = 1e-6, mechanism = "gaussian",
max_rows_per_person = c(visits = 6), longitudinal = "visits",
cross_table = TRUE, transition_parent = 1,
domain = "public", bounds = list(age = c(18, 100)))
res_tp <- synth_linked(
tables = list(patients = patients, visits = visits),
structures = list(patients = ~ id, visits = ~ id / visit_num),
keys = list(patients = "id", visits = c("id", "visit_num")),
privacy = dp_tp, seed = 1)
res_tp$privacy
#> <dp_accounting> Track B differential privacy
#> guarantee : (epsilon = 6 , delta = 1e-06 )-DP, person level
#> mechanism : gaussian (rho = 0.54 zCDP)
#> model : tree over 3 variables (linked: 2 tables)
#> histograms: 8 (per-table variable marginals + child count models, composed budget)
#> noise : Gaussian sd 5.446 per cell
#> - patients root, cap 1
#> - visits child of patients, <=6/parent, path cap 6, DP Markov over rows (initial state cond. on parent) (5 rows dropped)
#> transitions: order 1 + 0 cross-parent(s) + 1 parent-attr(s) (sensitivity cap - order)
#> parent-attrs: status ~ age
#> bin edges : public (public bounds; no budget spent)
The visits line reports initial state cond. on parent plus a `transitions: …
line and aparent-attrs: status ~ …` mapping — at the same
\((\epsilon, \delta)\) as the cross-conditioned-initial-state model without it.cross_table = TRUE conditions each child on its immediate parent only.rule()) are refused under DP, because data-dependent rejection
sampling would leak information the budget does not account for.Within those limits the guarantee is exact and reported. For anything labelled
differentially private, keep the accounting record with the release. The
discretisation adds no unaccounted leakage in either rigorous mode — supply
public bounds (domain = "public") to spend nothing on the domain, or let the
default domain = "dp" estimate the edges privately and charge them to the
budget.