Real clinical data rarely lives in one table. Here we build a small, fully synthetic cardiac example (never real patient data): patients, their cardiac admissions, and the labs taken during each admission.
patients (id)
└─ admissions (id / admission_id)
└─ labs (id / admission_id / lab_number)
n_pat <- 50
patients <- data.frame(
id = seq_len(n_pat),
sex = sample(c("F", "M"), n_pat, replace = TRUE, prob = c(0.45, 0.55)),
age = round(rnorm(n_pat, 64, 11)),
smoker = sample(c(FALSE, TRUE), n_pat, replace = TRUE, prob = c(0.7, 0.3))
)
adm <- do.call(rbind, lapply(patients$id, function(pid) {
k <- 1 + rpois(1, 0.6)
data.frame(id = pid, admission_id = seq_len(k),
los_days = 1 + rpois(k, 3))
}))
labs <- do.call(rbind, lapply(seq_len(nrow(adm)), function(i) {
k <- 1 + rpois(1, 2)
data.frame(id = adm$id[i], admission_id = adm$admission_id[i],
lab_number = seq_len(k),
analyte = sample(c("troponin", "LDL", "creatinine"), k, replace = TRUE),
value = round(abs(rnorm(k, 2, 1.5)), 2))
}))
tables <- list(patients = patients, admissions = adm, labs = labs)
vapply(tables, nrow, integer(1))
#> patients admissions labs
#> 50 80 244
synth_linked()synth_linked() takes a structures formula per table and the keys, and
reads the hierarchy from them: each table’s key is its parent’s key plus one more
column. Root tables are synthesised with the single-table engine; child tables
are generated from their synthetic parent, with foreign keys copied over so
referential integrity holds by construction.
structures <- list(
patients = ~ id,
admissions = ~ id / admission_id,
labs = ~ id / admission_id / lab_number
)
keys <- list(
patients = "id",
admissions = c("id", "admission_id"),
labs = c("id", "admission_id", "lab_number")
)
res <- synth_linked(tables, structures = structures, keys = keys, seed = 1)
res
#> <synth_linked_result>
#> track : A (high-utility; NOT differentially private)
#> datasets (m) : 1
#> tables : 3
#> - patients 50 rows (input 50) (root)
#> - admissions 82 rows (input 80) child of patients
#> - labs 234 rows (input 244) child of admissions
#> method : cart
#>
#> Get the tables with as.list(x)
#> Verify linkage with check_linkage(x)
The number of children per parent is drawn from a learned count model (including parents with no children), and child variables are conditioned on the parent’s synthesised attributes.
check_linkage() confirms key uniqueness and the absence of orphan child rows.
syn <- as.list(res)
check_linkage(res)
# every synthetic lab points at a real (synthetic) admission — no orphans
all(paste(syn$labs$id, syn$labs$admission_id) %in%
paste(syn$admissions$id, syn$admissions$admission_id))
#> [1] TRUE
diagnose() and disclosure_risk() accept the whole result and report one
block per table.
diagnose(tables, res)
#> <flexsynth_diagnostics_list> 3 tables
#>
#> == patients ==
#> <flexsynth_diagnostics>
#> rows : real 50 synthetic 50
#> variables : 4
#>
#> Univariate fit (smaller = closer):
#> variable type metric distance
#> id numeric ks 0.00
#> sex categorical tvd 0.08
#> age numeric ks 0.10
#> smoker logical tvd 0.10
#> mean distance: 0.0700 worst: age (0.1000)
#>
#> Correlation structure (2 numeric vars):
#> Frobenius diff: 0.6521 mean |diff|: 0.4611 max |diff|: 0.4611
#>
#> Categorical association (Cramer's V, 2 vars):
#> mean |diff|: 0.0000 max |diff|: 0.0000
#>
#> Propensity utility (pMSE, logistic; descriptive, in-sample):
#> pMSE: 0.00470 expected: 0.00500 ratio: 0.94 (1 = indistinguishable)
#>
#> == admissions ==
#> <flexsynth_diagnostics>
#> rows : real 80 synthetic 82
#> variables : 3
#>
#> Univariate fit (smaller = closer):
#> variable type metric distance
#> id numeric ks 0.0564
#> admission_id numeric ks 0.0152
#> los_days numeric ks 0.0680
#> mean distance: 0.0465 worst: los_days (0.0680)
#>
#> Correlation structure (3 numeric vars):
#> Frobenius diff: 0.3669 mean |diff|: 0.1243 max |diff|: 0.2419
#>
#> Propensity utility (pMSE, logistic; descriptive, in-sample):
#> pMSE: 0.00023 expected: 0.00229 ratio: 0.10 (1 = indistinguishable)
#>
#> == labs ==
#> <flexsynth_diagnostics>
#> rows : real 244 synthetic 234
#> variables : 5
#>
#> Univariate fit (smaller = closer):
#> variable type metric distance
#> id numeric ks 0.0711
#> admission_id numeric ks 0.0450
#> lab_number numeric ks 0.0345
#> analyte categorical tvd 0.0879
#> value numeric ks 0.0894
#> mean distance: 0.0656 worst: value (0.0894)
#>
#> Correlation structure (4 numeric vars):
#> Frobenius diff: 0.1646 mean |diff|: 0.0359 max |diff|: 0.1019
#>
#> Propensity utility (pMSE, logistic; descriptive, in-sample):
#> pMSE: 0.00348 expected: 0.00160 ratio: 2.17 (1 = indistinguishable)
#>
disclosure_risk() is most meaningful per table with the genuinely identifying
columns named (excluding surrogate keys such as id, which are regenerated):
disclosure_risk(patients, syn$patients,
quasi = c("age", "sex", "smoker"), seed = 1)
#> <flexsynth_disclosure>
#> rows : real 50 synthetic 50
#> quasi-identifiers: age, sex, smoker
#>
#> Replicated uniques (identity risk):
#> real sample-uniques : 33
#> reproduced in syn : 19 (57.58% of uniques, 38.00% of real rows)
#> syn rows copying a real row: 37 (74.00% of syn)
#>
#> Distance to closest record (Gower, 0 = exact copy):
#> syn->real : median 0.0000 5th pct 0.0000 exact copies 74.00%
#> real->real : median 0.0081 (baseline)
#> median syn distance is smaller than the real-neighbour baseline
#> descriptive only: inspect lower-tail distances and exact copies; this is not a safety guarantee
#>
#> Membership inference: not run (supply `holdout` of non-training records).
#>
#> Attribute disclosure: not run (supply `target` = a sensitive column).
Passing the whole objects (disclosure_risk(tables, res)) also works and
defaults to all shared columns per table — quick, but it will include surrogate
keys, so prefer naming quasi for a real release. Constraints (rule())
currently apply to single-table synth(); enforce cross-table logic on the
parent table before linking.
The example data here is synthetic. Track A output is high-utility but carries no formal privacy guarantee — always read the disclosure diagnostics before sharing.