--- title: "molting(): Hash-Based De-identification" subtitle: "Stripping identifiers while preserving relinkability" author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service" date: "`r Sys.Date()`" output: html_document: toc: true toc_depth: 3 toc_float: true theme: flatly pdf_document: toc: true toc_depth: 3 number_sections: true latex_engine: xelatex vignette: > %\VignetteIndexEntry{molting(): Hash-Based De-identification} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE) library(mudnester) ``` ## Why de-identify? Sharing a linelist externally — with collaborators, for CRAN examples, in a research supplement — requires that directly identifying information is removed. `molting()` automates this using a cryptographic hash: each row gets a unique fingerprint derived from its identifiers, the identifiers are stripped, and the fingerprint is stored in a lookup table that only authorised personnel hold. The name comes from the natural process of moulting: a bird sheds its distinctive, identifiable plumage and temporarily becomes more uniform. The old plumage is not destroyed — it re-grows from the same follicles. The lookup table is those follicles. --- ## What gets removed `molting()` uses regular-expression pattern matching to detect PII columns. The default patterns cover: - Names: `name`, `surname`, `firstname`, `lastname` - Dates: `dob`, `birth` - Identifiers: `mrn`, `urn`, `medicare`, `patient_id`, `subject_id`, `_id$` - Contact: `address`, `street`, `phone`, `email` Age **category** variables (`age2cat`, `age5cat`, `age10cat`, etc.) are automatically preserved — they are not directly identifying. ```{r detect-demo} patient_data <- data.frame( patient_name = c("John Doe", "Jane Smith"), dob = as.Date(c("1980-01-01", "1975-05-15")), mrn = c("12345", "67890"), age5cat = factor(c("18-64", "18-64")), # preserved automatically diagnosis = c("Condition A", "Condition B"), lab_value = c(120, 95) ) result <- suppressMessages(molting(patient_data)) names(result$deidentified) # hash + retained columns names(result$lookup) # hash + removed columns ``` --- ## The output list `molting()` returns a named list with two elements when `return_lookup = TRUE` (the default): - `$deidentified` — the de-identified data frame, with `row_hash` as the first column - `$lookup` — the lookup table, with `row_hash` plus all removed identifier columns ```{r output-structure} str(result, max.level = 1) head(result$lookup) ``` Store `$lookup` securely and separately from `$deidentified`. Consider encrypting the lookup file before archiving. In a Queensland Health context, the lookup table should remain within the Health Service network. --- ## Hash algorithm selection The default is SHA-256, which provides strong collision resistance for typical surveillance dataset sizes (tens of thousands of rows). For very large datasets where speed matters more than collision resistance, SHA-1 or MD5 are faster but should not be used where linkage integrity is critical. ```{r hash-algos, eval=FALSE} # SHA-256 (default, recommended) result_256 <- molting(patient_data, hash_method = "sha256") # MD5 (shorter hash, faster, lower collision resistance) result_md5 <- molting(patient_data, hash_method = "md5") # Blake3 (fast and cryptographically strong — good for large datasets) result_b3 <- molting(patient_data, hash_method = "blake3") ``` --- ## Controlling which columns are hashed By default, `molting()` hashes all detected PII columns. Supply `id_cols` to override this — useful when you want a shorter, more stable hash based only on a true unique identifier. ```{r id-cols} # Hash only on MRN and DOB — more stable if name variations exist result_ids <- suppressMessages( molting(patient_data, id_cols = c("mrn", "dob")) ) result_ids$lookup ``` --- ## Adding columns to the removal list Use `additional_pii_cols` for dataset-specific identifiers that don't match the default patterns. ```{r additional} patient_data2 <- patient_data patient_data2$study_code <- c("SC-001","SC-002") result2 <- suppressMessages( molting(patient_data2, additional_pii_cols = "study_code") ) names(result2$deidentified) ``` --- ## Irreversible de-identification If you genuinely do not need to relink (e.g. producing a public-use file), set `return_lookup = FALSE`. This is irreversible — there is no way to recover the original identifiers. ```{r no-lookup} deidentified_only <- suppressMessages( molting(patient_data, return_lookup = FALSE) ) class(deidentified_only) # a data frame, not a list ``` --- ## Hash collisions If two rows produce the same hash (extremely rare with SHA-256 for realistic dataset sizes but possible with very short hashes like CRC32), `molting()` warns you. If a collision is detected, switch to a stronger algorithm or add more columns to `id_cols`. --- ## What comes next Use [homing()] to relink the de-identified data when authorised (see `vignette("homing")`). For aggregated outputs — monthly counts from `roost()` — de-identification may not be necessary at all if the counts are not small enough to be re-identifying. The ABS cell-suppression threshold of 5 is a useful rule of thumb.