--- title: "homing(): Relinking De-identified Data" subtitle: "The authorised path back from anonymisation" author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service" date: "`r Sys.Date()`" output: html_document: toc: true toc_depth: 3 toc_float: true theme: flatly pdf_document: toc: true toc_depth: 3 number_sections: true latex_engine: xelatex vignette: > %\VignetteIndexEntry{homing(): Relinking De-identified Data} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE) library(mudnester) ``` ## When to use `homing()` `homing()` is the complement to `molting()`. It is used when: 1. A de-identified dataset has been shared or stored, and 2. An authorised person subsequently needs to recover the original identifiers — for example, to contact individuals for follow-up, to correct a data error, or for a notifiable disease regulatory obligation. The name is precise: homing pigeons navigate back to their loft regardless of where they were released, using an internal compass that only they carry. `homing()` navigates a de-identified dataset back to its identifiers using the lookup table — and only those who hold the lookup table can make that journey. --- ## The basic relink workflow ```{r workflow} # Construct sample data patient_data <- data.frame( patient_name = c("John Doe", "Jane Smith", "Alice Brown"), dob = as.Date(c("1980-01-01", "1975-05-15", "1992-11-30")), mrn = c("12345", "67890", "11111"), diagnosis = c("Condition A", "Condition B", "Condition C"), severity = c("mild", "moderate", "severe") ) # Step 1: de-identify (typically done at data collection / storage time) result <- suppressMessages(molting(patient_data)) # Step 2: share or archive result$deidentified # store result$lookup securely, separately # Step 3: relink when authorised relinked <- homing( deidentified_data = result$deidentified, lookup_table = result$lookup ) head(relinked) ``` The original identifiers (`patient_name`, `dob`, `mrn`) are joined back in via the `row_hash` column. --- ## Controlling the hash column name If `molting()` was called with a custom `hash_col_name`, pass the same name to `homing()`. ```{r custom-hash} result_custom <- suppressMessages( molting(patient_data, hash_col_name = "person_hash") ) relinked_custom <- homing( result_custom$deidentified, result_custom$lookup, hash_col_name = "person_hash" ) "patient_name" %in% names(relinked_custom) ``` --- ## Removing the hash after relinking If you want a clean re-identified dataset without the hash column: ```{r no-hash} relinked_clean <- homing( result$deidentified, result$lookup, keep_hash = FALSE ) names(relinked_clean) # no row_hash column ``` --- ## Partial matches and unmatched records If the lookup table is incomplete (e.g. some records were excluded from the lookup for a legitimate reason, or the wrong lookup was supplied), `homing()` warns you about unmatched rows and returns them with `NA` in the identifier columns rather than silently dropping them. ```{r partial} # Simulate a truncated lookup — only the first two rows partial_lookup <- result$lookup[1:2, ] relinked_partial <- homing( result$deidentified, partial_lookup ) # Third row has NA identifiers relinked_partial[, c("row_hash","patient_name","diagnosis")] ``` Always check the summary message for the matched count. A significantly lower matched count than expected usually means the wrong lookup was supplied. --- ## Security and governance checklist Before using `homing()` in a production workflow, ensure: - [ ] The relink is authorised by your organisation's privacy and governance framework - [ ] The lookup table was stored separately from the de-identified data, with access logging - [ ] The relinking event is documented in your data management plan - [ ] The re-identified dataset is treated as identifiable data from the moment `homing()` returns - [ ] The re-identified dataset is not stored in the same location as the de-identified data unless access controls are equivalent In Queensland Health, re-identification for notifiable disease follow-up typically falls under Public Health Act 2005 obligations and does not require separate ethics approval, but document the basis for re-identification in your outbreak log. --- ## What comes next After re-identification, the data is again fully identifiable. If you need to re-anonymise for a secondary analysis, run `molting()` again. See `vignette("molting")` for options. If the purpose of the relink was to add clinical follow-up data, the updated dataset can be re-cleaned with `clean_the_nest()` before further analysis.