--- title: "Getting started with trialdiff" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Getting started with trialdiff} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` ## Why trialdiff? Clinical trial datasets are regenerated many times during a study. Between two data cuts, observations are added, removed or corrected, derivations change, and metadata evolves. Generic tools such as `diffdf` and `waldo` can tell you *that* two datasets differ. `trialdiff` answers the clinical-programming question that comes next: > What changed, what does the change represent, and which downstream analyses or > outputs might be affected? `trialdiff` is organised into five layers, each usable on its own: 1. **Compare** - [compare_cut()] 2. **Classify** - [classify_changes()] 3. **Trace** - [define_lineage()], [trace_dependencies()] 4. **Assess** - [assess_impact()] 5. **Report** - [report_diff()] All five layers are deterministic and rule-based. Nothing is inferred by an opaque model, and no statistical significance is ever claimed. ## The example data The package ships synthetic `ADSL` and `ADLB` data cuts. No proprietary data is used. ```{r} library(trialdiff) str(adsl_cut1, max.level = 1) ``` ## 1. Compare two data cuts ```{r} diff <- compare_cut( old = adsl_cut1, new = adsl_cut2, by = "USUBJID", dataset = "ADSL" ) diff ``` The result is a `tdiff` object with `added`, `removed`, `modified` and `schema` tables. For example, the treatment-assignment change: ```{r} diff$modified[, c("USUBJID", "variable", "old_value", "new_value", "change")] ``` ## 2. Classify the changes ```{r} classified <- classify_changes(diff) table(classified$register$category_label) ``` Every classification carries a plain-language explanation: ```{r} classified$register$reason[classified$register$category == "treatment_assignment_change"] ``` ## 3. Define lineage Lineage is explicit metadata. Nodes are datasets (`ADSL`), variables (`ADSL.TRT01P`), analyses (`MMRM`) or outputs (`Table_14_2_1`). ```{r} lineage <- define_lineage( lineage_edge("ADSL.TRT01P", "ADLB.TRT01P", relationship = "groups_by"), lineage_edge("ADLB.AVAL", "MMRM", relationship = "models"), lineage_edge("MMRM", "Table_14_2_1", relationship = "reports") ) trace_dependencies(lineage, from = "ADSL.TRT01P") ``` ## 4. Assess impact ```{r} impact <- assess_impact(classified, adsl_adlb_lineage) impact$impacts[, c("node", "node_type", "level", "requires_rerun")] ``` Impact is graded as `definitely_affected`, `potentially_affected` or `unlikely`. Analyses and outputs are flagged as requiring review/rerun; the package never claims statistical impact. ## 5. Report ```{r} report <- report_diff(diff, impact = impact, output = "list") report$data$review_items ``` Use `output = "report.html"` to write a self-contained HTML report, or `as_json()` for machine-readable output suitable for automated QC pipelines. ## Next steps * `vignette("change-classification")` for the rule system. * `vignette("lineage-and-impact")` for the lineage model. * `vignette("ecosystem")` for how `trialdiff` complements existing tools.