--- title: "Worked Example Using Multiple Tables" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Worked Example Using Multiple Tables} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", # pharmaversesdtm is suggested, only evaluate if it is installed eval = requireNamespace("pharmaversesdtm", quietly = TRUE) ) ``` # Introduction This R Markdown document illustrates example usage of the RESIDE package with multiple related tables, using the Demographics (DM) and Adverse Events (AE) SDTM domains from the [pharmaversesdtm](https://pharmaverse.github.io/pharmaversesdtm/) package. The tables are linked by a subject identifier (`USUBJID`), DM has one row per subject and AE has one row per adverse event. # Setup Load the RESIDE package and set a seed for reproducibility and store the folder directory for export / import. ```{r load-library, message = FALSE} # Load the Library library(RESIDE) # Load dplyr for data manipulation library(dplyr) # Set the seed set.seed(1234) # Store the folder path used for import / export folder_path <- tempdir() ``` # Summarise original data Store the tables in a named list, the names are used to identify the tables in the marginal distributions. ```{r original-data} # Store the tables in a named list dfs <- list( dm = pharmaversesdtm::dm, ae = pharmaversesdtm::ae ) # Number of rows in each table sapply(dfs, nrow) # Number of subjects in each table sapply(dfs, function(df) length(unique(df$USUBJID))) # Treatment arms of the subjects table(dfs$dm$ARM) # Severity of the adverse events table(dfs$ae$AESEV) ``` # Fit a logistic regression model on the original data Join the adverse events to the demographics of the subject and fit a logistic regression model for the odds of an adverse event being moderate or severe, by age, sex and treatment arm. The same function is used for the original and synthesised data. ```{r original-model} # Join the adverse events to the demographics of each subject prepare_ae_data <- function(dm, ae) { ae |> select(USUBJID, AESEV) |> inner_join(select(dm, USUBJID, AGE, SEX, ARM), by = "USUBJID") |> # Remove screen failures and adverse events without a severity filter(ARM != "Screen Failure", AESEV != "") |> mutate( MOD_SEV = AESEV %in% c("MODERATE", "SEVERE"), ARM = relevel(factor(ARM), "Placebo") ) } ae_original <- prepare_ae_data(dfs$dm, dfs$ae) # Proportion of moderate or severe adverse events by treatment arm prop.table(table(ae_original$ARM, ae_original$MOD_SEV), 1) # Fit a logistic regression model glm.original <- glm( MOD_SEV ~ AGE + SEX + ARM, data = ae_original, family = binomial ) # Output the coefficients of the model summary(glm.original)$coefficients ``` **NB** Adverse events from the same subject are not independent, a mixed effects model would be more appropriate, however a logistic regression model is used here for simplicity. # Get the marginal distributions Use the `get_marginal_distributions()` function to get the marginal distributions of the tables, specifying the column that identifies each subject using the `subject_identifier` parameter. ```{r get-marginals} # Get the Marginal Distributions of both tables marginals <- get_marginal_distributions( dfs, subject_identifier = "USUBJID" ) # Summarise the Marginal Distributions summary(marginals) ``` Columns that are present in every table (other than the subject identifier), are identified as common columns, in this case the study identifier `STUDYID`. # Export the marginal distributions Export the marginal distributions using the `export_marginal_distributions()` function, using the `force` parameter to override any existing files. ```{r export-marginals} # Export the Marginal Distributions export_marginal_distributions(marginals, folder_path = folder_path, force = TRUE) ``` # Reimport marginal distributions Import the exported marginal distributions using the `import_marginal_distributions()` function ```{r import-marginals} # Import the Marginal Distributions imported_marginals <- import_marginal_distributions(folder_path = folder_path) ``` # Synthesise data from marginal distributions (without correlations) Synthesise data from the imported marginals using the `synthesise_data` function, for multiple tables a named list of data frames is returned. ```{r synthesis-data} # Synthesise the tables from the imported Marginal Distributions (without correlations) sim_dfs <- synthesise_data(imported_marginals) # Number of rows in each table sapply(sim_dfs, nrow) # Number of subjects in each table sapply(sim_dfs, function(df) length(unique(df$USUBJID))) ``` The number of rows and subjects of each table are maintained, synthesised subjects are identified by a number, which links the subjects between the tables. # Fit a logistic regression model on the synthesised data Fit the same logistic regression model as earlier except this time on the synthesised data. ```{r sim-data-model} ae_sim <- prepare_ae_data(sim_dfs$dm, sim_dfs$ae) # Proportion of moderate or severe adverse events by treatment arm prop.table(table(ae_sim$ARM, ae_sim$MOD_SEV), 1) # Fit a logistic regression model on the synthesised data glm.sim <- glm( MOD_SEV ~ AGE + SEX + ARM, data = ae_sim, family = binomial ) # Output the coefficients of the model summary(glm.sim)$coefficients ``` Without correlations the variables are synthesised independently, so there is no relationship between the treatment arm and the severity of the adverse events. # Synthesise data with correlations Synthesise data from the imported marginals with assumed correlations, using the `correlations` parameter to specify a list of correlations created with the `correlation()` function. Categorical variables are correlated using a single category, specified with `factor_name.x` or `factor_name.y`, and the tables of the variables can be specified with `df_name.x` and `df_name.y`. ```{r synthesis-data-cor} # Synthesise the tables specifying assumed correlations sim_dfs_cor <- synthesise_data( imported_marginals, correlations = list( # Subjects on placebo are more likely to have mild adverse events correlation( "ARM", "AESEV", 0.2, df_name.x = "dm", df_name.y = "ae", factor_name.x = "Placebo", factor_name.y = "MILD" ), # Male subjects are younger correlation("AGE", "SEX", -0.2, factor_name.y = "M") ) ) ``` Correlated variables are synthesised together, one row per subject, and joined to each table by subject. Therefore correlations between tables are between subjects, and a correlated variable takes a single value for each subject within a table, in this case every adverse event of a subject has the same severity. # Fit a logistic regression model on the synthesised data with correlations Using the synthesised data (with correlations) fit the same logistic regression model as earlier. ```{r sim-data-model-cor} ae_sim_cor <- prepare_ae_data(sim_dfs_cor$dm, sim_dfs_cor$ae) # Proportion of moderate or severe adverse events by treatment arm prop.table(table(ae_sim_cor$ARM, ae_sim_cor$MOD_SEV), 1) # Fit a logistic regression model on the synthesised data (with correlations) glm.sim.cor <- glm( MOD_SEV ~ AGE + SEX + ARM, data = ae_sim_cor, family = binomial ) # Output the coefficients of the model summary(glm.sim.cor)$coefficients ``` With the assumed correlation, adverse events of subjects on placebo are less likely to be moderate or severe, allowing the analysis to be tested before access to the original data is granted. **NB** The correlation specified is that of the underlying multivariate normal distribution (Copula), the correlation observed in the synthesised data will differ, particularly for categorical variables. It is not possible to entirely maintain all the marginal distributions when specifying correlations.