--- title: "plumage(): Comorbidity Detection from ICD-10-AM Coding" subtitle: "A dual/triple-code strategy for Queensland hospital admission data" author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service" date: "`r Sys.Date()`" output: html_document: toc: true toc_depth: 3 toc_float: true theme: flatly pdf_document: toc: true toc_depth: 3 number_sections: true latex_engine: xelatex vignette: > %\VignetteIndexEntry{plumage(): Comorbidity Detection from ICD-10-AM Coding} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE) library(mudnester) ``` ## Why a dual-code strategy? In Queensland hospital admission data, a chronic condition is not always coded consistently. A patient with COPD admitted for an exacerbation may be coded under the acute admission code `J44.1` without the supplementary U-code `U83.2` being recorded alongside it. A strategy that searches U-codes alone will systematically undercount comorbidity burden — sometimes substantially. `plumage()` solves this by flagging a condition as present if **any** of up to three independent code types match: an ICD-10-AM supplementary U-code, a matching acute principal or secondary ICD-10-AM admission code, or (optionally) an AR-DRG code. These are combined with OR logic — one match is enough. The name fits: just as experienced birders read a bird's plumage as a proxy for its underlying physiological condition, `plumage()` reads a patient record's clinical coding as a proxy for chronic condition burden. --- ## Conditions detected `plumage()` detects **29 chronic conditions** across **9 body-system categories**. ```{r conditions-table, echo=FALSE} knitr::kable( data.frame( Category = c( "Metabolic/Endocrine", "Metabolic/Endocrine", "Mental Health", "Mental Health", "Mental Health", "Mental Health", "Neurological", "Neurological", "Neurological", "Neurological", "Neurological", "Cardiovascular", "Cardiovascular", "Cardiovascular", "Respiratory", "Respiratory", "Respiratory", "Respiratory", "Respiratory", "Respiratory", "Gastrointestinal", "Gastrointestinal", "Gastrointestinal", "Musculoskeletal", "Musculoskeletal", "Musculoskeletal", "Musculoskeletal", "Renal", "Congenital", "Congenital" ), Condition = c( "Obesity", "Cystic fibrosis†", "Dementia", "Schizophrenia", "Depression", "Intellectual/developmental disability", "Parkinson's disease", "Multiple sclerosis", "Epilepsy", "Cerebral palsy", "Paralysis", "Ischaemic heart disease", "Heart failure", "Hypertension", "Emphysema", "COPD", "Asthma/chronic bronchitis", "Bronchiectasis", "Chronic respiratory failure", "Cystic fibrosis†", "Crohn's disease", "Ulcerative colitis", "Liver failure", "Rheumatoid arthritis", "Osteoarthritis", "SLE", "Osteoporosis", "Chronic kidney disease", "Spina bifida", "Down syndrome" ), `Column name` = c( "obesity", "cystic_fibrosis", "dementia", "schizophrenia", "depression", "intellectual_dev", "parkinsons", "multiple_sclerosis", "epilepsy", "cerebral_palsy", "paralysis", "ihd", "heart_failure", "hypertension", "emphysema", "copd", "asthma", "bronchiectasis", "respiratory_failure", "cystic_fibrosis", "crohns", "ulcerative_colitis", "liver_failure", "rheumatoid_arthritis", "osteoarthritis", "lupus", "osteoporosis", "kidney_disease", "spina_bifida", "downs" ), check.names = FALSE ), caption = "† Cystic fibrosis is intentionally counted in both Metabolic/Endocrine and Respiratory because it has clinically relevant manifestations in both systems." ) ``` --- ## Basic usage ```{r basic} hospital_data <- data.frame( patient_id = 1:5, icd_codes = c( "K29.70", # gastritis only — no chronic comorbidities "U78.1, U83.2, U82.3", # obesity + COPD + hypertension (all U-codes) "J44.1, U79.3", # COPD via acute ICD + depression via U-code "J43.2, J47", # emphysema + bronchiectasis (acute ICD, no U-codes) "E84.0, U80.3" # cystic fibrosis (acute ICD) + epilepsy (U-code) ) ) results <- plumage(hospital_data, "icd_codes") # View key columns results[, c("patient_id", "copd", "emphysema", "bronchiectasis", "cystic_fibrosis", "total_conditions", "conditions_category")] ``` Notice that row 3 has `copd = 1` detected purely from the acute code `J44.1` (no `U83.2` present), and row 4 has both emphysema and bronchiectasis detected from acute codes alone. This is exactly the dual-code advantage. --- ## The `conditions_category` summary Every run of `plumage()` produces a `conditions_category` ordered factor — a coarse summary of comorbidity burden useful for stratified analyses and tables. ```{r category} table(results$conditions_category) ``` The ordering (`0 < 1 < 2 < 3+`) is preserved in `gtsummary::tbl_summary()` and `ggplot2` without any extra setup. --- ## No-decimal code format Some Queensland datasets store ICD codes without decimal points (e.g. `U832` instead of `U83.2`). Set `decimal = FALSE` to match this format. ```{r decimal-false} df_nodot <- data.frame( icd = c("U832 U823", "J441 J431"), stringsAsFactors = FALSE ) plumage(df_nodot, "icd", decimal = FALSE)[, c("copd", "hypertension", "emphysema")] ``` --- ## Including AR-DRG codes When DRG codes are available in a separate column, `include_drg = TRUE` adds them as a third detection pathway — particularly useful for COPD, asthma, and bronchiectasis where DRGs are well-specified. ```{r drg} df_drg <- data.frame( patient_id = 1:3, icd_codes = c("K29.70", "J44.1", "K29.70"), # row 3: no respiratory ICD drg_codes = c("G07B", "E65A", "E65A") # row 3: COPD DRG only ) # Without DRG: row 3 missed entirely plumage(df_drg, "icd_codes", include_drg = FALSE)[, c("patient_id", "copd")] # With DRG: row 3 caught via E65A plumage(df_drg, "icd_codes", include_drg = TRUE, drg_column = "drg_codes")[, c("patient_id", "copd")] ``` --- ## Prefixing output columns When combining `plumage()` output with other flag columns, use `prefix` to avoid name collisions. ```{r prefix} res_prefixed <- plumage(hospital_data, "icd_codes", prefix = "chr_") names(res_prefixed)[grepl("^chr_", names(res_prefixed))] |> head(8) ``` --- ## Lean output with `drop_eggs` For downstream modelling where you only need summary counts, `drop_eggs = TRUE` removes the 29 individual binary columns and retains only the 11 summary columns — a substantial reduction in width for large datasets. ```{r drop-eggs} res_lean <- plumage(hospital_data, "icd_codes", drop_eggs = TRUE) names(res_lean) ``` --- ## After `plumage()`: what comes next The output of `plumage()` integrates naturally with the rest of the `mudnester` pipeline: ```r # Typical hospitalisation workflow df_hosp <- clean_the_nest(hosp_raw, data_type = "hospital", ...) df_hosp <- plumage(df_hosp, icd_column = "icd_code") df_hosp <- preening(df_hosp, age_col = "age", scheme = "geriatric_fine") # Stratify comorbidity burden by age group before aggregation roost(df_hosp, date_col = "admission_date", time_unit = "month", group_cols = c("age_group", "conditions_category")) ``` Before sharing or archiving the enriched dataset, pass it through `molting()` (see `vignette("molting")`). The `conditions_category` column is retained by default — it matches the `age\d+cat` preservation pattern — but check that `total_conditions` and the individual binary columns are appropriately handled for your sharing context. --- ## Notes for SCPHU practice - **ICD-10-AM vs ICD-10**: the acute ICD stems and U-codes are specific to the Australian modification (ICD-10-AM). Do not apply this function to datasets coded under the international ICD-10 without reviewing the code mappings — the U-code block (`U78`–`U88`) does not exist in ICD-10. - **Multiple admissions per patient**: `plumage()` operates row-wise. If your dataset has multiple rows per patient (one per admission), a patient will be flagged for a condition in any row where the relevant code appears. Aggregate across admissions first (e.g. `group_by(patient_id) |> summarise(copd = max(copd))`) if you want one row per patient. - **Code completeness**: neither U-codes nor DRG codes are 100% consistently recorded. The dual/triple-code strategy mitigates but does not eliminate undercounting. For high-stakes analyses, validate against a clinical reference standard on a subset. See `vignette("mudnester-getting-started")` for the full pipeline context.