Why a dual-code strategy?

In Queensland hospital admission data, a chronic condition is not always coded consistently. A patient with COPD admitted for an exacerbation may be coded under the acute admission code J44.1 without the supplementary U-code U83.2 being recorded alongside it. A strategy that searches U-codes alone will systematically undercount comorbidity burden — sometimes substantially.

plumage() solves this by flagging a condition as present if any of up to three independent code types match: an ICD-10-AM supplementary U-code, a matching acute principal or secondary ICD-10-AM admission code, or (optionally) an AR-DRG code. These are combined with OR logic — one match is enough.

The name fits: just as experienced birders read a bird’s plumage as a proxy for its underlying physiological condition, plumage() reads a patient record’s clinical coding as a proxy for chronic condition burden.


Conditions detected

plumage() detects 29 chronic conditions across 9 body-system categories.

† Cystic fibrosis is intentionally counted in both Metabolic/Endocrine and Respiratory because it has clinically relevant manifestations in both systems.
Category Condition Column name
Metabolic/Endocrine Obesity obesity
Metabolic/Endocrine Cystic fibrosis† cystic_fibrosis
Mental Health Dementia dementia
Mental Health Schizophrenia schizophrenia
Mental Health Depression depression
Mental Health Intellectual/developmental disability intellectual_dev
Neurological Parkinson’s disease parkinsons
Neurological Multiple sclerosis multiple_sclerosis
Neurological Epilepsy epilepsy
Neurological Cerebral palsy cerebral_palsy
Neurological Paralysis paralysis
Cardiovascular Ischaemic heart disease ihd
Cardiovascular Heart failure heart_failure
Cardiovascular Hypertension hypertension
Respiratory Emphysema emphysema
Respiratory COPD copd
Respiratory Asthma/chronic bronchitis asthma
Respiratory Bronchiectasis bronchiectasis
Respiratory Chronic respiratory failure respiratory_failure
Respiratory Cystic fibrosis† cystic_fibrosis
Gastrointestinal Crohn’s disease crohns
Gastrointestinal Ulcerative colitis ulcerative_colitis
Gastrointestinal Liver failure liver_failure
Musculoskeletal Rheumatoid arthritis rheumatoid_arthritis
Musculoskeletal Osteoarthritis osteoarthritis
Musculoskeletal SLE lupus
Musculoskeletal Osteoporosis osteoporosis
Renal Chronic kidney disease kidney_disease
Congenital Spina bifida spina_bifida
Congenital Down syndrome downs

Basic usage

hospital_data <- data.frame(
  patient_id = 1:5,
  icd_codes  = c(
    "K29.70",                        # gastritis only — no chronic comorbidities
    "U78.1, U83.2, U82.3",           # obesity + COPD + hypertension (all U-codes)
    "J44.1, U79.3",                  # COPD via acute ICD + depression via U-code
    "J43.2, J47",                    # emphysema + bronchiectasis (acute ICD, no U-codes)
    "E84.0, U80.3"                   # cystic fibrosis (acute ICD) + epilepsy (U-code)
  )
)

results <- plumage(hospital_data, "icd_codes")

# View key columns
results[, c("patient_id", "copd", "emphysema", "bronchiectasis",
            "cystic_fibrosis", "total_conditions", "conditions_category")]
#>   patient_id copd emphysema bronchiectasis cystic_fibrosis total_conditions
#> 1          1    0         0              0               0                0
#> 2          2    1         0              0               0                3
#> 3          3    1         0              0               0                2
#> 4          4    0         1              1               0                2
#> 5          5    0         0              0               1                2
#>   conditions_category
#> 1                   0
#> 2                  3+
#> 3                   2
#> 4                   2
#> 5                   2

Notice that row 3 has copd = 1 detected purely from the acute code J44.1 (no U83.2 present), and row 4 has both emphysema and bronchiectasis detected from acute codes alone. This is exactly the dual-code advantage.


The conditions_category summary

Every run of plumage() produces a conditions_category ordered factor — a coarse summary of comorbidity burden useful for stratified analyses and tables.

table(results$conditions_category)
#> 
#>  0  1  2 3+ 
#>  1  0  3  1

The ordering (0 < 1 < 2 < 3+) is preserved in gtsummary::tbl_summary() and ggplot2 without any extra setup.


No-decimal code format

Some Queensland datasets store ICD codes without decimal points (e.g. U832 instead of U83.2). Set decimal = FALSE to match this format.

df_nodot <- data.frame(
  icd = c("U832 U823", "J441 J431"),
  stringsAsFactors = FALSE
)
plumage(df_nodot, "icd", decimal = FALSE)[, c("copd", "hypertension", "emphysema")]
#>   copd hypertension emphysema
#> 1    1            1         0
#> 2    1            0         1

Including AR-DRG codes

When DRG codes are available in a separate column, include_drg = TRUE adds them as a third detection pathway — particularly useful for COPD, asthma, and bronchiectasis where DRGs are well-specified.

df_drg <- data.frame(
  patient_id = 1:3,
  icd_codes  = c("K29.70",  "J44.1",  "K29.70"),  # row 3: no respiratory ICD
  drg_codes  = c("G07B",    "E65A",   "E65A")      # row 3: COPD DRG only
)

# Without DRG: row 3 missed entirely
plumage(df_drg, "icd_codes", include_drg = FALSE)[, c("patient_id", "copd")]
#>   patient_id copd
#> 1          1    0
#> 2          2    1
#> 3          3    0

# With DRG: row 3 caught via E65A
plumage(df_drg, "icd_codes", include_drg = TRUE,
        drg_column = "drg_codes")[, c("patient_id", "copd")]
#>   patient_id copd
#> 1          1    0
#> 2          2    1
#> 3          3    1

Prefixing output columns

When combining plumage() output with other flag columns, use prefix to avoid name collisions.

res_prefixed <- plumage(hospital_data, "icd_codes", prefix = "chr_")
names(res_prefixed)[grepl("^chr_", names(res_prefixed))] |> head(8)
#> [1] "chr_obesity"            "chr_cystic_fibrosis"    "chr_dementia"          
#> [4] "chr_schizophrenia"      "chr_depression"         "chr_intellectual_dev"  
#> [7] "chr_parkinsons"         "chr_multiple_sclerosis"

Lean output with drop_eggs

For downstream modelling where you only need summary counts, drop_eggs = TRUE removes the 29 individual binary columns and retains only the 11 summary columns — a substantial reduction in width for large datasets.

res_lean <- plumage(hospital_data, "icd_codes", drop_eggs = TRUE)
names(res_lean)
#>  [1] "patient_id"                        "icd_codes"                        
#>  [3] "total_conditions"                  "total_metabolic_conditions"       
#>  [5] "total_mental_health_conditions"    "total_neurological_conditions"    
#>  [7] "total_cardiovascular_conditions"   "total_respiratory_conditions"     
#>  [9] "total_gastrointestinal_conditions" "total_musculoskeletal_conditions" 
#> [11] "total_renal_conditions"            "total_congenital_conditions"      
#> [13] "conditions_category"

After plumage(): what comes next

The output of plumage() integrates naturally with the rest of the mudnester pipeline:

# Typical hospitalisation workflow
df_hosp <- clean_the_nest(hosp_raw, data_type = "hospital", ...)
df_hosp <- plumage(df_hosp, icd_column = "icd_code")
df_hosp <- preening(df_hosp, age_col = "age", scheme = "geriatric_fine")

# Stratify comorbidity burden by age group before aggregation
roost(df_hosp, date_col = "admission_date", time_unit = "month",
      group_cols = c("age_group", "conditions_category"))

Before sharing or archiving the enriched dataset, pass it through molting() (see vignette("molting")). The conditions_category column is retained by default — it matches the age\d+cat preservation pattern — but check that total_conditions and the individual binary columns are appropriately handled for your sharing context.


Notes for SCPHU practice

  • ICD-10-AM vs ICD-10: the acute ICD stems and U-codes are specific to the Australian modification (ICD-10-AM). Do not apply this function to datasets coded under the international ICD-10 without reviewing the code mappings — the U-code block (U78–U88) does not exist in ICD-10.
  • Multiple admissions per patient: plumage() operates row-wise. If your dataset has multiple rows per patient (one per admission), a patient will be flagged for a condition in any row where the relevant code appears. Aggregate across admissions first (e.g. group_by(patient_id) |> summarise(copd = max(copd))) if you want one row per patient.
  • Code completeness: neither U-codes nor DRG codes are 100% consistently recorded. The dual/triple-code strategy mitigates but does not eliminate undercounting. For high-stakes analyses, validate against a clinical reference standard on a subset.

See vignette("mudnester-getting-started") for the full pipeline context.