--- title: "dataprep: descriptive statistics and diagnostic plots" output: rmarkdown::html_vignette vignette: > %\documentclass{article} %\VignetteIndexEntry{dataprep: descriptive statistics and diagnostic plots} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", fig.align = "center", fig.width = 6, fig.height = 5.5, out.width = "80%", fig.retina = 2 ) ``` ```{r} library(dataprep) library(ggplot2) ``` ## Overview `dataprep` provides two families of plotting helpers, both built on top of the C++ descriptive backends: * **`descplot()`** — descriptive statistics (`n`, `na`, `mean`, `sd`, `median`, `trimmed`, `min`, `max`, `IQR`) computed in C++ and displayed as line or bar charts. * **`percplot()`** — top and bottom percentile summaries, useful for detecting heavy tails and percentile-based outlier cutoffs. The geometry depends on the column names: lines when they are numeric, grouped bars otherwise. Both share the same interface style: a data frame, a numeric range, and optional grouping. Use `data1` (7,640 rows × 7 columns) for quick demos, and `data` (7,640 × 65) for full-size examples. Under the hood, `descplot()` calls `descdata()` (which calls `desc_stats_cpp()`), and `percplot()` calls `percdata()` (which calls `quantile()` from base R on each column). Both return a `ggplot` object, so all usual `ggplot2` layers apply. ## Descriptive statistics ### Line plot, numeric variable names When variable names are essentially numeric (e.g. particle diameters such as 3.16, 3.55, ...), `descplot()` draws a line plot with a log-scaled x axis. ```{r, fig.height = 5} descplot(data1, cols = 3:7) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ### Selected statistics Pass a subset of statistics by index or by name to focus the plot. ```{r, fig.height = 4} descplot(data1, cols = 3:7, stats = c("na", "min", "max", "IQR")) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` The `stats` argument accepts both forms — numeric indices (`1:9`) and character names (`"na"`, `"min"`, `"max"`, `"IQR"`) — and can mix them. ### Bar chart, character variable names When variable names are character (e.g. aerosol mode names `Nucleation`, `Aitken`, `Accumulation`), `descplot()` falls back to a bar chart. ```{r, fig.height = 5} descplot(data1, cols = 3:7) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ### Control facet layout ```{r, fig.height = 3} descplot(data1, cols = 3:7, stats = c("min", "max", "IQR")) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ### Full-size data ```{r, fig.height = 4.5} descplot(data, cols = 5:65) ``` ### The underlying table If you only need the numbers (not the plot), call `descdata()` directly. It returns a data frame with one row per variable and one column per statistic. ```{r, fig.height = 4} descdata(data1, cols = 3:7, stats = c(2, 3, 4, 7:9)) ``` ## Percentile plots ### Full percentile range `percplot()` computes the extreme percentiles (0 to 0.5 and 99.5 to 100 by default) and draws them against the variable axis. This is the visual companion to `condextr()` and `percoutl()`. ```{r, fig.height = 6, fig.width = 7} percplot(data1, cols = 3:7, group = 2) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ### Top percentiles only ```{r, fig.height = 3, fig.width = 7} percplot(data1, cols = 3:7, group = 2, part = "top") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ### Bottom percentiles only ```{r, fig.height = 3, fig.width = 7} percplot(data1, cols = 3:7, group = 2, part = "bottom") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ### Percentile curves of the raw data When the column names are numeric, `percplot()` draws curves instead of bars. The 61 size-bin columns of `data` (`cols = 5:65`) have numeric names, so the default `num_xaxis = "auto"` selects a log-scaled x axis: ```{r, fig.height = 4, fig.width = 7} percplot(data, cols = 5:65, group = 4) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ### Numeric axis control For numeric variable names, the x axis can be forced to linear scale with `num_xaxis = "numeric"`. ```{r, fig.height = 4, fig.width = 7} percplot(data, cols = 5:65, group = 4, num_xaxis = "numeric") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` The `num_xaxis` argument controls how the x axis is treated when column names are numeric: * `"auto"` (default) — uses a log scale if the column names are evenly spaced on a log scale with a range ratio of at least 1000; otherwise keeps them as factor levels. * `"log"` / `"numeric"` — force the corresponding scale. * `"character"` / `"factor"` / `"keep"` / `FALSE` — keep the column names as factor levels. ### The underlying table `percdata()` returns the same table that `percplot()` draws. ```{r} percdata(data1, cols = 3:7, group = 2, part = "top") ``` ## Combining with `ggplot2` Both `descplot()` and `percplot()` return `ggplot` objects, so all usual `ggplot2` layers apply. ```{r, fig.height = 4, fig.width = 7} percplot(data1, cols = 3:7, group = 2) + ggplot2::theme_bw(base_size = 11) + ggplot2::labs(title = "Percentile plots by month", x = "Variable", y = "Value") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1)) ``` ## A diagnostic workflow A typical diagnostic workflow combines `data_report()`, `na_diagnose()`, and the two plot families: ```{r, eval = FALSE} # 1. Overview of the whole table data_report(data, cols = 5:65) # 2. Per-column NA run statistics na_diagnose(data, cols = 5:65) # 3. Descriptive statistics of the raw data descplot(data, cols = 5:65) # 4. Percentile curves of the raw data percplot(data, cols = 5:65, group = 4) ``` After running `dataprep()` you can compare the raw and cleaned versions in the same plot by stacking them with a `g` column: ```{r, eval = FALSE} cleaned <- dataprep(data, cols = 5:65, group = 4) percplot( rbind( transform(data[names(cleaned)], g = "original"), transform(cleaned, g = "preprocessed") ), cols = 5:ncol(cleaned), group = ncol(cleaned) + 1 ) ``` ## Where to go next * **Design philosophy** — `vignette("dataprep-philosophy")`. Why the cleaning pipeline has the shape it does. * **Cleaning pipeline walkthrough** — `vignette("dataprep-cleaning")`. Step-by-step execution of the four cleaning steps on `data`. * **Performance and cross-engine consistency** — `vignette("dataprep-performance")`. Benchmark tables and 8-engine consistency checks. * **Upgrading from 0.1.5 to 0.1.8** — `vignette("dataprep-migration")`. Behaviour changes and the migration checklist. * **Fast reshaping** — `vignette("dataprep-melt-dcast")`. ## Session info ```{r} sessionInfo() ```