Package {tidymatrix}


Title: Tidyverse-Style Operations on Matrices with Row and Column Metadata
Version: 0.1.0
Description: Provides a unified data structure for matrices with associated row and column metadata, enabling 'tidyverse'-style data manipulation. Following the approach of 'tidygraph', users can activate rows, columns, or the matrix itself and operate on it with familiar 'dplyr' verbs, while the matrix and both metadata tables are kept consistent.
License: MIT + file LICENSE
URL: https://raivokolde.github.io/tidymatrix/, https://github.com/raivokolde/tidymatrix
BugReports: https://github.com/raivokolde/tidymatrix/issues
Depends: R (≥ 4.1.0)
Encoding: UTF-8
LazyData: true
Imports: dplyr, rlang, stats, tibble, utils
Suggests: testthat (≥ 3.0.0), ggplot2, knitr, pheatmap, rmarkdown, Rtsne, umap
Config/testthat/edition: 3
VignetteBuilder: knitr
Config/roxygen2/version: 8.1.0
NeedsCompilation: no
Packaged: 2026-09-28 13:06:52 UTC; raivokolde
Author: Raivo Kolde [aut, cre]
Maintainer: Raivo Kolde <rkolde@gmail.com>
Repository: CRAN
Date/Publication: 2026-10-08 11:00:02 UTC

Activate different components of a tidymatrix

Description

Switch the active context of a tidymatrix to operate on rows, columns, or the matrix itself. This determines which component will be affected by subsequent dplyr operations.

Usage

activate(.data, what)

Arguments

.data

A tidymatrix object

what

Which component to activate. One of "rows", "columns", or "matrix"

Value

A tidymatrix object with the specified component activated

Examples

mat <- matrix(rnorm(12), nrow = 4, ncol = 3)
row_data <- data.frame(id = 1:4, group = c("A", "A", "B", "B"))
col_data <- data.frame(id = 1:3, type = c("x", "y", "z"))
tm <- tidymatrix(mat, row_data, col_data)

# Activate rows to filter/mutate row metadata
tm |> activate(rows)

# Activate columns to work with column metadata
tm |> activate(columns)

# Activate matrix to work with the matrix directly
tm |> activate(matrix)

Get the active component of a tidymatrix

Description

Get the active component of a tidymatrix

Usage

active(.data)

Arguments

.data

A tidymatrix object

Value

A character string indicating the active component

Examples

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

active(tm)
active(activate(tm, rows))

Add statistics to metadata

Description

Compute statistics for the active dimension and add them as columns to the corresponding metadata. Requires rows or columns to be active.

Usage

add_stats(.data, ..., .fns = NULL, .names = NULL)

Arguments

.data

A tidymatrix object with rows or columns active

...

Functions to compute statistics (unquoted names like mean, var, sd)

.fns

Alternative way to specify functions as a list

.names

Names for the new columns. If NULL, uses function names.

Value

A tidymatrix object with statistics added to metadata

Examples

mat <- matrix(rnorm(20), nrow = 4, ncol = 5)
tm <- tidymatrix(mat)

# Add row statistics
tm <- tm |>
  activate(rows) |>
  add_stats(mean, var, sd)

# Now row_data has columns: mean, var, sd

# Add column statistics
tm <- tm |>
  activate(columns) |>
  add_stats(median, min, max)

# Custom names
tm <- tm |>
  activate(rows) |>
  add_stats(mean, var, .names = c("row_mean", "row_var"))

Analysis Management for tidymatrix

Description

Functions to store, retrieve, and manage analysis results attached to tidymatrix objects.


Arrange the active component of a tidymatrix

Description

Reorder rows based on the active metadata component. When rows are active, arranges by row_data and reorders matrix rows accordingly. When columns are active, arranges by col_data and reorders matrix columns accordingly. Cannot arrange when matrix is active.

Usage

## S3 method for class 'tidymatrix'
arrange(.data, ...)

Arguments

.data

A tidymatrix object

...

Variables to arrange by

Details

Reordering invalidates any stored analyses (see get_analysis), since objects such as hclust and prcomp refer to the original row/column positions.

Value

A tidymatrix object with reordered data

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

# Order items by trait; the matrix columns are reordered to match
tm |>
  activate(columns) |>
  arrange(trait, position)

Convert tidymatrix metadata to data.frame

Description

Converts the active metadata component (row_data or col_data) to a data.frame. This method only works when rows or columns are active, not when the matrix is active.

Usage

## S3 method for class 'tidymatrix'
as.data.frame(x, ...)

Arguments

x

A tidymatrix object

...

Additional arguments (currently unused)

Value

A data.frame containing the active metadata

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
row_data <- data.frame(id = 1:10, group = rep(c("A", "B"), each = 5))
tm <- tidymatrix(mat, row_data)

# Convert row metadata to data.frame
df <- tm |>
  activate(rows) |>
  compute_prcomp(center = TRUE, scale. = TRUE) |>
  as.data.frame()

Convert tidymatrix metadata to tibble

Description

Converts the active metadata component (row_data or col_data) to a tibble. This method only works when rows or columns are active, not when the matrix is active.

Usage

## S3 method for class 'tidymatrix'
as_tibble(x, ...)

Arguments

x

A tidymatrix object

...

Additional arguments (currently unused)

Value

A tibble containing the active metadata

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
row_data <- data.frame(id = 1:10, group = rep(c("A", "B"), each = 5))
tm <- tidymatrix(mat, row_data)

# Convert row metadata to a tibble, e.g. for plotting
pcs <- tm |>
  activate(rows) |>
  compute_prcomp(center = TRUE, scale. = TRUE) |>
  as_tibble()

if (requireNamespace("ggplot2", quietly = TRUE)) {
  library(ggplot2)
  ggplot(pcs, aes(x = row_pca_PC1, y = row_pca_PC2, color = group)) +
    geom_point()
}

Simulated Big Five personality survey

Description

A simulated personality questionnaire in which 400 respondents answer 30 Likert items, six for each of the Big Five personality traits. The data come in the three pieces that make up a tidymatrix, and can be combined with tidymatrix(big5_responses, big5_respondents, big5_items).

Usage

big5_responses

big5_respondents

big5_items

Format

big5_responses

An integer matrix with 400 rows (respondents) and 30 columns (items). Values range from 1 (strongly disagree) to 5 (strongly agree). Row and column names are respondent and item IDs.

big5_respondents

A data frame with one row per respondent:

respondent_id

Respondent ID, e.g. "R001".

age

Age in years (18–79).

gender

"Female", "Male" or "Non-binary".

education

Highest completed education, a factor with levels Basic < Secondary < Bachelor < Master < PhD.

occupation

Occupational group, e.g. "Student", "Professional", "Retired".

country

Two-letter country code.

life_satisfaction

Self-rated life satisfaction, 0–10.

completion_min

Time taken to complete the survey, minutes.

big5_items

A data frame with one row per item:

item_id

Item ID: trait letter and number, e.g. "E1".

trait

Big Five trait measured by the item.

reversed

TRUE for reverse-keyed items.

item_text

Statement shown to the respondent.

position

Position of the item in the questionnaire; traits are interleaved.

Details

The data are simulated, but built to behave like real survey data:

The script that generates the data is in the data-raw folder of the package source.

See Also

big5_countries for country-level information to join to the respondents.

Examples

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)
tm


Country information for the Big Five survey

Description

A small country-level table to join to big5_respondents. It is deliberately incomplete: Germany ("DE") has respondents but no row here, and Norway ("NO") has a row but no respondents, which makes it useful for illustrating the different kinds of joins.

Usage

big5_countries

Format

A data frame with 6 rows and 5 columns:

country

Two-letter country code.

country_name

Country name.

region

"Baltic" or "Nordic".

language

Main official language.

population_m

Approximate population, millions.

Examples

big5_countries

Center the active dimension of an object

Description

Generic for centering. See center.tidymatrix for the tidymatrix method.

Usage

center(x, ...)

Arguments

x

An object to center

...

Arguments passed to methods

Value

An object of the same class as x, centered

Examples

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

# Center each item (column) on its mean
tm |>
  activate(columns) |>
  center()

Center rows or columns

Description

Center the active dimension by subtracting the mean. Requires rows or columns to be active.

Usage

## S3 method for class 'tidymatrix'
center(x, ...)

Arguments

x

A tidymatrix object with rows or columns active

...

Not used

Value

A tidymatrix object with centered matrix

Examples

mat <- matrix(rnorm(20, mean = 10), nrow = 4, ncol = 5)
tm <- tidymatrix(mat)

# Center rows (mean of each row = 0)
tm_centered <- tm |>
  activate(rows) |>
  center()

# Center columns (mean of each column = 0)
tm_centered <- tm |>
  activate(columns) |>
  center()

Check analysis validity

Description

Check if stored analyses are still valid given the current data dimensions.

Usage

check_analyses(x)

Arguments

x

A tidymatrix object

Value

Invisibly returns x. Reports the status of each analysis as a message.

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(matrix(rnorm(100), 10, 10))
tm <- tm |>
  activate(columns) |>
  compute_prcomp(name = "pca")

check_analyses(tm)
# Analysis 'pca': VALID (10 rows x 10 columns)

tm <- tm |> activate(rows) |> filter(row_number() <= 5)
check_analyses(tm)
# Analysis 'pca': dimensions changed (was 10 rows, now 5 rows)

Clip matrix values

Description

Cap matrix values at specified minimum and/or maximum. Requires matrix to be active.

Usage

clip_values(.data, min = NULL, max = NULL)

Arguments

.data

A tidymatrix object with matrix active

min

Minimum value (values below are set to this)

max

Maximum value (values above are set to this)

Value

A tidymatrix object with clipped matrix values

Examples

mat <- matrix(rnorm(20), nrow = 4, ncol = 5)
tm <- tidymatrix(mat)

# Clip to [-2, 2] range
tm_clipped <- tm |>
  activate(matrix) |>
  clip_values(min = -2, max = 2)

# Only set floor
tm_floor <- tm |>
  activate(matrix) |>
  clip_values(min = 0)

Apply a function across matrix dimensions with metadata access

Description

Applies a user-defined function to each row (when rows are active) or each column (when columns are active), providing both the values and the opposite dimension's metadata. This enables complex statistical modeling where you need access to all annotations.

Usage

compute_across(
  .data,
  fn,
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE,
  ...
)

Arguments

.data

A tidymatrix object with rows or columns active (not matrix)

fn

A function with signature ⁠function(values, metadata, ...)⁠ where:

  • values: Numeric vector of row/column values

  • metadata: Complete col_data (rows active) or row_data (columns active)

  • ...: Additional arguments from compute_across() Must return a named list or named vector

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

...

Additional arguments passed to fn

Value

If add_to_data = FALSE: data.frame/tibble with one row per matrix row/column, containing identifiers and computed statistics. If add_to_data = TRUE: modified tidymatrix with results added to metadata.

Examples

# T-test example
mat <- matrix(rnorm(100, mean = 10), nrow = 10, ncol = 10)
col_data <- data.frame(
  sample = paste0("S", 1:10),
  condition = rep(c("Control", "Treatment"), each = 5)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# Run t-test on each row
results <- tm |>
  activate(rows) |>
  compute_across(
    fn = function(vals, meta) {
      test <- t.test(vals ~ meta$condition)
      list(
        p.value = test$p.value,
        log2fc = log2(mean(vals[meta$condition == "Treatment"]) /
                      mean(vals[meta$condition == "Control"]))
      )
    }
  )

# Linear model with multiple predictors
col_data2 <- data.frame(
  sample = paste0("S", 1:10),
  condition = rep(c("Control", "Treatment"), each = 5),
  batch = factor(rep(1:2, 5)),
  age = rnorm(10, 50, 10)
)
tm2 <- tidymatrix(mat, row_data, col_data2)

lm_results <- tm2 |>
  activate(rows) |>
  compute_across(
    fn = function(vals, meta) {
      fit <- lm(vals ~ condition + batch + age, data = meta)
      summ <- summary(fit)
      coef_summ <- coef(summ)

      list(
        condition_pval = coef_summ["conditionTreatment", "Pr(>|t|)"],
        condition_coef = coef_summ["conditionTreatment", "Estimate"],
        r.squared = summ$r.squared
      )
    }
  )

# Add results to metadata
tm_with_stats <- tm |>
  activate(rows) |>
  compute_across(
    fn = function(vals, meta) {
      test <- t.test(vals ~ meta$condition)
      list(p.value = test$p.value)
    },
    add_to_data = TRUE,
    prefix = "ttest"
  )

Convenience wrapper for ANOVA

Description

Performs one-way ANOVA for each row (or column) of a tidymatrix to test for differences across multiple groups.

Usage

compute_anova(
  .data,
  group_col,
  adjust = "fdr",
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE
)

Arguments

.data

A tidymatrix object with rows or columns active

group_col

Character. Name of column in metadata containing group labels

adjust

Character. Method for p-value adjustment. Default "fdr". Use "none" for no adjustment

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

Value

A data.frame/tibble with columns: identifiers, f.statistic, p.value, df_between, df_within, and p.adj (if adjustment applied)

Examples

mat <- matrix(rnorm(150), nrow = 10, ncol = 15)
col_data <- data.frame(
  sample = paste0("S", 1:15),
  condition = rep(c("A", "B", "C"), each = 5)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# ANOVA
results <- tm |>
  activate(rows) |>
  compute_anova(group_col = "condition")

Convenience wrapper for correlation tests

Description

Performs correlation tests between each row (or column) and a continuous variable in the metadata.

Usage

compute_correlation(
  .data,
  var,
  method = "pearson",
  adjust = "fdr",
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE
)

Arguments

.data

A tidymatrix object with rows or columns active

var

Character. Name of continuous variable in metadata to correlate with

method

Character. Correlation method: "pearson" (default), "spearman", or "kendall"

adjust

Character. Method for p-value adjustment. Default "fdr". Use "none" for no adjustment

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

Value

A data.frame/tibble with columns: identifiers, correlation, p.value, and p.adj (if adjustment applied)

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
col_data <- data.frame(
  sample = paste0("S", 1:10),
  age = rnorm(10, 50, 10)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# Pearson correlation
results <- tm |>
  activate(rows) |>
  compute_correlation(var = "age", method = "pearson")

# Spearman correlation
results <- tm |>
  activate(rows) |>
  compute_correlation(var = "age", method = "spearman")

Compute hierarchical clustering on tidymatrix

Description

Perform hierarchical clustering on the matrix, adding cluster assignments to metadata and optionally storing the full hclust object.

Usage

compute_hclust(
  x,
  k = NULL,
  h = NULL,
  name = NULL,
  store = TRUE,
  method = "complete",
  dist_method = "euclidean",
  ...
)

Arguments

x

A tidymatrix object

k

Number of clusters to cut the tree into. If NULL, no cluster assignments are added (only dendrogram is stored).

h

Height at which to cut the tree. Alternative to k.

name

Name for this analysis. Default is "row_hclust" or "column_hclust" depending on active component.

store

If TRUE, stores the full hclust object for later retrieval with get_analysis(). Default is TRUE.

method

Agglomeration method for hclust. Default is "complete". Options: "ward.D", "ward.D2", "single", "complete", "average", "mcquitty", "median", "centroid".

dist_method

Distance method for dist(). Default is "euclidean". Options: "euclidean", "maximum", "manhattan", "canberra", "binary", "minkowski".

...

Additional arguments passed to stats::dist()

Details

This function wraps stats::hclust() and stats::dist(), passing additional parameters directly to them.

Value

A tidymatrix object with cluster assignments added to metadata

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
row_data <- data.frame(id = 1:10, group = rep(c("A", "B"), each = 5))
tm <- tidymatrix(mat, row_data)

# Cluster rows into 3 groups
tm <- tm |>
  activate(rows) |>
  compute_hclust(k = 3, method = "ward.D2")

# Now row_data has row_hclust_cluster column

# Get full hclust object for plotting
hc <- get_analysis(tm, "row_hclust")
plot(hc)

# Multiple clusterings with different k
tm <- tm |>
  activate(rows) |>
  compute_hclust(k = 3, name = "gene_k3") |>
  compute_hclust(k = 5, name = "gene_k5")

Compute k-means clustering on tidymatrix

Description

Perform k-means clustering on the matrix, adding cluster assignments to metadata and optionally storing the full kmeans object.

Usage

compute_kmeans(x, centers, name = NULL, store = TRUE, ...)

Arguments

x

A tidymatrix object

centers

Number of clusters (k) or a set of initial cluster centers.

name

Name for this analysis. Default is "row_kmeans" or "column_kmeans" depending on active component.

store

If TRUE, stores the full kmeans object for later retrieval with get_analysis(). Default is TRUE.

...

Additional arguments passed to stats::kmeans(), such as iter.max, nstart, algorithm, etc.

Details

This function wraps stats::kmeans(), passing additional parameters directly to it.

Value

A tidymatrix object with cluster assignments added to metadata

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
row_data <- data.frame(id = 1:10)
tm <- tidymatrix(mat, row_data)

# K-means clustering with k=3
tm <- tm |>
  activate(rows) |>
  compute_kmeans(centers = 3, nstart = 25)

# Now row_data has row_kmeans_cluster column

# Get full kmeans object
km <- get_analysis(tm, "row_kmeans")
km$tot.withinss  # Total within-cluster sum of squares

Convenience wrapper for Kruskal-Wallis test

Description

Performs Kruskal-Wallis tests for each row (or column) of a tidymatrix. This is a non-parametric alternative to one-way ANOVA.

Usage

compute_kruskal(
  .data,
  group_col,
  adjust = "fdr",
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE
)

Arguments

.data

A tidymatrix object with rows or columns active

group_col

Character. Name of column in metadata containing group labels

adjust

Character. Method for p-value adjustment. Default "fdr". Use "none" for no adjustment

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

Value

A data.frame/tibble with columns: identifiers, statistic, p.value, df, and p.adj (if adjustment applied)

Examples

mat <- matrix(rnorm(150), nrow = 10, ncol = 15)
col_data <- data.frame(
  sample = paste0("S", 1:15),
  condition = rep(c("A", "B", "C"), each = 5)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# Kruskal-Wallis test
results <- tm |>
  activate(rows) |>
  compute_kruskal(group_col = "condition")

Convenience wrapper for linear models

Description

Fits linear models for each row (or column) of a tidymatrix. Can return statistics for a specific coefficient or overall model fit statistics.

Usage

compute_lm(
  .data,
  formula_rhs,
  coef = NULL,
  adjust = "fdr",
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE,
  ...
)

Arguments

.data

A tidymatrix object with rows or columns active

formula_rhs

Right-hand side of formula (e.g., ~ condition + batch + age)

coef

Character. Name of coefficient to extract. If NULL, returns overall model statistics

adjust

Character. Method for p-value adjustment. Default "fdr". Use "none" for no adjustment

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

...

Additional arguments passed to lm()

Value

A data.frame/tibble with model statistics. If coef specified: estimate, p.value, se, and p.adj. If coef = NULL: r.squared, f.statistic, p.value, and p.adj

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
col_data <- data.frame(
  sample = paste0("S", 1:10),
  condition = factor(rep(c("Control", "Treatment"), each = 5)),
  batch = factor(rep(1:2, 5)),
  age = rnorm(10, 50, 10)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# Extract specific coefficient
results <- tm |>
  activate(rows) |>
  compute_lm(
    formula_rhs = ~ condition + batch + age,
    coef = "conditionTreatment"
  )

# Overall model statistics
results <- tm |>
  activate(rows) |>
  compute_lm(formula_rhs = ~ condition + batch + age)

Convenience wrapper for simple linear regression

Description

Performs simple linear regression (one predictor) for each row (or column) of a tidymatrix. For multiple predictors, use compute_lm() instead.

Usage

compute_lm_simple(
  .data,
  predictor,
  adjust = "fdr",
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE
)

Arguments

.data

A tidymatrix object with rows or columns active

predictor

Character. Name of predictor variable in metadata

adjust

Character. Method for p-value adjustment. Default "fdr". Use "none" for no adjustment

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

Value

A data.frame/tibble with columns: identifiers, slope, intercept, r.squared, p.value, and p.adj (if adjustment applied)

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
col_data <- data.frame(
  sample = paste0("S", 1:10),
  age = rnorm(10, 50, 10)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# Simple linear regression
results <- tm |>
  activate(rows) |>
  compute_lm_simple(predictor = "age")

Compute MDS on tidymatrix

Description

Perform Classical Multidimensional Scaling on the matrix, adding MDS coordinates to metadata and optionally storing the distance matrix and result.

Usage

compute_mds(
  x,
  name = NULL,
  k = 2,
  store = TRUE,
  dist_method = "euclidean",
  eig = FALSE,
  ...
)

Arguments

x

A tidymatrix object

name

Name for this analysis. Default is "row_mds" or "column_mds" depending on active component.

k

Number of dimensions for MDS embedding. Default is 2.

store

If TRUE, stores the MDS result for later retrieval with get_analysis(). Default is TRUE.

dist_method

Distance method for dist(). Default is "euclidean". Options: "euclidean", "maximum", "manhattan", "canberra", "binary", "minkowski".

eig

If TRUE, return eigenvalues and GOF statistics (passed to cmdscale).

...

Additional arguments passed to stats::dist() or stats::cmdscale()

Details

This function wraps stats::cmdscale() and stats::dist().

Value

A tidymatrix object with MDS coordinates added to metadata

Examples

mat <- matrix(rnorm(500), nrow = 50, ncol = 10)
row_data <- data.frame(id = 1:50)
tm <- tidymatrix(mat, row_data)

# MDS on rows
tm <- tm |>
  activate(rows) |>
  compute_mds(k = 2)

# Now row_data has row_mds_1, row_mds_2 columns

# Get MDS result
mds_obj <- get_analysis(tm, "row_mds")

Compute PCA on tidymatrix

Description

Perform Principal Component Analysis on the matrix, adding PC scores to metadata and optionally storing the full prcomp object.

Usage

compute_prcomp(x, name = NULL, n_components = NULL, store = TRUE, ...)

Arguments

x

A tidymatrix object

name

Name for this analysis. Default is "row_pca" or "column_pca" depending on active component. Used as prefix for column names.

n_components

Number of PC components to add to metadata. Default is all components. Use a smaller number for large datasets.

store

If TRUE, stores the full prcomp object for later retrieval with get_analysis(). Default is TRUE.

...

Additional arguments passed to stats::prcomp(), such as center, scale., tol, etc.

Details

This function wraps stats::prcomp() and passes all additional parameters directly to it. The PC scores are added as columns to the active metadata (row_data or col_data).

Value

A tidymatrix object with PC scores added to metadata

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
row_data <- data.frame(id = 1:10, group = rep(c("A", "B"), each = 5))
col_data <- data.frame(id = 1:10, type = rep(c("x", "y"), 5))
tm <- tidymatrix(mat, row_data, col_data)

# PCA on columns (samples)
tm <- tm |>
  activate(columns) |>
  compute_prcomp(center = TRUE, scale. = TRUE)

# Now col_data has column_pca_PC1, column_pca_PC2, etc.

# Custom name and limited components
tm <- tm |>
  activate(rows) |>
  compute_prcomp(name = "gene_pca", n_components = 3, center = TRUE)

# Get full prcomp object
pca_obj <- get_analysis(tm, "gene_pca")
summary(pca_obj)
plot(pca_obj$sdev^2 / sum(pca_obj$sdev^2))  # Variance explained

Compute t-SNE on tidymatrix

Description

Perform t-distributed Stochastic Neighbor Embedding on the matrix, adding t-SNE coordinates to metadata and optionally storing the full Rtsne object.

Usage

compute_tsne(x, name = NULL, dims = 2, store = TRUE, perplexity = 30, ...)

Arguments

x

A tidymatrix object

name

Name for this analysis. Default is "row_tsne" or "column_tsne" depending on active component.

dims

Number of dimensions for t-SNE embedding. Default is 2.

store

If TRUE, stores the full Rtsne object for later retrieval with get_analysis(). Default is TRUE.

perplexity

Perplexity parameter (default 30). Should be less than the number of samples. Typical values are between 5 and 50.

...

Additional arguments passed to Rtsne::Rtsne(), such as theta, max_iter, verbose, etc.

Details

This function wraps Rtsne::Rtsne() and passes all additional parameters directly to it. Note that t-SNE is stochastic, so use set.seed() before calling for reproducible results.

Value

A tidymatrix object with t-SNE coordinates added to metadata

Examples

if (requireNamespace("Rtsne", quietly = TRUE)) {
mat <- matrix(rnorm(500), nrow = 50, ncol = 10)
row_data <- data.frame(id = 1:50)
tm <- tidymatrix(mat, row_data)

# t-SNE on rows
set.seed(42)  # For reproducibility
tm <- tm |>
  activate(rows) |>
  compute_tsne(dims = 2, perplexity = 10)

# Now row_data has row_tsne_1, row_tsne_2 columns

# Get full Rtsne object
tsne_obj <- get_analysis(tm, "row_tsne")
}

Convenience wrapper for t-tests

Description

Performs t-tests comparing two groups for each row (or column) of a tidymatrix. Automatically detects groups and applies multiple testing correction.

Usage

compute_ttest(
  .data,
  group_col,
  control = NULL,
  treatment = NULL,
  log2 = TRUE,
  adjust = "fdr",
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE,
  ...
)

Arguments

.data

A tidymatrix object with rows or columns active

group_col

Character. Name of column in metadata containing group labels

control

Character. Label for the control group. If NULL, it is inferred: when group_col is a factor, the first level present in the data; otherwise the first value in sorted order. If only treatment is given, control is the other group.

treatment

Character. Label for the treatment group. If NULL, the group that is not control. When either group is inferred, a message reports the comparison being made. Set both explicitly to silence it.

log2

Logical. If TRUE (default), computes log2 fold change. If FALSE, computes raw fold change

adjust

Character. Method for p-value adjustment. Default "fdr". See ?p.adjust for options. Use "none" for no adjustment

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

...

Additional arguments passed to t.test()

Value

A data.frame/tibble with columns: identifiers, p.value, log2fc (or fc), and p.adj (if adjustment applied)

Examples

mat <- matrix(rnorm(100, mean = 10), nrow = 10, ncol = 10)
col_data <- data.frame(
  sample = paste0("S", 1:10),
  condition = rep(c("Control", "Treatment"), each = 5)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# Simple t-test
results <- tm |>
  activate(rows) |>
  compute_ttest(group_col = "condition")

# With specific group labels
results <- tm |>
  activate(rows) |>
  compute_ttest(
    group_col = "condition",
    control = "Control",
    treatment = "Treatment"
  )

Compute UMAP on tidymatrix

Description

Perform Uniform Manifold Approximation and Projection on the matrix, adding UMAP coordinates to metadata and optionally storing the full umap object.

Usage

compute_umap(
  x,
  name = NULL,
  n_components = 2,
  store = TRUE,
  n_neighbors = 15,
  min_dist = 0.1,
  metric = "euclidean",
  random_state = NULL,
  ...
)

Arguments

x

A tidymatrix object

name

Name for this analysis. Default is "row_umap" or "column_umap" depending on active component.

n_components

Number of dimensions for UMAP embedding. Default is 2.

store

If TRUE, stores the full umap object for later retrieval with get_analysis(). Default is TRUE.

n_neighbors

Size of local neighborhood (default 15). Larger values preserve more global structure, smaller values preserve more local structure.

min_dist

Minimum distance between points in low-dimensional space (default 0.1). Smaller values create tighter, more separated clusters.

metric

Distance metric to use (default "euclidean"). Options include "manhattan", "cosine", "correlation", etc.

random_state

Seed for reproducibility (default NULL). Set to an integer for reproducible results.

...

Additional configuration passed via umap.defaults

Details

This function wraps umap::umap() and passes additional parameters through the config parameter. UMAP is generally faster than t-SNE and better preserves global structure.

Value

A tidymatrix object with UMAP coordinates added to metadata

Examples

if (requireNamespace("umap", quietly = TRUE)) {
mat <- matrix(rnorm(500), nrow = 50, ncol = 10)
row_data <- data.frame(id = 1:50)
tm <- tidymatrix(mat, row_data)

# UMAP on rows
tm <- tm |>
  activate(rows) |>
  compute_umap(n_components = 2, random_state = 42)

# Now row_data has row_umap_1, row_umap_2 columns

# Get full umap object
umap_obj <- get_analysis(tm, "row_umap")
}

Convenience wrapper for Wilcoxon rank-sum test

Description

Performs Wilcoxon rank-sum tests (Mann-Whitney U test) comparing two groups for each row (or column) of a tidymatrix. This is a non-parametric alternative to the t-test that doesn't assume normal distribution.

Usage

compute_wilcox(
  .data,
  group_col,
  control = NULL,
  treatment = NULL,
  log2 = TRUE,
  adjust = "fdr",
  add_to_data = FALSE,
  prefix = NULL,
  return_tibble = TRUE,
  ...
)

Arguments

.data

A tidymatrix object with rows or columns active

group_col

Character. Name of column in metadata containing group labels

control

Character. Label for the control group. If NULL, it is inferred: when group_col is a factor, the first level present in the data; otherwise the first value in sorted order. If only treatment is given, control is the other group.

treatment

Character. Label for the treatment group. If NULL, the group that is not control. When either group is inferred, a message reports the comparison being made. Set both explicitly to silence it.

log2

Logical. If TRUE (default), computes log2 fold change. If FALSE, computes raw fold change

adjust

Character. Method for p-value adjustment. Default "fdr". See ?p.adjust for options. Use "none" for no adjustment

add_to_data

Logical. If TRUE, adds results to row_data/col_data and returns modified tidymatrix. If FALSE (default), returns data.frame

prefix

Character. Optional prefix for result column names when add_to_data = TRUE

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame. Only applies when add_to_data = FALSE

...

Additional arguments passed to wilcox.test()

Value

A data.frame/tibble with columns: identifiers, p.value, log2fc (or fc), median_diff, and p.adj (if adjustment applied)

Examples

mat <- matrix(rnorm(100, mean = 10), nrow = 10, ncol = 10)
col_data <- data.frame(
  sample = paste0("S", 1:10),
  condition = rep(c("Control", "Treatment"), each = 5)
)
row_data <- data.frame(gene = paste0("Gene", 1:10))
tm <- tidymatrix(mat, row_data, col_data)

# Wilcoxon test
results <- tm |>
  activate(rows) |>
  compute_wilcox(group_col = "condition")

Count observations by group

Description

Count the number of observations in each group, aggregating the matrix. Returns a tidymatrix (not a tibble).

Usage

## S3 method for class 'tidymatrix'
count(
  x,
  ...,
  wt = NULL,
  sort = FALSE,
  name = NULL,
  .drop = TRUE,
  .matrix_fn = NULL,
  .matrix_args = list()
)

Arguments

x

A tidymatrix object

...

Variables to group by

wt

Frequency weights (not yet implemented)

sort

If TRUE, sort output in descending order of n

name

Name of count column (default: "n")

.drop

Drop groups with zero observations

.matrix_fn

Function to aggregate matrix values. Default is mean for numeric matrices. Required for non-numeric matrices.

.matrix_args

List of additional arguments to pass to .matrix_fn

Value

A tidymatrix object with counts and aggregated matrix

Examples

library(dplyr, warn.conflicts = FALSE)
mat <- matrix(rnorm(20), nrow = 10, ncol = 2)
row_data <- data.frame(
  id = 1:10,
  group = rep(c("A", "B"), each = 5),
  subgroup = rep(c("x", "y"), 5)
)
tm <- tidymatrix(mat, row_data)

# Count by single variable
tm |>
  activate(rows) |>
  count(group)

# Count by multiple variables
tm |>
  activate(rows) |>
  count(group, subgroup)

# Use different aggregation for matrix
tm |>
  activate(rows) |>
  count(group, .matrix_fn = median)

Filter rows or columns of a tidymatrix

Description

Filter the active component of a tidymatrix object. When rows are active, filters row_data and the corresponding matrix rows. When columns are active, filters col_data and the corresponding matrix columns. Cannot filter when matrix is active.

Usage

## S3 method for class 'tidymatrix'
filter(.data, ..., .preserve = FALSE)

Arguments

.data

A tidymatrix object

...

Logical predicates for filtering

.preserve

Not used (for compatibility with dplyr)

Value

A tidymatrix object with filtered data

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

# Keep respondents aged 60 or over; the matrix rows follow
tm |>
  activate(rows) |>
  filter(age >= 60)

# Keep only the Extraversion items
tm |>
  activate(columns) |>
  filter(trait == "Extraversion")

Get stored analysis object

Description

Retrieve a full analysis object (prcomp, hclust, etc.) that was stored using a compute_* function.

Usage

get_analysis(x, name)

Arguments

x

A tidymatrix object

name

Name of the analysis to retrieve

Value

The stored analysis object (class depends on analysis type)

Examples

tm <- tidymatrix(matrix(rnorm(100), 10, 10))
tm <- tm |>
  activate(columns) |>
  compute_prcomp(name = "pca")

# Get the full prcomp object
pca_obj <- get_analysis(tm, "pca")
summary(pca_obj)
plot(pca_obj$sdev)  # Scree plot

Get suggestions for matrix aggregation functions

Description

Get suggestions for matrix aggregation functions

Usage

get_matrix_fn_suggestions(type)

Group a tidymatrix by variables in metadata

Description

Create a grouped tidymatrix for use with summarize(). Groups are created based on the active metadata component (row_data or col_data).

Usage

## S3 method for class 'tidymatrix'
group_by(.data, ..., .add = FALSE, .drop = TRUE)

Arguments

.data

A tidymatrix object

...

Variables to group by (unquoted names or expressions)

.add

When FALSE (default), group_by() will override existing groups. When TRUE, add to existing groups.

.drop

Drop groups with zero observations

Value

A grouped_tidymatrix object

Examples

library(dplyr, warn.conflicts = FALSE)
mat <- matrix(rnorm(20), nrow = 5, ncol = 4)
row_data <- data.frame(
  id = 1:5,
  group = c("A", "A", "B", "B", "C")
)
tm <- tidymatrix(mat, row_data)

# Group by a variable
tm_grouped <- tm |>
  activate(rows) |>
  group_by(group)

# Multiple grouping variables
row_data2 <- data.frame(
  id = 1:5,
  condition = c("ctrl", "ctrl", "treat", "treat", "treat"),
  batch = c(1, 2, 1, 2, 1)
)
tm2 <- tidymatrix(mat, row_data2)
tm_grouped2 <- tm2 |>
  activate(rows) |>
  group_by(condition, batch)

Get grouping variable names

Description

Get grouping variable names

Usage

## S3 method for class 'grouped_tidymatrix'
group_vars(x)

Arguments

x

A grouped_tidymatrix object

Value

A character vector of grouping variable names

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  group_by(country, gender) |>
  group_vars()

Get grouping variables

Description

Get grouping variables

Usage

## S3 method for class 'grouped_tidymatrix'
groups(x)

Arguments

x

A grouped_tidymatrix object

Value

A list of grouping variable names

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  group_by(country, gender) |>
  groups()

Check if object is a grouped tidymatrix

Description

Check if object is a grouped tidymatrix

Usage

is_grouped_tidymatrix(x)

Arguments

x

An object to test

Value

TRUE if the object is a grouped_tidymatrix, FALSE otherwise

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

is_grouped_tidymatrix(tm)
is_grouped_tidymatrix(tm |> activate(rows) |> group_by(country))

Check if an object is a tidymatrix

Description

Check if an object is a tidymatrix

Usage

is_tidymatrix(x)

Arguments

x

An object to test

Value

TRUE if the object is a tidymatrix, FALSE otherwise

Examples

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

is_tidymatrix(tm)
is_tidymatrix(big5_responses)

Join tidymatrix with another data frame

Description

These functions are tidymatrix methods for dplyr's join functions. They join the row_data or col_data (depending on which is active) with an external data.frame, and appropriately update the matrix dimensions.

Usage

## S3 method for class 'tidymatrix'
left_join(
  x,
  y,
  by = NULL,
  copy = FALSE,
  suffix = c(".x", ".y"),
  ...,
  keep = NULL
)

## S3 method for class 'tidymatrix'
right_join(
  x,
  y,
  by = NULL,
  copy = FALSE,
  suffix = c(".x", ".y"),
  ...,
  keep = NULL
)

## S3 method for class 'tidymatrix'
inner_join(
  x,
  y,
  by = NULL,
  copy = FALSE,
  suffix = c(".x", ".y"),
  ...,
  keep = NULL
)

## S3 method for class 'tidymatrix'
full_join(
  x,
  y,
  by = NULL,
  copy = FALSE,
  suffix = c(".x", ".y"),
  ...,
  keep = NULL
)

## S3 method for class 'tidymatrix'
semi_join(x, y, by = NULL, copy = FALSE, ...)

## S3 method for class 'tidymatrix'
anti_join(x, y, by = NULL, copy = FALSE, ...)

Arguments

x

A tidymatrix object

y

A data frame or tibble to join with

by

A character vector of variables to join by. If NULL, uses all variables that appear in both tables.

copy

If y is not a data frame or tibble, this controls whether to copy it or not.

suffix

If there are non-joined duplicate variables in x and y, these suffixes will be added to disambiguate them.

...

Additional arguments passed to the corresponding dplyr join function

keep

Control which join keys to preserve in the output (see left_join).

Details

Joins work on the active dimension (rows or columns). Use activate() to specify which metadata to join.

When joins add new rows/columns (e.g., right_join, full_join), the matrix is expanded with NA values for the new entries.

When joins remove rows/columns (e.g., inner_join, semi_join, anti_join), the matrix is subset accordingly.

All joins invalidate stored analyses, as the matrix dimensions may have changed.

Value

A tidymatrix object with joined metadata and updated matrix

Examples

library(dplyr)

# Create example tidymatrix
mat <- matrix(rnorm(50), nrow = 10, ncol = 5)
row_data <- data.frame(gene_id = paste0("Gene_", 1:10))
col_data <- data.frame(sample_id = paste0("Sample_", 1:5))
tm <- tidymatrix(mat, row_data, col_data)

# Create external annotation data
annotations <- data.frame(
  gene_id = paste0("Gene_", c(1:8, 15:17)),
  pathway = sample(c("A", "B"), 11, replace = TRUE)
)

# Left join - keep all genes from tidymatrix
tm_left <- tm |>
  activate(rows) |>
  left_join(annotations, by = "gene_id")

# Inner join - keep only matching genes
tm_inner <- tm |>
  activate(rows) |>
  inner_join(annotations, by = "gene_id")

# Full join - keep all genes from both
tm_full <- tm |>
  activate(rows) |>
  full_join(annotations, by = "gene_id")


List stored analyses

Description

Get the names of all analyses stored in a tidymatrix object.

Usage

list_analyses(x)

Arguments

x

A tidymatrix object

Value

Character vector of analysis names

Examples

tm <- tidymatrix(matrix(rnorm(100), 10, 10))
tm <- tm |>
  activate(columns) |>
  compute_prcomp(name = "pca") |>
  compute_hclust(k = 3, name = "clusters")

list_analyses(tm)
# [1] "pca" "clusters"

Log transform matrix

Description

Apply log transformation to matrix values. Requires matrix to be active. This is a convenience wrapper around transform_matrix().

Usage

log_transform(.data, base = 2, offset = 1)

Arguments

.data

A tidymatrix object with matrix active

base

Logarithm base (2, 10, or "natural" for ln)

offset

Value to add before log transform (default 1, for log(x + 1))

Value

A tidymatrix object with log-transformed matrix

Examples

mat <- matrix(abs(rnorm(20)), nrow = 4, ncol = 5)
tm <- tidymatrix(mat)

# Log2 transform with pseudocount
tm_log <- tm |>
  activate(matrix) |>
  log_transform(base = 2, offset = 1)

# Natural log
tm_ln <- tm |>
  activate(matrix) |>
  log_transform(base = "natural")

Matrix Operations for tidymatrix

Description

Functions for transforming and manipulating the matrix component of tidymatrix objects.


Mutate the active component of a tidymatrix

Description

Add or modify columns in the active metadata component. When rows are active, mutates row_data. When columns are active, mutates col_data. Cannot mutate when matrix is active.

Usage

## S3 method for class 'tidymatrix'
mutate(.data, ...)

Arguments

.data

A tidymatrix object

...

Name-value pairs for new or modified columns

Value

A tidymatrix object with mutated metadata

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  mutate(age_group = if_else(age < 40, "young", "older"))

Create a pheatmap from tidymatrix

Description

Generate a heatmap using the pheatmap package, automatically using tidymatrix metadata for annotations and stored clustering results.

Usage

plot_pheatmap(
  x,
  row_names = NULL,
  col_names = NULL,
  row_annotation = NULL,
  col_annotation = NULL,
  row_cluster = NULL,
  col_cluster = NULL,
  ...
)

Arguments

x

A tidymatrix object

row_names

Column name from row_data to use for row names. If NULL, uses sequential numbers.

col_names

Column name from col_data to use for column names. If NULL, uses sequential numbers.

row_annotation

Character vector of column names from row_data to include as row annotations. If NULL, includes all columns except the name column. Set to FALSE to exclude row annotations.

col_annotation

Character vector of column names from col_data to include as column annotations. If NULL, includes all columns except the name column. Set to FALSE to exclude column annotations.

row_cluster

Name of stored hclust analysis to use for row clustering, or TRUE to let pheatmap cluster, or FALSE for no clustering, or NULL to auto-detect stored clustering. Default NULL (auto-detect).

col_cluster

Name of stored hclust analysis to use for column clustering, or TRUE to let pheatmap cluster, or FALSE for no clustering, or NULL to auto-detect stored clustering. Default NULL (auto-detect).

...

Additional arguments passed to pheatmap::pheatmap()

Value

A pheatmap object

Examples

if (requireNamespace("pheatmap", quietly = TRUE)) {
mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
row_data <- data.frame(
  gene = paste0("Gene_", 1:10),
  type = rep(c("A", "B"), each = 5)
)
col_data <- data.frame(
  sample = paste0("Sample_", 1:10),
  condition = rep(c("Control", "Treatment"), 5)
)
tm <- tidymatrix(mat, row_data, col_data)

# Basic heatmap (auto-detects stored clustering if available)
plot_pheatmap(tm, row_names = "gene", col_names = "sample")

# With stored clustering (auto-detected)
tm <- tm |>
  activate(rows) |>
  compute_hclust(k = 2, name = "gene_clusters") |>
  activate(columns) |>
  compute_hclust(k = 2, name = "sample_clusters")

# Auto-detects and uses gene_clusters and sample_clusters
plot_pheatmap(tm, row_names = "gene", col_names = "sample")

# Explicitly specify which clustering to use
plot_pheatmap(tm,
  row_names = "gene",
  col_names = "sample",
  row_cluster = "gene_clusters",
  col_cluster = "sample_clusters"
)

# No clustering (explicit)
plot_pheatmap(tm,
  row_names = "gene",
  col_names = "sample",
  row_cluster = FALSE,
  col_cluster = FALSE
)

# Automatic pheatmap clustering (not using stored)
plot_pheatmap(tm,
  row_names = "gene",
  col_names = "sample",
  row_cluster = TRUE,
  col_cluster = TRUE
)
}

Print a grouped tidymatrix

Description

Print a grouped tidymatrix

Usage

## S3 method for class 'grouped_tidymatrix'
print(x, ...)

Arguments

x

A grouped_tidymatrix object

...

Additional arguments (currently unused)

Value

x, invisibly.

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  group_by(country) |>
  print()

Print a tidymatrix object

Description

Print a tidymatrix object

Usage

## S3 method for class 'tidymatrix'
print(x, ...)

Arguments

x

A tidymatrix object

...

Additional arguments (currently unused)

Value

x, invisibly.

Examples

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

print(tm)
print(activate(tm, columns))

Extract a column from metadata as a vector

Description

Extract a single column from the active metadata component as a vector.

Usage

## S3 method for class 'tidymatrix'
pull(.data, var = -1, ...)

Arguments

.data

A tidymatrix object

var

Column name to extract (can be unquoted)

...

Not used

Value

A vector containing the column values

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(columns) |>
  pull(trait) |>
  table()

Extract the active component from a tidymatrix

Description

Returns the currently active component of a tidymatrix object. This can be the matrix itself, the row metadata, or the column metadata, depending on what is currently activated.

Usage

pull_active(.data)

Arguments

.data

A tidymatrix object

Value

The active component:

Examples

mat <- matrix(rnorm(100), nrow = 10, ncol = 10)
row_data <- data.frame(id = 1:10, group = rep(c("A", "B"), each = 5))
tm <- tidymatrix(mat, row_data)

# Extract matrix
tm |>
  activate(matrix) |>
  pull_active()

# Extract row metadata, e.g. for plotting
pcs <- tm |>
  activate(rows) |>
  compute_prcomp(center = TRUE, scale. = TRUE) |>
  pull_active()

if (requireNamespace("ggplot2", quietly = TRUE)) {
  library(ggplot2)
  ggplot(pcs, aes(x = row_pca_PC1, y = row_pca_PC2, color = group)) +
    geom_point()
}

Objects exported from other packages

Description

These objects are imported from other packages. Follow the links below to see their documentation.

tibble

as_tibble()


Relocate columns in metadata

Description

Change the order of columns in the active metadata component.

Usage

## S3 method for class 'tidymatrix'
relocate(.data, ..., .before = NULL, .after = NULL)

Arguments

.data

A tidymatrix object

...

Columns to move

.before

Column to move before

.after

Column to move after

Value

A tidymatrix object with reordered metadata columns

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  relocate(country, .after = respondent_id)

Remove all stored analyses

Description

Internal helper to remove all analyses when data is modified. Used by filter, slice, and other data modification operations.

Usage

remove_all_analyses(x, operation = "data modification")

Arguments

x

A tidymatrix object

operation

Name of the operation causing removal (for warning message)

Value

A tidymatrix object with analyses removed


Remove stored analysis

Description

Remove a stored analysis object from a tidymatrix. This removes the full analysis object but keeps any metadata columns that were added.

Usage

remove_analysis(x, name = NULL)

Arguments

x

A tidymatrix object

name

Name of the analysis to remove. If NULL, removes all analyses.

Value

A tidymatrix object with the analysis removed

Examples

tm <- tidymatrix(matrix(rnorm(100), 10, 10))
tm <- tm |>
  activate(columns) |>
  compute_prcomp(name = "pca")

# Remove specific analysis
tm <- remove_analysis(tm, "pca")

# Remove all analyses
tm <- remove_analysis(tm, NULL)

Rename columns in metadata

Description

Rename columns in the active metadata component. When rows are active, renames columns in row_data. When columns are active, renames columns in col_data.

Usage

## S3 method for class 'tidymatrix'
rename(.data, ...)

Arguments

.data

A tidymatrix object

...

Name-value pairs for renaming (new_name = old_name)

Value

A tidymatrix object with renamed metadata columns

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(columns) |>
  rename(item = item_id)

Scale rows or columns

Description

Perform z-score scaling (center and scale to unit variance) on the active dimension. Requires rows or columns to be active.

Usage

## S3 method for class 'tidymatrix'
scale(x, center = TRUE, scale = TRUE)

Arguments

x

A tidymatrix object with rows or columns active

center

If TRUE (default), center to mean = 0

scale

If TRUE (default), scale to sd = 1

Value

A tidymatrix object with scaled matrix

Examples

mat <- matrix(rnorm(20, mean = 10, sd = 5), nrow = 4, ncol = 5)
tm <- tidymatrix(mat)

# Scale rows (z-score per row)
tm_scaled <- tm |>
  activate(rows) |>
  scale()

# Scale columns
tm_scaled <- tm |>
  activate(columns) |>
  scale()

# Only center, don't scale
tm_centered <- tm |>
  activate(rows) |>
  scale(center = TRUE, scale = FALSE)

Select columns from row or column metadata

Description

Select columns from the active metadata component. When rows are active, selects from row_data. When columns are active, selects from col_data. Cannot select when matrix is active.

Usage

## S3 method for class 'tidymatrix'
select(.data, ...)

Arguments

.data

A tidymatrix object

...

Column selection expressions

Value

A tidymatrix object with selected metadata columns

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  select(respondent_id, age, country)

Slice rows or columns by position

Description

Select rows or columns by their integer positions. When rows are active, slices row_data and the corresponding matrix rows. When columns are active, slices col_data and the corresponding matrix columns.

Usage

## S3 method for class 'tidymatrix'
slice(.data, ..., .preserve = FALSE)

Arguments

.data

A tidymatrix object

...

Integer positions or expressions to select

.preserve

Not used (for compatibility with dplyr)

Value

A tidymatrix object with sliced data

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  slice(1:10)

Select first or last rows/columns

Description

Select first or last rows/columns

Usage

## S3 method for class 'tidymatrix'
slice_head(.data, n, prop, ...)

## S3 method for class 'tidymatrix'
slice_tail(.data, n, prop, ...)

Arguments

.data

A tidymatrix object

n

Number of rows/columns to select

prop

Proportion of rows/columns to select

...

Not used

Value

A tidymatrix object with selected data

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  slice_head(n = 5)

tm |>
  activate(columns) |>
  slice_tail(prop = 0.1)

Select a random sample of rows/columns

Description

Select a random sample of rows/columns

Usage

## S3 method for class 'tidymatrix'
slice_sample(.data, n, prop, weight_by = NULL, replace = FALSE, ...)

Arguments

.data

A tidymatrix object

n

Number of rows/columns to select

prop

Proportion of rows/columns to select

weight_by

Sampling weights (not yet implemented)

replace

Sample with replacement

...

Not used

Value

A tidymatrix object with sampled data

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

tm |>
  activate(rows) |>
  slice_sample(n = 20)

Store an analysis object

Description

Internal function to store analysis results as an attribute.

Usage

store_analysis(x, name, object, active)

Arguments

x

A tidymatrix object

name

Name for the analysis

object

The analysis object to store

active

Which component was active ("rows" or "columns")

Value

A tidymatrix object with the analysis stored


Summarize grouped tidymatrix

Description

Aggregate grouped rows or columns, applying summary functions to metadata and aggregating the matrix. For numeric matrices, the default aggregation is mean(). For non-numeric matrices, you must specify .matrix_fn.

Usage

## S3 method for class 'grouped_tidymatrix'
summarize(.data, ..., .matrix_fn = NULL, .matrix_args = list(), .groups = NULL)

## S3 method for class 'grouped_tidymatrix'
summarise(.data, ..., .matrix_fn = NULL, .matrix_args = list(), .groups = NULL)

Arguments

.data

A grouped_tidymatrix object

...

Name-value pairs of summary functions for metadata

.matrix_fn

Function to aggregate matrix values within each group. Default is mean for numeric matrices. Required for non-numeric matrices.

.matrix_args

List of additional arguments to pass to .matrix_fn (e.g., list(na.rm = TRUE))

.groups

Grouping structure of result (same as dplyr::summarize)

Value

An ungrouped tidymatrix object with aggregated data. Stored analysis objects are removed (with a warning), because they describe the data before aggregation; metadata columns are kept.

Examples

library(dplyr, warn.conflicts = FALSE)
mat <- matrix(rnorm(20), nrow = 10, ncol = 2)
row_data <- data.frame(
  id = 1:10,
  group = rep(c("A", "B"), each = 5)
)
tm <- tidymatrix(mat, row_data)

# Summarize with default (mean) for numeric matrix
tm |>
  activate(rows) |>
  group_by(group) |>
  summarize(n = n(), avg_id = mean(id))

# Use different aggregation function
tm |>
  activate(rows) |>
  group_by(group) |>
  summarize(n = n(), .matrix_fn = median)

# With additional arguments
mat_na <- mat
mat_na[1, 1] <- NA
tm_na <- tidymatrix(mat_na, row_data)
tm_na |>
  activate(rows) |>
  group_by(group) |>
  summarize(n = n(), .matrix_fn = mean, .matrix_args = list(na.rm = TRUE))

Summarize columns

Description

Summarize columns

Usage

summarize_columns(.data, ..., .matrix_fn, .matrix_args, .groups)

Summarize rows

Description

Summarize rows

Usage

summarize_rows(.data, ..., .matrix_fn, .matrix_args, .groups)

Transpose a tidymatrix

Description

Transpose the matrix and swap row and column metadata. This operation works regardless of which component is active. Uses base R's t() generic, so it works correctly even when tidyverse is loaded.

Usage

## S3 method for class 'tidymatrix'
t(x)

Arguments

x

A tidymatrix object

Value

A tidymatrix object with transposed matrix and swapped metadata

Examples

mat <- matrix(1:12, nrow = 3, ncol = 4)
row_data <- data.frame(gene = paste0("G", 1:3))
col_data <- data.frame(sample = paste0("S", 1:4))
tm <- tidymatrix(mat, row_data, col_data)

# Transpose: genes × samples → samples × genes
tm_t <- t(tm)
dim(tm_t$matrix)  # Now 4 × 3

Count observations in groups

Description

Count observations within existing groups. This is a wrapper around summarize(n = n()).

Usage

## S3 method for class 'grouped_tidymatrix'
tally(x, wt = NULL, sort = FALSE, name = NULL)

Arguments

x

A grouped_tidymatrix object

wt

Frequency weights (not yet implemented)

sort

If TRUE, sort output in descending order of n

name

Name of count column (default: "n")

Details

The matrix is aggregated with mean(), which requires a numeric matrix. dplyr::tally() takes no ..., so the aggregation function cannot be overridden here; call summarize(n = n(), .matrix_fn = ...) directly instead.

Value

A tidymatrix object with counts and aggregated matrix

Examples

library(dplyr, warn.conflicts = FALSE)
mat <- matrix(rnorm(20), nrow = 10, ncol = 2)
row_data <- data.frame(
  id = 1:10,
  group = rep(c("A", "B"), each = 5)
)
tm <- tidymatrix(mat, row_data)

# Tally within groups
tm |>
  activate(rows) |>
  group_by(group) |>
  tally()

Create a tidymatrix object

Description

A tidymatrix combines a matrix with row and column metadata, enabling tidyverse-style manipulation. The object stores three components: the matrix data, row annotations, and column annotations.

Usage

tidymatrix(matrix, row_data = NULL, col_data = NULL)

Arguments

matrix

A numeric matrix

row_data

A data.frame with row metadata. Must have same number of rows as the matrix. If NULL, a data.frame with row indices is created.

col_data

A data.frame with column metadata. Must have same number of rows as the matrix has columns. If NULL, a data.frame with column indices is created.

Value

A tidymatrix object

Examples

# Create a simple tidymatrix
mat <- matrix(rnorm(12), nrow = 4, ncol = 3)
row_data <- data.frame(
  person_id = 1:4,
  age = c(25, 30, 35, 40),
  gender = c("M", "F", "F", "M")
)
col_data <- data.frame(
  question_id = 1:3,
  type = c("numeric", "categorical", "numeric")
)
tm <- tidymatrix(mat, row_data, col_data)

Convert tidymatrix to long-format data.frame

Description

Converts a tidymatrix into a long-format data.frame where each row represents a single matrix cell with its associated row and column metadata. This is useful for plotting individual data points or performing analyses that require long-format data.

Usage

to_long(.data, return_tibble = TRUE)

Arguments

.data

A tidymatrix object

return_tibble

Logical. If TRUE (default), returns tibble. If FALSE, returns data.frame.

Details

The conversion always processes the entire matrix regardless of which component is active. If column names conflict between row_data and col_data, all row metadata columns are prefixed with "row." and all column metadata columns are prefixed with "col." to avoid ambiguity.

Matrix values are unwrapped in column-major order (R's default), meaning all values from column 1, then all values from column 2, etc.

Value

A data.frame/tibble with m*n rows (where m and n are matrix dimensions) containing:

Examples

# Basic conversion to long format
mat <- matrix(1:12, nrow = 4, ncol = 3)
row_data <- data.frame(
  gene_id = paste0("Gene", 1:4),
  gene_type = c("A", "A", "B", "B")
)
col_data <- data.frame(
  sample_id = paste0("Sample", 1:3),
  condition = c("Control", "Treatment", "Control")
)
tm <- tidymatrix(mat, row_data, col_data)

long <- to_long(tm)
head(long)

# Use in ggplot2 workflow
if (requireNamespace("ggplot2", quietly = TRUE)) {
  library(ggplot2)
  tm |>
    to_long() |>
    ggplot(aes(x = sample_id, y = value, color = condition)) +
    geom_point() +
    facet_wrap(~gene_id)
}

# Statistics added to the metadata carry over to the long format
big5 <- tidymatrix(big5_responses, big5_respondents, big5_items) |>
  activate(columns) |>
  compute_ttest(
    group_col = "gender", control = "Female", treatment = "Male",
    add_to_data = TRUE
  )
long_big5 <- to_long(big5)
head(long_big5[long_big5$p.adj < 0.05, ])

Apply a function to the matrix

Description

Applies a function to the matrix with behavior determined by the active component:

Usage

transform_matrix(.data, fn, ...)

Arguments

.data

A tidymatrix object

fn

A function to apply. When matrix is active, receives the full matrix. When rows or columns are active, receives one row or column vector at a time. Must return values with the same dimensions.

...

Additional arguments passed to fn. With rows active they may refer to columns of col_data; with columns active, to columns of row_data. Avoid argument names that partially match fn (such as f), as R would match them to fn.

Details

Additional arguments in ... are passed on to fn. With rows or columns active they are evaluated with the metadata of the other dimension as a data mask, in the same way as in dplyr::mutate(). A row vector has one element per matrix column, so with rows active a column of col_data lines up element by element with the vector that fn receives (and likewise for columns and row_data). This makes it possible to transform values depending on their metadata, e.g. to reverse-score some questionnaire items (see examples). Use .env$x to refer to a variable x in the calling environment when a metadata column has the same name. With the matrix active, arguments are evaluated normally.

Value

A tidymatrix object with the transformed matrix

Examples

mat <- matrix(1:12, nrow = 3, ncol = 4)
tm <- tidymatrix(mat)

# Element-wise: apply log to entire matrix
tm |>
  activate(matrix) |>
  transform_matrix(log)

# Row-wise: rank values within each row
tm |>
  activate(rows) |>
  transform_matrix(rank)

# Column-wise: min-max normalize each column to [0, 1]
tm |>
  activate(columns) |>
  transform_matrix(\(x) (x - min(x)) / (max(x) - min(x)))

# Passing extra arguments: round to 2 decimal places
tm |>
  activate(matrix) |>
  transform_matrix(round, digits = 2)

# Arguments can use the metadata of the other dimension. Reverse-score
# the questionnaire items flagged in the column metadata (1 <-> 5):
big5 <- tidymatrix(big5_responses, big5_respondents, big5_items)
big5 |>
  activate(rows) |>
  transform_matrix(\(x, flip) ifelse(flip, 6L - x, x), flip = reversed)

Remove grouping from a grouped tidymatrix

Description

Remove grouping from a grouped tidymatrix

Usage

## S3 method for class 'grouped_tidymatrix'
ungroup(x, ...)

Arguments

x

A grouped_tidymatrix object

...

Not used

Value

A tidymatrix object (ungrouped)

Examples

library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

grouped <- tm |>
  activate(rows) |>
  group_by(country)

ungroup(grouped)

Validate a tidymatrix object

Description

Validate a tidymatrix object

Usage

validate_tidymatrix(tm)

Arguments

tm

A tidymatrix object

Value

The tidymatrix object (invisibly) if valid, otherwise throws an error