hilldiv3 measures and compares the diversity of
biological communities (OTU/ASV/MAG tables) using Hill
numbers, a single family of metrics that unifies richness,
Shannon and Simpson diversity through one parameter, the diversity
order q.
If Hill numbers are new to you, the key idea is the effective number of taxa: every value answers “how many equally-abundant taxa would give this much diversity?”. A community of 10 taxa where one dominates and the rest are rare behaves, in practice, like far fewer than 10 — and the Hill number says exactly how many. Because all the metrics below share this one currency, you can compare them directly across samples, studies and diversity types.
From the same framework hilldiv3 derives diversity
partitioning, (dis)similarity,
profiles, evenness and
redundancy, for three flavours of diversity:
| Flavour | What it accounts for | How you ask for it |
|---|---|---|
| Neutral | abundances only | counts |
| Phylogenetic | evolutionary relatedness | counts + tree |
| Functional | trait dissimilarity | counts + dist |
The diversity type is inferred from the inputs you pass —
you call the same functions either way. You can also state it explicitly
with type = "neutral" | "phylogenetic" | "functional" to
have it validated against your inputs.
Every function takes a count table with taxa (OTUs/ASVs/MAGs) in rows and samples in columns. A matrix is the simplest form:
counts <- matrix(
c(10, 0, 5,
2, 8, 1,
3, 4, 0,
6, 2, 7),
nrow = 3, byrow = FALSE,
dimnames = list(c("t1", "t2", "t3"),
c("s1", "s2", "s3", "s4"))
)
counts
#> s1 s2 s3 s4
#> t1 10 2 3 6
#> t2 0 8 4 2
#> t3 5 1 0 7Data frames, tibbles, phyloseq objects and
TreeSummarizedExperiment objects work too — see Preparing
your data.
The package also ships a small simulated
gut-microbiome example — gut_counts (a MAG count table),
gut_tree (a phylogeny) and gut_traits (a trait
table) — used throughout the website articles.
hilldiv() returns Hill numbers per sample — the
diversity within each community, traditionally called
alpha diversity. By default it computes orders
q = 0 (richness), q = 1 (Shannon diversity)
and q = 2 (Simpson diversity):
hilldiv(counts)
#> Computing "neutral" Hill numbers of "q0", "q1", and "q2".
#> ℹ 3 taxa across 4 samples.
#> <hilldiv3 result: neutral>
#> 12 rows x 3 cols
#>
#> q sample value
#> 1 0 s1 2.000000
#> 2 1 s1 1.889882
#> 3 2 s1 1.800000
#> 4 0 s2 3.000000
#> 5 1 s2 2.137309
#> 6 2 s2 1.753623
#> 7 0 s3 2.000000
#> 8 1 s3 1.979626
#> 9 2 s3 1.960000
#> 10 0 s4 3.000000
#> 11 1 s4 2.693484
#> 12 2 s4 2.528090Higher q down-weights rare taxa, so qD
decreases as q grows unless the sample is perfectly even.
Add a tree or a distance matrix to layer on more flavours: a
tree adds phylogenetic diversity alongside neutral, and
supplying both a tree and a dist returns all
three types at once (a type column tells them apart).
Restrict the output with type =.
tree <- ape::read.tree(text = "((t1:1,t2:1):1,t3:2);")
hilldiv(counts, tree = tree) # neutral + phylogenetic
#> Computing "neutral" and "phylogenetic" Hill numbers of "q0", "q1", and "q2".
#> ℹ 3 taxa across 4 samples.
#> <hilldiv3 result: neutral, phylogenetic>
#> 24 rows x 4 cols
#>
#> q sample type value
#> 1 0 s1 neutral 2.000000
#> 2 1 s1 neutral 1.889882
#> 3 2 s1 neutral 1.800000
#> 4 0 s2 neutral 3.000000
#> 5 1 s2 neutral 2.137309
#> 6 2 s2 neutral 1.753623
#> 7 0 s3 neutral 2.000000
#> 8 1 s3 neutral 1.979626
#> 9 2 s3 neutral 1.960000
#> 10 0 s4 neutral 3.000000
#> 11 1 s4 neutral 2.693484
#> 12 2 s4 neutral 2.528090
#> 13 0 s1 phylogenetic 2.000000
#> 14 1 s1 phylogenetic 1.889882
#> 15 2 s1 phylogenetic 1.800000
#> 16 0 s2 phylogenetic 2.500000
#> 17 1 s2 phylogenetic 1.702490
#> 18 2 s2 phylogenetic 1.423529
#> 19 0 s3 phylogenetic 1.500000
#> 20 1 s3 phylogenetic 1.406992
#> 21 2 s3 phylogenetic 1.324324
#> 22 0 s4 phylogenetic 2.500000
#> 23 1 s4 phylogenetic 2.318405
#> 24 2 s4 phylogenetic 2.227723
hilldiv(counts, tree = tree, type = "phylogenetic") # phylogenetic only
#> Computing "phylogenetic" Hill numbers of "q0", "q1", and "q2".
#> ℹ 3 taxa across 4 samples.
#> <hilldiv3 result: phylogenetic>
#> 12 rows x 3 cols
#>
#> q sample value
#> 1 0 s1 2.000000
#> 2 1 s1 1.889882
#> 3 2 s1 1.800000
#> 4 0 s2 2.500000
#> 5 1 s2 1.702490
#> 6 2 s2 1.423529
#> 7 0 s3 1.500000
#> 8 1 s3 1.406992
#> 9 2 s3 1.324324
#> 10 0 s4 2.500000
#> 11 1 s4 2.318405
#> 12 2 s4 2.227723hillpart() splits diversity across samples into
alpha (within-sample), gamma (pooled)
and beta (gamma / alpha, the number of
effectively distinct communities):
hillpart(counts)
#> Partitioning neutral Hill numbers of "q0", "q1", and "q2".
#> <hilldiv3 result: neutral>
#> 9 rows x 3 cols
#>
#> q component value
#> 1 0 alpha 2.500000
#> 2 1 alpha 2.154269
#> 3 2 alpha 1.968927
#> 4 0 gamma 3.000000
#> 5 1 gamma 2.905735
#> 6 2 gamma 2.828374
#> 7 0 beta 1.200000
#> 8 1 beta 1.348827
#> 9 2 beta 1.436505hilldiss() and hillsim() turn beta into
bounded dissimilarity / similarity metrics (Sorensen-, Jaccard-, and
UniFrac-type), and hillpair() returns a dist
object of pairwise dissimilarities ready for ordination:
The website carries in-depth articles:
hillprof(),
hilleven(), hillred().phyloseq/TreeSummarizedExperiment,
tss(), traits2dist() and
match_data().