DiversityAndDissimilarity provides ecology-style diversity indices for species
abundance data, linguistic and demography category counts, observation vectors,
and pairwise assemblage comparisons for dissimilarity or divergence.
This is an early package, and still needs additional testing and verification. All the main pathways appear to be good, but there may be (i) corner cases that are unexplored and (ii) differences in conventions for folks coming to this package from different disciplines. Testing is ongoing and comments are welcome through the Issues tab.
The package is small and has very limited dependencies.
It is designed around dispatchable
index types such as Richness(), Shannon(), Renyi(), Tsallis(),
Simpson(), Jaccard(), BrayCurtis(), Hellinger(), and
JensenShannon(), plus convenience functions for common workflows.
Where other fields use the same mathematical object under a different name,
the package exposes those conventions explicitly; for example,
LinguisticDiversityIndex() and GreenbergDiversityIndex() are aliases by
interpretation for Gini-Simpson diversity.
The maintained manual has seven top-level pages:
- Introduction
- Data Input Formats
- Diversity Indices
- Dissimilarity Indices
- Framework
- Additional Information
- API Reference
The Additional Information page includes the R vegan translation notes, estimator guidance, validation examples, scaling notes, and local documentation build instructions.
From the Julia package prompt:
pkg> add DiversityAndDissimilarityOr, while developing from this repository:
pkg> activate .
pkg> testDiversityAndDissimilarity supports Julia 1.10 and newer 1.x releases.
using DiversityAndDissimilarity
assemblage = Dict(:oak => 12, :ash => 5, :elm => 3)
richness(assemblage)
shannon_entropy(assemblage)
shannon_diversity(assemblage)
renyi_entropy(assemblage, 2)
tsallis_entropy(assemblage, 2)
simpson_index(assemblage)
chao1(assemblage)
sample_coverage(assemblage)
inverse_simpson_index(assemblage)
pielou_evenness(assemblage)
fisher_alpha(assemblage)
alpha_diversity(assemblage)Shannon, Renyi, and Tsallis entropy default to base=2, so entropy values are
reported in bits. Pass another base to the entropy index or convenience
function when you need a different logarithm base.
For a more explicit style, pass an entropy index to entropy or an effective
diversity index to diversity:
diversity(Richness(), assemblage)
entropy(Shannon(; base=2), assemblage)
entropy(Renyi(2), assemblage)
entropy(Tsallis(2), assemblage)
diversity(Shannon(; base=2), assemblage)
diversity(Renyi(2), assemblage)
diversity(Tsallis(2), assemblage)
effective_diversity(Shannon(; base=2), assemblage)
effective_diversity(Renyi(2), assemblage)
diversity(Hill(1), assemblage)Noting that the term "species" comes from ecology but is a proxy for the label used for each category of data, data can be input as
- A dictionary is interpreted to relate "species" to their abundance count;
- A numerical vector, interpreted as giving a sequence of unlabelled abundance counts;
- A non-numerical vector, interpreted as sequence of raw observations;
- A matrix, interpreted as giving abundances for samples in rows and taxa/categories in columns.
- A dataframe, where the column labels give species names.
For data input that does not have species labels, these can be provided with additional keyword arguments to most functions.
In more detail (with examples):
Dictionaries are interpreted as species => abundance:
abundance = Dict("oak" => 12, "ash" => 5, "elm" => 3)
proportions(abundance)Numeric vectors are interpreted as abundance vectors by default:
abundance = [12, 5, 3]
entropy(Shannon(), abundance)Non-numeric vectors are interpreted as raw observations:
observations = ["oak", "ash", "oak", "elm"]
counts(observations)
richness(observations)Pass frequencies=false to treat a numeric vector as raw observations rather
than abundance values:
numeric_observations = [1, 2, 1, 3]
richness(numeric_observations; frequencies=false)Community matrices are interpreted as samples in rows and taxa/categories in columns. Alpha-diversity functions return one value per row:
community = [
1 1 2 0 5
3 0 1 1 0
]
richness(community)
shannon_entropy(community)
chao1(community)
ace(community)
sample_coverage(community)
alpha_diversity(community)Tables.jl-compatible inputs, including DataFrames, can be converted with
community_matrix. By default numeric columns are treated as species columns;
pass species explicitly when the table also contains numeric metadata.
community_matrix(table; species=[:oak, :ash, :elm])
richness(table; species=[:oak, :ash, :elm])
shannon_entropy(table; species=[:oak, :ash, :elm])The Data Input Formats page covers all formats, orientation rules, validation, keyword parameters, and the pre-validated pipeline in full detail.
The generic API is built around small index types:
entropy(index, data)evaluates entropy-family indexes in entropy units.diversity(index, data)evaluates alpha diversity indexes; for entropy families it returns effective diversity.effective_diversity(index, data)returns the effective number of species where that transformation is standard.similarity(index, left, right)anddissimilarity(index, left, right)compare two assemblages.index_metadata(index),index_bounds(index), and trait helpers expose package conventions for generic workflows.
is_metric(Jaccard())
is_semimetric(BrayCurtis())
is_symmetric(KullbackLeibler())
index_bounds(JensenShannon())The descriptor helpers include is_finite, is_symmetric, is_nonnegative,
is_bounded, is_metric, is_triangular, is_pseudometric,
is_quasimetric, is_metametric, is_semimetric, is_premetric,
is_supermetric, is_similarity, and is_dissimilarity. Unknown or
unencoded properties return :unknown.
Supported alpha diversity indices include:
Richness()Shannon(; base=2, estimator=Plugin())Renyi(q; base=2)Tsallis(q; base=2)Simpson()GiniSimpson()GreenbergDiversityIndex()LinguisticDiversityIndex()InverseSimpson()Hill(q)Chao1()ACE(; threshold=10)SampleCoverage()PielouEvenness()FisherAlpha()
The main entropy and effective-diversity formulas are:
Supported pairwise indices include:
Jaccard()SorensenDice()Overlap()BrayCurtis()Ruzicka()TotalVariation()Manhattan()Euclidean()Canberra()Hellinger()Chord()Bhattacharyya()KullbackLeibler()ShannonDifference()JensenDifference()JensenShannon()MorisitaHorn()
Incidence-style comparisons work with dictionaries or observation vectors:
plot_a = Dict(:oak => 12, :ash => 5)
plot_b = Dict(:ash => 4, :elm => 7)
similarity(Jaccard(), plot_a, plot_b)
jaccard_distance(plot_a, plot_b)
sorensen_dice_index(plot_a, plot_b)
overlap_similarity(plot_a, plot_b)Abundance-style comparisons use aligned species abundances:
bray_curtis_dissimilarity(plot_a, plot_b)
ruzicka_similarity(plot_a, plot_b)
morisita_horn_similarity(plot_a, plot_b)When passing numeric abundance vectors to BrayCurtis(), the vectors must have
the same length and positions are treated as shared species. Probability-style
distances normalize the aligned abundances first:
left = [12, 5, 0]
right = [0, 4, 7]
dissimilarity(BrayCurtis(), left, right)
hellinger_distance(left, right)
kullback_leibler_divergence(left, right)
shannon_difference(left, right)
jensen_difference(left, right)
jensen_shannon_similarity(left, right)
jensen_shannon_distance(left, right)
total_variation_distance(left, right)As a rule of thumb, use Jaccard or Sorensen-Dice for presence/absence turnover,
Bray-Curtis for a practical ecological abundance default, Ruzicka when you want
a quantitative Jaccard, Hellinger or Chord for composition comparisons before
ordination, and Jensen-Shannon for a symmetric information-theoretic comparison.
Use Shannon difference when the question is whether two assemblages have
different entropy, and Jensen difference when you want the entropy gap between
the average composition and the average entropy of the two assemblages.
Kullback-Leibler divergence is available when an asymmetric
D_{KL}(left \\Vert right) comparison is intended; it can be infinite when
the right assemblage has zero probability where the left assemblage is
positive.
Low-sample corrections for KL and Jensen-Shannon-style divergences use the same estimator objects as Shannon entropy:
kullback_leibler_divergence(left, right; estimator=MillerMadow())
kullback_leibler_divergence(left, right; estimator=AddGamma(1)) # Laplace
kullback_leibler_divergence(left, right; estimator=AddGamma(0.5)) # Jeffreys
kullback_leibler_divergence(left, right; estimator=HausserStrimmer())
kullback_leibler_divergence(left, right; estimator=ChaoShen())
jensen_shannon_divergence(left, right; estimator=MillerMadow())
jensen_shannon_distance(left, right; estimator=AddGamma(0.5), support=10)AddGamma and HausserStrimmer use the aligned observed support unless a
larger known support is supplied. ChaoShen() applies a Good-Turing-style
unseen-mass correction.
Passing a community matrix or Tables.jl-compatible table as the only data argument returns a pairwise matrix across rows:
distance(BrayCurtis(), community)
distance(Hellinger(), community)
jensen_shannon_distance(community)
jaccard_distance(community)
bray_curtis_distance(table; species=[:oak, :ash, :elm])
labeled_distance(BrayCurtis(), table; label=:site, species=[:oak, :ash, :elm])For scaling, alpha-diversity summaries are generally linear in observed support
or in the number of entries in a community matrix. Full pairwise distance
matrices scale quadratically in the number of samples/sites because they return
an M x M result. Known-support estimators may also scale with the supplied
support size, and bootstrap intervals scale roughly linearly with nboot.
Shannon entropy estimation is represented by dispatchable estimator types. The
default is Plugin(), the direct empirical estimator. Other estimators include
MillerMadow(), HausserStrimmer(), Basharin(), AddGamma(gamma), and
ChaoShen().
Use support when the finite support is known but some categories may have zero
observed counts. Pass an integer support size, or pass a collection of category
labels. Leave support=nothing when only the observed support is known, or when
using ChaoShen() for possible unseen species.
entropy(Shannon(; estimator=Plugin()), assemblage)
entropy(Shannon(; estimator=MillerMadow()), assemblage)
entropy(Shannon(; estimator=AddGamma(1)), assemblage; support=10)
entropy(Shannon(; estimator=ChaoShen()), assemblage)
shannon_diversity(assemblage; estimator=MillerMadow())
shannon_diversity(assemblage; estimator=AddGamma(1), support=10)Analytic or approximate Shannon uncertainty is available for selected estimators, and bootstrap/jackknife helpers are available for integer count data:
shannon_variance(assemblage; estimator=Basharin(), support=10)
shannon_confint(assemblage; estimator=ChaoShen())
bootstrap(Shannon(; estimator=AddGamma(1)), assemblage; support=10)
jackknife(Shannon(; estimator=MillerMadow()), assemblage)Renyi and Tsallis entropies are available as dispatchable index types and
convenience functions. The q = 1 case is evaluated as Shannon entropy, and
effective_diversity returns the corresponding Hill number.
renyi_entropy(assemblage, 2)
renyi_diversity(assemblage, 2)
tsallis_entropy(assemblage, 2)
tsallis_diversity(assemblage, 2)Chao1, ACE, and Good-Turing sample coverage are available for abundance count data:
chao1(assemblage)
ace(assemblage)
ace(assemblage; threshold=10)
sample_coverage(assemblage)These functions also work row-wise on community matrices.
Documenter.jl documentation lives
in docs/.
julia --project=docs docs/make.jlThe generated site is written to docs/build/.
Useful local entry points are:
- Introduction
- Data Input Formats
- Diversity Indices
- Dissimilarity Indices
- Framework
- Additional Information
- API Reference
Repository-level notes include ARCHITECTURE.md, CONTRIBUTIONS.md, the benchmark notes, the validation notes, and the annotated bibliography.
Reproducible benchmark scripts live in benchmark/. They generate a
deterministic simulated community matrix at runtime, so no large dataset is
committed and Git LFS is not needed.
julia --project=. benchmark/julia_benchmark.jl
python3 benchmark/python_benchmark.py
Rscript benchmark/r_vegan_benchmark.R
benchmark/run_report.shThe Python benchmark requires NumPy and uses SciPy for Bray-Curtis when available. The R benchmark requires vegan.
The ordinary package tests are run with:
julia --project=. -e 'using Pkg; Pkg.test()'Additional package-quality checks use Aqua.jl and JET.jl from a dedicated test environment:
julia --project=test/quality -e 'using Pkg; Pkg.develop(PackageSpec(path=pwd())); Pkg.instantiate()'
julia --project=test/quality test/quality/runtests.jlGitHub Actions runs the standard Julia compatibility matrix separately from the Aqua and JET quality workflows, which run only on the latest stable Julia release.
This package was developed with assistance from Claude Code (claude-sonnet-4-6), an AI coding assistant from Anthropic. AI assistance was used for code generation, refactoring, documentation, and benchmarking tasks across multiple development sessions. All design decisions were human-mediated and the resulting code was manually reviewed.