Notes for a hypothetical graduate course in information theory in the Department of Statistics at the University of Auckland.
(c) 2026 Brendon J. Brewer
LICENCE: CC-BY-SA 4.0 International. See LICENCE file for details.
- Review of probability theory and probability distributions.
- Historical background.
- Definition and basic properties of Shannon entropy (e.g., non-negativity).
- Interpretation of Shannon entropy in discrete and continuous cases.
- Joint entropy, conditional entropy, and mutual information.
- Relative entropy, cross entropy, and Kullback-Leibler divergence.
- Derivation of entropies of some common distributions.
- Simple applications --- quantifying uncertainty and relevance.
- Infinitesimal KL divergence and the Fisher metric.
- Jeffreys priors.
- Markov chains and entropy rates.
- Asymptotic equipartition property.
- Communication channels.
- Noisy channel theorem.
- Binary symmetric channel and AWGN channel examples.
- Linear block codes and generator matrices.
- Hamming codes and syndrome decoding.
- Repetition codes and simple decoding rules.
- Relationship between codes and channel capacity.
- The principle of maximum entropy.
- Updating probabilities with maximum entropy.
- Canonical distributions.
- Bayesian updating, Jeffrey conditionalisation, entropic priors.
- Statistical mechanics.
- Liouville's theorem.
- The second law of thermodynamics.
- Lossy compression and distortion measures.
- Rate-distortion function and its interpretation.
- Trade‑offs between fidelity and compression.
- Applications in modern statistics and machine learning.
- Variational inference and ELBO.
- Bayesian experimental design.
- Logarithmic scoring rules.
- Foundations of probability: ordering of statements.
- Foundations of entropy: ordering of questions.
- Summary and unifying themes.