Skip to main content
Ctrl+K
scikit-plots homepage scikit-plots homepage
  • Introduction
  • User Guide
  • APIs Reference
  • Tutorials
  • Learn
    • Tags
    • Code of Conduct
    • Community
    • Developer's Guide
    • Governance Process
    • Release History
    • Roadmap
    • About Us | Project
  • Social Media Platforms
  • PyPI
  • GitHub
  • Introduction
  • User Guide
  • APIs Reference
  • Tutorials
  • Learn
  • Tags
  • Code of Conduct
  • Community
  • Developer's Guide
  • Governance Process
  • Release History
  • Roadmap
  • About Us | Project
  • Social Media Platforms
  • PyPI
  • GitHub

Section Navigation

  • Terminology
    • Beta Distribution
    • Confidence Level
    • Correlation
    • Critical Value
    • Cumulative Distribution Function (CDF)
    • Frequentist
    • IID (Independent and Identically Distributed)
    • Likelihood
    • Margin of Error (MoE)
    • Mean
    • Median
    • Normal Distribution
    • Outlier
    • Population Proportion
    • Probability
    • Probability Density
    • Probability Distribution
    • Probability Mass
    • Proportion
    • Regression Coefficient
    • Sample Mean
    • Sample Standard Deviation
    • Standard Error (SE)
    • Statistical Significance
    • Statistically Significant
    • True Mean (Population Mean)
    • True Population Parameter
    • Z-Score
    • A Priori Power Analysis
    • Chi-square (χ²) Test
    • Clopper–Pearson Interval
    • Compromise Power Analysis
    • Confidence Intervals (CIs)
    • Effect Size (δ)
    • Hypothesis Testing
    • Kolmogorov–Smirnov (KS) Test
    • Minimum Detectable Lift (MDL)
    • P-Value (probability value)
    • Post Hoc Power Analysis
    • Power (1 – β)
    • Power Analysis
    • Sample size
    • Significance Level (α)
    • Statistical Power
    • Statistical Tests
    • T-Test
    • Trivial Effects
    • Two-Proportion Z-Test
    • Type I Error
    • Wilson Score Interval
    • Z-Test
    • AI (Artificial Intelligence)
    • Classification Models
    • Computer Vision (CV)
    • Decision Trees
    • Linear Models
    • LLMs (Large Language Models)
    • Logistic Regression
    • Machine Learning (ML)
    • Medical AI
    • Natural Language Processing (NLP)
    • Neural Networks
    • Regression Models
    • Support Vector Machines (SVMs)
    • Target Variable
    • Class Weighting
    • Cluster-based undersampling
    • Downsampling
    • NearMiss (Distance-based Undersampling)
    • Oversampling
    • Random Undersampling
    • SMOTE (Synthetic Minority Over-sampling Technique)
    • Subsampling
    • Upsampling
    • Accuracy
    • AUC (Area Under the Curve)
    • Average Precision (AP)
    • Binary Classification
    • Classification Probability
    • Discriminatory Power
    • F1-score
    • Gini Coefficient
    • Harmonic Mean
    • Log Loss (also called Logarithmic Loss or Cross-Entropy Loss)
    • Macro AUC
    • Macro AUROC (Macro-Averaged AUROC)
    • Macro Averaging
    • Macro F1
    • Macro Precision
    • Macro Recall
    • Micro AUC
    • Micro AUROC
    • Micro Averaging
    • Micro F1
    • Micro Precision
    • Micro Recall
    • Model Score
    • Multi-label Classification
    • Multiclass AUROC
    • Multiclass Classification
    • Multiclass Precision
    • Multilabel Precision
    • One-vs-Rest (OvR)
    • One-vs-Rest (OvR) AUROC
    • Partial AUC (pAUC)
    • Per-class Precision (sometimes called class-wise precision)
    • Precision (a.k.a. Positive Predictive Value, PPV)
    • Precision–Recall AUC (PR-AUC)
    • Recall
    • ROC Curve (Receiver Operating Characteristic)
    • ROC-AUC (Receiver Operating Characteristic – Area Under Curve, = AUROC)
    • Single-label Classification
    • Weighted Averaging
    • Average Absolute Error (AAE)
    • Baseline Heuristics
    • Bootstrap
    • Bootstrap Confidence Intervals (CIs)
    • Coverage
    • Cramér’s V
    • DeLong’s Test
    • KS Statistic (Kolmogorov–Smirnov Statistic)
    • Likelihood Ratio (LR)
    • Mann–Whitney U Test (also called the Wilcoxon rank-sum test)
    • MASE (Mean Absolute Scaled Error)
    • Mean Absolute Error (MAE)
    • Mean Absolute Percentage Error (MAPE)
    • Mean Squared Error (MSE)
    • Relative accuracy
    • RMSLE (Root Mean Squared Logarithmic Error)
    • Root Mean Squared Error (RMSE)
    • R² (R-squared)
    • sMAPE (Symmetric Mean Absolute Percentage Error)
    • WAPE (Weighted Absolute Percentage Error)
    • WMAPE (Weighted Mean Absolute Percentage Error)
    • A/B Testing
    • A/B/n Test
    • Bayesian Sequential Testing
    • Bayesian Stopping Rules
    • Conversion Rate Uplift
    • Fixed-Horizon Testing
    • Group Sequential Testing
    • Multivariate Test (MVT)
    • Online Experimentation Platforms
    • Optimizely
    • Risk of Peeking
    • Sequential Testing (also called sequential analysis)
    • Stopping Rules
    • Traditional A/B Test (Fixed-Horizon A/B Test)
    • Treatment Effect
    • True Conversion Rate
    • Blended CAC (Customer Acquisition Cost)
    • CAC (Customer Acquisition Cost)
    • Cannibalization
    • Channel-Specific CAC (Customer Acquisition Cost)
    • Churn
    • Cohort
    • Cohort-Based LTV (Simple Version)
    • Conversion Rate (CR)
    • Cost-Per-Click (CPC) Models
    • Cross-Selling
    • CTR (Click-Through Rate)
    • Customer Lifetime
    • Customer Segmentation
    • D2C (Direct-to-Consumer)
    • FTEs
    • Fully Loaded CAC (Customer Acquisition Cost)
    • Gross LTV (Customer Lifetime Value)
    • Gross Margin
    • KYC
    • Lagging Indicators
    • Lead-Gen Software
    • Leading Indicators
    • LTV (Customer Lifetime Value)
    • LTV:CAC Ratio
    • Net LTV (sometimes called Contribution LTV)
    • OpEx
    • Organic CAC (Customer Acquisition Cost)
    • Paid CAC (Customer Acquisition Cost)
    • Predictive LTV (pLTV)
    • Retention
    • Revenue per User (RPU / ARPU)
    • ROI (Return on Investment)
    • SaaS (Software as a Service)
    • Session Length
    • Upselling
    • Valuation Metric
    • Blocked Splits (Single Holdout)
    • Cross-Validation (CV)
    • Data Leakage
    • Evaluation Set
    • Expanding Window Cross-Validation
    • k-fold cross-validation
    • k-fold Stratified Cross-Validation (Stratified CV)
    • Multiclass stratified CV
    • Sliding Window (Rolling Window) Cross-Validation
    • Stratified Group K-Fold
    • Stratified Shuffle Split
    • Time-based splits (a.k.a. Temporal Cross-Validation, Rolling Window Validation)
    • Active Learning
    • Binary Cross-Entropy (BCE)
    • Deep Ensembles
    • Early Stopping
    • Ensemble
    • Epochs
    • FLOPs
    • Full Annotation
    • Hyperparameter
    • Label Noise
    • Log-Odds
    • Logit Space
    • Logits
    • Loss Functions
    • Model Distillation (Knowledge Distillation)
    • Model Weights
    • Quantization
    • Sigmoid Function
    • Softmax Function
    • Squashing Function
    • Underflow
    • Weak Supervision
    • AWS SageMaker
    • Google Experiments
    • Kaggle
    • ONNX (Open Neural Network Exchange)
    • OpenAI API (ML API)
    • TPU Clusters
    • Vertex AI
    • Backorder Rate
    • Crew Overtime
    • Demand Forecasting
    • Fill Rate
    • Long Lead Times
    • Long-Tail Items
    • Lost Sales Value
    • Overstock %
    • Real-Time Inventory Tracking
    • Reorder Point (ROP) Optimization
    • Safety Stock
    • SKU
    • Slow-Moving SKUs
    • Stockout Rate
    • Stockouts
    • Supplier Constraints
    • Supplier Management
    • Advanced Sorting in Spreadsheets
    • Encode (in Feature Engineering)
    • Normalize (in Feature Engineering)
    • Sensitivity in Feature Engineering
    • Demographic Parity (Statistical Parity)
    • Equal Opportunity (Fairness)
    • Equalized Odds (Fairness)
    • Fairness Guardrails
    • Fairness parity
    • Four-Fifths (80%) Rule
    • Predictive Parity (Calibration)
    • Selection Rate
    • Bayes’ Theorem
    • Bayesian Correction
    • Bayesian Decision Theory (BDT)
    • Bayesian Inference.
    • Bayesian Neural Networks (BNNs)
    • Binomial Likelihood
    • Gaussian Processes (GPs)
    • Marginal Likelihood (also called The Model Evidence or Integrated Likelihood)
    • MCMC (Markov Chain Monte Carlo)
    • Parameter(s) of Interest
    • Posterior
    • Posterior belief
    • Posterior Probability
    • Posterior probability of uplift
    • Prior Belief (or Prior Probability)
    • Variational Inference (VI)
    • Bandit Algorithms
    • O’Brien–Fleming (OBF) Method
    • Pocock Method
    • Sequential Probability Ratio Test (SPRT)
    • Sequential Settings
    • Thompson Sampling (TS) in Bandits (Multi-Armed Bandit Problem (MAB))
    • ARIMA (AutoRegressive Integrated Moving Average)
    • Bayesian Time Series
    • Forecast Error
    • Forecasting Benchmarks
    • Forecasting Competitions
    • Log-Space
    • Low-pass Filtering
    • LSTM — Long Short-Term Memory Networks
    • M-Competitions (Makridakis Competitions)
    • Naïve Baseline Forecast
    • Prophet — Time Series Forecasting by Facebook (Meta)
    • Seasonal Lag
    • Seasonality
    • Signal Processing
    • Simple Baseline Methods
    • Temporal autocorrelation (Serial Correlation)
    • Time Series
    • Time Series Forecasting
    • Windows (in Time-Series)
    • AWS SageMaker Endpoints
    • Caching
    • Cloud Inference
    • Cloud Inference with Big Payloads
    • Compute budgets
    • Continuous Retraining
    • Feature Values
    • Guardrails (in ML & Data Systems)
    • Inference Cost (Inference $)
    • Latency Guardrails
    • Manual review minutes
    • Model KPIs (Key Performance Indicators)
    • Model Stability
    • Monitoring Pipelines
    • Ops Health Dashboard
    • Re-scoring
    • Recalibrate Thresholds
    • Recalibration
    • Reweighting
    • SLA (Service Level Agreement)
    • SLA Breach Rate
    • SLA Breaches
    • SLI (Service Level Indicator)
    • SLOs (Service Level Objectives)
    • Cardinality in Categorical Data
    • Categorical Drift
    • Categorical Explosions
    • Classifier Two-Sample Tests (C2STs)
    • Concept Drift
    • Covariate Drift (a.k.a. Covariate Shift)
    • Data Drift
    • Dataset Shift
    • Drift Detection
    • Drift Guardrails
    • Energy Distance
    • Jensen–Shannon (JS) Divergence
    • KS shift (Kolmogorov–Smirnov shift)
    • Kullback–Leibler (KL) Divergence
    • Label Drift (a.k.a. Target Drift)
    • Macro Shifts
    • Maximum Mean Discrepancy (MMD)
    • Off-Distribution
    • PSI (Population Stability Index)
    • Representation Shift
    • Autoencoder
    • Embedding
    • Embedding Similarity
    • Frozen Encoder
    • Balanced Interleaving
    • DCG (Discounted Cumulative Gain)
    • Interleaving Tests
    • Mean Average Precision (MAP)
    • NDCG (Normalized Discounted Cumulative Gain)
    • Probabilistic Interleaving
    • Ranking Algorithms
    • Team Draft Interleaving (TDI)
    • TREC (Text REtrieval Conference)
    • AUUC (Area Under the Uplift Curve)
    • Causal Effect
    • Causal Impact
    • Causal Inference
    • Causal ML (Causal Machine Learning)
    • Causal Trees
    • Cumulative Incremental Gain (CIG)
    • Cumulative Uplift
    • Incremental Conversions
    • Incremental Gain
    • Incremental Recovery Rate (IRR)
    • Incremental Revenue
    • Incremental Sales
    • Qini Coefficient
    • Qini Curve
    • Random Targeting Strategy
    • Revenue net of treatment cost
    • Total Incremental Benefit (TIB)
    • Treatment Cost
    • Uplift
    • Uplift Curve
    • Uplift Models
    • Uplift Random Forests
    • Uplift Score
    • Uplift@k
    • Continuous Probabilistic Forecasts
    • Continuous Ranked Probability Score (CRPS)
    • Deterministic forecasts
    • Full Distribution
    • Pinball Loss (a.k.a. Quantile Loss)
    • Point Forecasts
    • Predicting Percentiles
    • Prediction Intervals (PI)
    • Probabilistic Forecasts
    • Probabilistic Scoring
    • Probability Forecasts
    • Quantile Forecasts
    • Quantile Level
    • Quantile Regression
    • Return Distribution
    • Risk Forecast
    • Risk-Based Decisions
    • Strictly Proper Scoring Rules
    • Value-at-Risk (VaR)
    • Catalog Coverage
    • Cosine Similarity of Item Features
    • Diminishing Utility
    • Diversity (in Recommender Systems)
    • Dominating in Recommender Systems
    • Genre Overlap
    • Hit Rate (HR)
    • Intra-List Diversity (ILD)
    • Item Coverage
    • Jaccard index
    • Novelty (in Recommender Systems)
    • Relevance in Recommender Systems
    • Self-Information of Popularity
    • User Coverage
    • Adaptive ECE (Expected Calibration Error with Adaptive Binning)
    • Brier Score
    • Calibration quality (Model Calibration)
    • Expected Calibration Error (ECE)
    • Isotonic Regression
    • Maximum Calibration Error (MCE)
    • Murphy’s Decomposition
    • Overconfident
    • Platt Scaling
    • Reliability Curves (also called Calibration Curves)
    • Temperature Scaling
    • Underconfident
    • Basel III
    • Counterfactual Explanations
    • Fair Lending laws
    • High-Stakes Domains
    • LIME (Local Interpretable Model-agnostic Explanations)
    • Post-hoc Explainability
    • SHAP (SHapley Additive exPlanations)
  • External Resources
    • Data Resources
    • Documentation Resources
    • Plot Resources
    • Model Resources
    • Research Resources
    • Youtube Resources
  • Data Analytics
    • 🌱 Foundations
    • 🎯 Data-Driven Decisions
    • 📦 Data Preparation
    • 🧽 Data Cleaning & Preparation
    • 📊 Analyze Data
    • 🎨 Data Visualization
    • 🐍 Data Analysis Using Python
    • 💼 Job Search
  • Data Preparation & Analysis
    • Why Do We Analyze Data?
    • The Process of Data Analysis
    • CRISP-DM for Data Science
    • Big Data: Definition, Characteristics, Evolution, and Business Impact
    • The First Step in Knowing Your Data
    • IEEE 754 Floating-Point Standard
    • Discovering Associations Through Data: From Everyday Patterns to Chicago Taxi Trips (September 2022)
    • Taxi Trips – 2022 dataset from the City of Chicago open data portal
    • Objective Selection of the Bin Width for a Time Histogram
    • Measuring Associations in Data
    • Measuring Associations Between Two Continuous Variables
    • Correlation Coefficients in Python (Pearson, Spearman, Kendall)
    • Karl Pearson
    • Harald Cramér
    • What Are Statistical Tests?
    • Eta Squared (η²): Effect Size in ANOVA
    • Understanding Market Baskets and Ideal Customers
    • What Can Association Rules Tell Us?
    • How Association Rules Are Discovered: Concepts, Scale, Measures, and the Apriori Approach
    • Apriori: Frequent Itemsets via the Apriori Algorithm
    • association_rules: Generating Association Rules from Frequent Itemsets (mlxtend)
    • Cross-Selling
    • Stratified Random Sampling
    • Linear Congruential Random Number Generator (LCG)
    • Partitioning Observations to Train Objective Models
    • Putting Similar Observations into Clusters
    • Clustering
    • Recency, Frequency, and Monetary Value (RFM)
    • RFM Analysis
    • Creating Segments of Observations for Business Reasons (RFM)
    • Least Squares Regression
    • Multiple Linear Regression
    • Feature Importance in Linear Regression
    • Forward Selection: Definition and Core Idea
    • Forward Selection and Model Interpretation in Linear Regression
    • Understanding Forward and Backward Stepwise Regression
    • How Shapley Values Work
    • Logistic Regression: Modeling Binary Outcomes via Odds and Log-Odds
    • Maximum Likelihood (MLE): Fitting a Distribution to Observed Data
    • Assessing Model Fit in Logistic Regression
    • Complete and Quasi-Complete Separation in Logistic Regression
    • Forward Selection with Nested Models and Deviance Tests
    • Interpreting and Assessing a Forward-Selection Logistic Regression Model for College Student Retention
    • Motivation of Decision Trees: An Incremental Model of Decision-Making
    • The CART Algorithm
    • Decision Trees as Piecewise Models and Their Predictive Structure
    • How CART Decision Trees Model Interactions
    • Cluster Profiling Using Decision Trees
    • Using Decision Trees to Explain Clustering Results
    • Assessing the Quality of Prediction Models
    • Binary Classification Models – Conceptual Framework and Evaluation Metrics
    • Nominal Classification Models: Model State and Evaluation Metrics
    • Binary Classification Model Evaluation and Threshold Optimization
    • Identifying Outliers Using Residuals and Studentized Residuals
    • AUC–ROC Curve: Evaluating Classification Model Performance
    • Lift Analysis for Direct Mail Campaigns: Concept, Process, and Business Value
  • Bayesian Data Analysis
    • The three steps of Bayesian data analysis
    • General Notation for Statistical Inference
    • Bayesian Inference
    • Discrete Bayesian Examples – Genetics and Spell Checking (with θ)
    • Probability as a Measure of Uncertainty
    • Example — Probabilities from Football Point Spreads
    • Example — Calibration for Record Linkage
    • Some Useful Results from Probability Theory
    • Computation and Software
    • Bayesian Inference in Applied Statistics
    • Estimating a Probability from Binomial Data
    • Posterior as a Compromise Between Data and Prior Information
    • Summarizing Posterior Inference
    • Informative Prior Distributions
    • Normal Distribution with Known Variance
    • Other Standard Single-Parameter Models
    • Informative Prior Distribution for Cancer Rates
    • Noninformative Prior Distributions
    • Weakly Informative Prior Distributions
    • Averaging Over Nuisance Parameters
    • Normal Data with a Noninformative Prior Distribution
    • Normal Data with a Conjugate Prior Distribution
    • Multinomial Model for Categorical Data
    • Multivariate Normal Model with Known Variance
    • Multivariate Normal with Unknown Mean and Variance
    • Example: Bayesian analysis of a bioassay experiment (logistic, nonconjugate)
    • Summary of Elementary Modeling and Computation
    • Normal Approximations to the Posterior Distribution
    • Large-Sample Theory
    • Counterexamples to large-sample (asymptotic) Bayesian theorems
    • Frequency Evaluations of Bayesian Inferences
    • Bayesian interpretations of other statistical methods
    • Constructing a Parameterized Prior Distribution
    • Exchangeability and hierarchical models
    • Bayesian analysis of conjugate hierarchical models
    • Normal model with exchangeable parameters
    • Example: parallel experiments in eight schools
    • Hierarchical modeling applied to a meta-analysis
    • Weakly Informative Priors for Variance Parameters
    • The Place of Model Checking in Applied Bayesian Statistics
    • Do the Inferences from the Model Make Sense?
    • Posterior predictive checking
    • Graphical posterior predictive checks
    • Model checking for the educational testing example
    • Measures of predictive accuracy
    • Model comparison based on predictive performance
    • Model comparison using Bayes factors
    • Continuous model expansion
    • Implicit assumptions and model expansion: an example
    • Bayesian inference requires a model for data collection
    • Data-collection models and ignorability
    • Sample surveys
    • Designed experiments
    • Sensitivity and the role of randomization
    • Observational studies
    • Censoring and truncation
    • Bayesian decision theory in different contexts
    • Using regression predictions: survey incentives
    • Multistage decision making: medical screening
    • Hierarchical decision analysis for home radon
    • Personal vs. institutional decision analysis
    • Numerical integration
    • Distributional approximations
    • Direct simulation and rejection sampling
    • Importance sampling
    • How many simulation draws are needed?
    • Computing environments
    • Debugging Bayesian computing
    • Gibbs sampler
    • Metropolis and Metropolis-Hastings algorithms
    • Using Gibbs and Metropolis as building blocks
    • Inference and assessing convergence
    • Effective number of simulation draws
    • Example: hierarchical normal model
    • Efficient Gibbs samplers
    • Efficient Metropolis jumping rules
    • Further extensions to Gibbs and Metropolis
    • Hamiltonian Monte Carlo
    • Hamiltonian Monte Carlo for a hierarchical model
    • Stan: developing a computing environment
    • Finding posterior modes
    • Boundary-avoiding priors for modal summaries
    • Normal and related mixture approximations
    • Finding marginal posterior modes using EM
    • Conditional and marginal posterior approximations
    • Example: hierarchical normal model (continued)
    • Variational inference
    • Expectation propagation
    • Other approximations
    • Unknown normalizing factors
    • Conditional modeling
    • Bayesian analysis of classical regression
    • Regression for causal inference: incumbency and voting
    • Goals of regression analysis
    • Assembling the matrix of explanatory variables
    • Regularization and dimension reduction
    • Unequal variances and correlations
    • Including numerical prior information
    • Regression coefficients exchangeable in batches
    • Example: forecasting U.S. presidential elections
    • Interpreting a normal prior distribution as extra data
    • Varying intercepts and slopes
    • Computation: batching and transformation
    • Analysis of variance and the batching of coefficients
    • Hierarchical models for batches of variance components
    • Standard generalized linear model likelihoods
    • Working with generalized linear models
    • Weakly informative priors for logistic regression
    • Overdispersed Poisson regression for police stops
    • State-level opinons from national polls
    • Models for multivariate and multinomial responses
    • Loglinear models for multivariate discrete data
    • Aspects of robustness
    • Overdispersed versions of standard models
    • Posterior inference and computation
    • Robust inference for the eight schools
    • Robust regression using t-distributed errors
    • Notation
    • Multiple imputation
    • Missing data in the multivariate normal and t models
    • Example: multiple imputation for a series of polls
    • Missing values with counted data
    • Example: an opinion poll in Slovenia
    • Example: serial dilution assay
    • Example: population toxicokinetics
    • Splines and weighted sums of basis functions
    • Basis selection and shrinkage of coefficients
    • Non-normal models and regression surfaces
    • Gaussian process regression
    • Example: birthdays and birthdates
    • Latent Gaussian process models
    • Functional data analysis
    • Density estimation and regression
    • Setting up and interpreting mixture models
    • Example: reaction times and schizophrenia
    • Label switching and posterior computation
    • Unspecified number of mixture components
    • Mixture models for classification and regression
    • Bayesian histograms
    • Dirichlet process prior distributions
    • Dirichlet process mixtures
    • Beyond density estimation
    • Hierarchical dependence
    • Density regression
  • Time Series
    • What Are Time Series, and How Are They Used?
    • Getting Started with R
    • A Gentle Introduction to Stationarity
    • Weak and Strong Stationarity
    • Linear Processes
    • Understanding ARMA Processes
    • Computing ACFs of Causal AR(2) Processes Using Difference Equations
    • Understanding ACFs via Difference Equations for AR(p) and ARMA(p, q)
    • Best Linear Predictor of a Stationary Process
    • Sample ACF and Sample PACF
    • Preliminary Estimation for AR Models and the Yule–Walker Equations
    • Maximum Likelihood Estimation for ARMA Models (Gaussian MLE)
    • Diagnostics After Fitting a Time Series Model
    • Order Selection for Time Series Models
    • ARIMA Models: How Nonstationary Models Are Built from Stationary Ones
    • SARIMA Models: Seasonal ARIMA
    • Beyond One-Step Ahead Predictions
    • Exponential Smoothing Models
  • Deep Learning
    • What is a Neural Network?
    • Supervised Learning and Neural Networks
    • Why Deep Learning is Taking Off
    • Geoffrey Hinton Interview
    • Binary Classification and Logistic Regression (Neural Network Basics)
    • Logistic Regression (Binary Classification Model)
    • Logistic Regression – Loss Function and Cost Function
    • Gradient Descent in Logistic Regression
    • Derivatives
    • More Derivative Examples
    • Computation Graph
    • Derivatives with a Computation Graph
    • Logistic Regression Gradient Descent
    • Gradient Descent on m Training Examples
    • Vectorization in Logistic Regression
    • More Vectorization Examples
    • Vectorizing Logistic Regression
  • Hands-On Materials
    • BilimEdtech Labs
  • Cheatsheet
    • Md Cheatsheet
    • RST Cheatsheet
  • Glossary
    • https://scikit-learn.org/stable/glossary.html
    • https://ml-cheatsheet.readthedocs.io/en/latest/glossary.html
  • ☀️ Learning Hub • ✨ AI-Powered
  • Bayesian Data Analysis
  • Normal and related mixture approximations

Normal and related mixture approximations#

Part 3 · Stage 10 · 🎛️ Modal & Variational Approximation · Lesson 083 of 144 · intermediate

◀ Previous · Boundary-avoiding priors for modal summaries · Next · Finding marginal posterior modes using EM ▶ · ↑ Section

Important

✨ AI-generated content. This page was written with the assistance of an AI language model and is provided as a learning aid. Despite careful review, it may still contain mistakes, omissions, or out-of-date information. Whether you are new to the topic, a team lead, or a senior practitioner, treat it as a starting point rather than an authoritative reference: read it critically and independently verify anything you act on (code, commands, figures, and factual claims) against official documentation and primary sources before relying on it.

One normal is rarely enough#

The Laplace approximation places a single normal at a single mode. Real posteriors are skewed, heavy-tailed, or multimodal — a mixture likelihood has one mode per labelling, a weakly identified model has ridges. The natural repair is to approximate with a mixture of normals, one component per mode.

Building the mixture#

Run the optimiser from many dispersed starting points, collect the distinct modes \(\hat{\theta}_k\) and their curvatures \(\Sigma_k\), and weight each component by the posterior mass it accounts for — the Laplace estimate of the local integral:

\[p(\theta \mid y) \;\approx\; \sum_{k} \omega_k \, \mathrm{N}\bigl(\theta \mid \hat{\theta}_k, \Sigma_k\bigr), \qquad \omega_k \;\propto\; p(\hat{\theta}_k \mid y) \; |\Sigma_k|^{1/2} .\]

The determinant factor is doing real work: a broad, shallow mode can hold more probability than a narrow, tall one, so mixture weights must account for width, not just height.

Use heavier tails#

If the approximation will serve as a proposal or an importance-sampling density, replace each normal by a multivariate \(t\) with a few degrees of freedom. Light tails are the failure mode that matters: a proposal thinner than the target produces weights with infinite variance, and the estimate degrades without warning (Stage 8).

import numpy as np
from scipy import optimize, stats

modes = [optimize.minimize(neg_log_post, x0, method="BFGS") for x0 in dispersed_starts]
modes = dedupe(modes)                                    # distinct optima only
logw = [-m.fun + 0.5 * np.log(np.linalg.det(m.hess_inv)) for m in modes]
w = np.exp(logw - max(logw)); w /= w.sum()               # mixture weights
comps = [stats.multivariate_t(m.x, m.hess_inv, df=4) for m in modes]   # heavy tails

Honest limits#

Two. Modes found are modes searched for: an optimiser started nowhere near a component will never report it, and in high dimensions dispersed starts cover the space poorly. And the approximation’s quality has no internal diagnostic — importance weights (with their \(\hat{k}\)) give one from outside.

Still, the construction is the conceptual bridge to what follows. Fitting the best member of a chosen family to a posterior, by optimising a divergence rather than by curvature at a mode, is precisely variational inference — and it is the workhorse for models too large to sample.

Hint

Related lessons: Distributional approximations · Finding posterior modes · Variational inference · Importance sampling

See also

Source article Adapted (context, re-expressed) in our own words from: https://insightful-data-lab.com/2025/11/22/normal-and-related-mixture-approximations/ (insightful-data-lab.com).

Tags: purpose: reference topic: data analysis domain: bayesian level: intermediate

previous

Boundary-avoiding priors for modal summaries

next

Finding marginal posterior modes using EM

On this page
  • One normal is rarely enough
  • Building the mixture
  • Use heavier tails
  • Honest limits
Edit on GitHub
Show Source

© Copyright 2024 - 2026 scikit-plots developers (BSD-3 Clause License) 0.5.dev0+git.20260720.c8f33de 2026-07-20T16:01Z.