Skip to main content
Back to top
Ctrl
+
K
Introduction
User Guide
APIs Reference
Tutorials
Learn
More
Tags
Code of Conduct
Community
Developer's Guide
Governance Process
Release History
Roadmap
About Us | Project
Search
Ctrl
+
K
System Settings
Light
Dark
Choose version
GitHub
Search
Ctrl
+
K
Introduction
User Guide
APIs Reference
Tutorials
Learn
Tags
Code of Conduct
Community
Developer's Guide
Governance Process
Release History
Roadmap
About Us | Project
System Settings
Light
Dark
Choose version
GitHub
Collapse Sidebar
Expand Sidebar
Section Navigation
Tags
component: model (1)
plot_calibration with examples
domain: bayesian (145)
The three steps of Bayesian data analysis
General Notation for Statistical Inference
Bayesian Inference
Discrete Bayesian Examples – Genetics and Spell Checking (with θ)
Probability as a Measure of Uncertainty
Example — Probabilities from Football Point Spreads
Example — Calibration for Record Linkage
Some Useful Results from Probability Theory
Computation and Software
Bayesian Inference in Applied Statistics
Estimating a Probability from Binomial Data
Posterior as a Compromise Between Data and Prior Information
Summarizing Posterior Inference
Informative Prior Distributions
Normal Distribution with Known Variance
Other Standard Single-Parameter Models
Informative Prior Distribution for Cancer Rates
Noninformative Prior Distributions
Weakly Informative Prior Distributions
Averaging Over Nuisance Parameters
Normal Data with a Noninformative Prior Distribution
Normal Data with a Conjugate Prior Distribution
Multinomial Model for Categorical Data
Multivariate Normal Model with Known Variance
Multivariate Normal with Unknown Mean and Variance
Example: Bayesian analysis of a bioassay experiment (logistic, nonconjugate)
Summary of Elementary Modeling and Computation
Normal Approximations to the Posterior Distribution
Large-Sample Theory
Counterexamples to large-sample (asymptotic) Bayesian theorems
Frequency Evaluations of Bayesian Inferences
Bayesian interpretations of other statistical methods
Constructing a Parameterized Prior Distribution
Exchangeability and hierarchical models
Bayesian analysis of conjugate hierarchical models
Normal model with exchangeable parameters
Example: parallel experiments in eight schools
Hierarchical modeling applied to a meta-analysis
Weakly Informative Priors for Variance Parameters
The Place of Model Checking in Applied Bayesian Statistics
Do the Inferences from the Model Make Sense?
Posterior predictive checking
Graphical posterior predictive checks
Model checking for the educational testing example
Measures of predictive accuracy
Model comparison based on predictive performance
Model comparison using Bayes factors
Continuous model expansion
Implicit assumptions and model expansion: an example
Bayesian inference requires a model for data collection
Data-collection models and ignorability
Sample surveys
Designed experiments
Sensitivity and the role of randomization
Observational studies
Censoring and truncation
Bayesian decision theory in different contexts
Using regression predictions: survey incentives
Multistage decision making: medical screening
Hierarchical decision analysis for home radon
Personal vs. institutional decision analysis
Numerical integration
Distributional approximations
Direct simulation and rejection sampling
Importance sampling
How many simulation draws are needed?
Computing environments
Debugging Bayesian computing
Gibbs sampler
Metropolis and Metropolis-Hastings algorithms
Using Gibbs and Metropolis as building blocks
Inference and assessing convergence
Effective number of simulation draws
Example: hierarchical normal model
Efficient Gibbs samplers
Efficient Metropolis jumping rules
Further extensions to Gibbs and Metropolis
Hamiltonian Monte Carlo
Hamiltonian Monte Carlo for a hierarchical model
Stan: developing a computing environment
Finding posterior modes
Boundary-avoiding priors for modal summaries
Normal and related mixture approximations
Finding marginal posterior modes using EM
Conditional and marginal posterior approximations
Example: hierarchical normal model (continued)
Variational inference
Expectation propagation
Other approximations
Unknown normalizing factors
Conditional modeling
Bayesian analysis of classical regression
Regression for causal inference: incumbency and voting
Goals of regression analysis
Assembling the matrix of explanatory variables
Regularization and dimension reduction
Unequal variances and correlations
Including numerical prior information
Regression coefficients exchangeable in batches
Example: forecasting U.S. presidential elections
Interpreting a normal prior distribution as extra data
Varying intercepts and slopes
Computation: batching and transformation
Analysis of variance and the batching of coefficients
Hierarchical models for batches of variance components
Standard generalized linear model likelihoods
Working with generalized linear models
Weakly informative priors for logistic regression
Overdispersed Poisson regression for police stops
State-level opinons from national polls
Models for multivariate and multinomial responses
Loglinear models for multivariate discrete data
Aspects of robustness
Overdispersed versions of standard models
Posterior inference and computation
Robust inference for the eight schools
Robust regression using t-distributed errors
Notation
Multiple imputation
Missing data in the multivariate normal and t models
Example: multiple imputation for a series of polls
Missing values with counted data
Example: an opinion poll in Slovenia
Example: serial dilution assay
Example: population toxicokinetics
Splines and weighted sums of basis functions
Basis selection and shrinkage of coefficients
Non-normal models and regression surfaces
Gaussian process regression
Example: birthdays and birthdates
Latent Gaussian process models
Functional data analysis
Density estimation and regression
Setting up and interpreting mixture models
Example: reaction times and schizophrenia
Label switching and posterior computation
Unspecified number of mixture components
Mixture models for classification and regression
Bayesian histograms
Dirichlet process prior distributions
Dirichlet process mixtures
Beyond density estimation
Hierarchical dependence
Density regression
Bayesian Data Analysis
domain: cython (11)
Cython quickstart: compile_and_load
Browse and compile templates
Build profiles: fast-debug, release, annotate
Cache and restart reuse
Pin/Alias: stable handles for cached builds
Multi-module package builds (5 package examples)
Multi-file builds: .pxi includes and external headers
C++ mode basics: cppclass and libcpp containers
Vector ops without NumPy: array(‘d’) + memoryviews
Workflow templates (train / hpo / predict) + CLI entry template
Cython: Realtime compile_and_load (.pyx)
domain: mlflow (1)
MLflow
domain: neural network (9)
Visualkeras: Spam Classification Conv1D Dense Example
visualkeras: Spam Dense example
visualkeras: autoencoder example
visualkeras: custom vgg16 example
visualkeras: custom vgg16 show dimension example
visualkeras: EfficientNetV2 example
visualkeras: ResNetV2 example
visualkeras: custom VGG example
visualkeras: Vector Index DB
domain: statistics (5)
plot_ks_statistic with examples
plot_lift with examples
plot_residuals_distribution with examples
Gaussian Mixture Models — AIC, AICc, and BIC Model Selection
plot_residuals_distribution with examples
internal: needs-review (1)
Tag Glossary
level: advanced (264)
Conditional modeling
Bayesian analysis of classical regression
Regression for causal inference: incumbency and voting
Goals of regression analysis
Assembling the matrix of explanatory variables
Regularization and dimension reduction
Unequal variances and correlations
Including numerical prior information
Regression coefficients exchangeable in batches
Example: forecasting U.S. presidential elections
Interpreting a normal prior distribution as extra data
Varying intercepts and slopes
Computation: batching and transformation
Analysis of variance and the batching of coefficients
Hierarchical models for batches of variance components
Standard generalized linear model likelihoods
Working with generalized linear models
Weakly informative priors for logistic regression
Overdispersed Poisson regression for police stops
State-level opinons from national polls
Models for multivariate and multinomial responses
Loglinear models for multivariate discrete data
Aspects of robustness
Overdispersed versions of standard models
Posterior inference and computation
Robust inference for the eight schools
Robust regression using t-distributed errors
Notation
Multiple imputation
Missing data in the multivariate normal and t models
Example: multiple imputation for a series of polls
Missing values with counted data
Example: an opinion poll in Slovenia
Example: serial dilution assay
Example: population toxicokinetics
Splines and weighted sums of basis functions
Basis selection and shrinkage of coefficients
Non-normal models and regression surfaces
Gaussian process regression
Example: birthdays and birthdates
Latent Gaussian process models
Functional data analysis
Density estimation and regression
Setting up and interpreting mixture models
Example: reaction times and schizophrenia
Label switching and posterior computation
Unspecified number of mixture components
Mixture models for classification and regression
Bayesian histograms
Dirichlet process prior distributions
Dirichlet process mixtures
Beyond density estimation
Hierarchical dependence
Density regression
Logistic Regression: Modeling Binary Outcomes via Odds and Log-Odds
Maximum Likelihood (MLE): Fitting a Distribution to Observed Data
Assessing Model Fit in Logistic Regression
Complete and Quasi-Complete Separation in Logistic Regression
Forward Selection with Nested Models and Deviance Tests
Interpreting and Assessing a Forward-Selection Logistic Regression Model for College Student Retention
Motivation of Decision Trees: An Incremental Model of Decision-Making
The CART Algorithm
Decision Trees as Piecewise Models and Their Predictive Structure
How CART Decision Trees Model Interactions
Cluster Profiling Using Decision Trees
Using Decision Trees to Explain Clustering Results
Assessing the Quality of Prediction Models
Binary Classification Models – Conceptual Framework and Evaluation Metrics
Nominal Classification Models: Model State and Evaluation Metrics
Binary Classification Model Evaluation and Threshold Optimization
Identifying Outliers Using Residuals and Studentized Residuals
AUC–ROC Curve: Evaluating Classification Model Performance
Lift Analysis for Direct Mail Campaigns: Concept, Process, and Business Value
Low-pass Filtering
Signal Processing
Time Series
Predictive Parity (Calibration)
Equalized Odds (Fairness)
Equal Opportunity (Fairness)
Demographic Parity (Statistical Parity)
Thompson Sampling (TS) in Bandits (Multi-Armed Bandit Problem (MAB))
Bayesian Decision Theory (BDT)
Bayesian Time Series
Posterior probability of uplift
Gaussian Processes (GPs)
Bayesian Neural Networks (BNNs)
Variational Inference (VI)
MCMC (Markov Chain Monte Carlo)
Sequential Settings
Binomial Likelihood
Posterior belief
Marginal Likelihood (also called The Model Evidence or Integrated Likelihood)
Posterior
Prior Belief (or Prior Probability)
Parameter(s) of Interest
Bayes’ Theorem
Posterior Probability
Sequential Probability Ratio Test (SPRT)
Pocock Method
O’Brien–Fleming (OBF) Method
Ranking Algorithms
Probabilistic Interleaving
Team Draft Interleaving (TDI)
Balanced Interleaving
Causal Impact
Bandit Algorithms
Causal Inference
Temporal autocorrelation (Serial Correlation)
Re-scoring
Drift Detection
AWS SageMaker Endpoints
Cloud Inference with Big Payloads
Cloud Inference
Recalibration
Reweighting
Continuous Retraining
Monitoring Pipelines
Bayesian Correction
Recalibrate Thresholds
Guardrails (in ML & Data Systems)
Model KPIs (Key Performance Indicators)
Windows (in Time-Series)
Autoencoder
Frozen Encoder
Embedding
Representation Shift
Classifier Two-Sample Tests (C2STs)
Energy Distance
Maximum Mean Discrepancy (MMD)
Cardinality in Categorical Data
Categorical Drift
Macro Shifts
Categorical Explosions
Off-Distribution
Model Stability
Feature Values
Four-Fifths (80%) Rule
SLI (Service Level Indicator)
Treatment Cost
Incremental Revenue
Incremental Recovery Rate (IRR)
Incremental Sales
Random Targeting Strategy
Causal ML (Causal Machine Learning)
Cumulative Uplift
Incremental Gain
Total Incremental Benefit (TIB)
Cumulative Incremental Gain (CIG)
Qini Curve
Uplift Score
Uplift Models
Ops Health Dashboard
SLA Breach Rate
SLA (Service Level Agreement)
Prophet — Time Series Forecasting by Facebook (Meta)
LSTM — Long Short-Term Memory Networks
ARIMA (AutoRegressive Integrated Moving Average)
Return Distribution
Value-at-Risk (VaR)
Risk Forecast
Probabilistic Scoring
Full Distribution
Continuous Probabilistic Forecasts
Quantile Forecasts
Point Forecasts
Strictly Proper Scoring Rules
Probability Forecasts
Probabilistic Forecasts
Deterministic forecasts
M-Competitions (Makridakis Competitions)
Forecasting Benchmarks
Seasonal Lag
Simple Baseline Methods
Naïve Baseline Forecast
Forecast Error
Forecasting Competitions
Predicting Percentiles
Prediction Intervals (PI)
Quantile Regression
Quantile Level
Time Series Forecasting
Log-Space
Self-Information of Popularity
Relevance in Recommender Systems
Genre Overlap
Jaccard index
Cosine Similarity of Item Features
Intra-List Diversity (ILD)
Dominating in Recommender Systems
Catalog Coverage
User Coverage
Item Coverage
Diminishing Utility
DCG (Discounted Cumulative Gain)
TREC (Text REtrieval Conference)
Adaptive ECE (Expected Calibration Error with Adaptive Binning)
Maximum Calibration Error (MCE)
Murphy’s Decomposition
Temperature Scaling
Platt Scaling
Isotonic Regression
Underconfident
Overconfident
Risk-Based Decisions
Causal Trees
Uplift Random Forests
Uplift Curve
Causal Effect
Embedding Similarity
Jensen–Shannon (JS) Divergence
Kullback–Leibler (KL) Divergence
Seasonality
Concept Drift
Data Drift
Fair Lending laws
Basel III
High-Stakes Domains
Counterfactual Explanations
LIME (Local Interpretable Model-agnostic Explanations)
SHAP (SHapley Additive exPlanations)
Post-hoc Explainability
Caching
Drift Guardrails
Latency Guardrails
Fairness Guardrails
Dataset Shift
Fairness parity
Bayesian Inference.
Interleaving Tests
Compute budgets
Manual review minutes
Inference Cost (Inference $)
Label Drift (a.k.a. Target Drift)
Covariate Drift (a.k.a. Covariate Shift)
KS shift (Kolmogorov–Smirnov shift)
PSI (Population Stability Index)
Selection Rate
SLOs (Service Level Objectives)
Revenue net of treatment cost
Incremental Conversions
Uplift@k
AUUC (Area Under the Uplift Curve)
Qini Coefficient
SLA Breaches
Continuous Ranked Probability Score (CRPS)
Pinball Loss (a.k.a. Quantile Loss)
Novelty (in Recommender Systems)
Diversity (in Recommender Systems)
Hit Rate (HR)
NDCG (Normalized Discounted Cumulative Gain)
Mean Average Precision (MAP)
Expected Calibration Error (ECE)
Reliability Curves (also called Calibration Curves)
Brier Score
Calibration quality (Model Calibration)
Uplift
Preliminary Estimation for AR Models and the Yule–Walker Equations
Maximum Likelihood Estimation for ARMA Models (Gaussian MLE)
Diagnostics After Fitting a Time Series Model
Order Selection for Time Series Models
ARIMA Models: How Nonstationary Models Are Built from Stationary Ones
SARIMA Models: Seasonal ARIMA
Beyond One-Step Ahead Predictions
Exponential Smoothing Models
level: beginner (177)
annoy.Annoy legacy c-api with examples
annoy.Index python-api with examples
Index (cython) python-api benchmark with examples
Index (cython) python-api with examples
Approximate Nearest Neighbors with Annoy — A Hamlet Example
annoy.Index to NPY or CSV with examples
Mmap annoy.AnnoyIndex with examples
Precision annoy.AnnoyIndex with examples
Simple annoy.AnnoyIndex with examples
plot_calibration with examples
plot_classifier_eval with examples
plot_confusion_matrix with examples
plot_feature_importances with examples
plot_learning_curve with examples
plot_precision_recall with examples
plot_roc_curve with examples
plot_elbow with examples
plot_silhouette with examples
corpus A Tale of Two Cities .mp3 with examples
corpus Knowledge and Information local .png with examples
corpus WHO European Region local or url per file with examples
corpus WHO European Region YouTube shorts with examples
corpus WHO European Region local .zip with examples
Cython: Realtime compile_and_load (.pyx)
plot_cumulative_gain with examples
plot_ks_statistic with examples
plot_lift with examples
Introduction to modelplotpy (legacy)
Introduction to modelplotpy
plot_report with examples
plot_pca_2d_projection with examples
plot_pca_component_variance with examples
annoy impute with examples
Memory-Mapping Showcase – Basic / Medium / Advanced
Misc Showcase
MLflow
plot_aucplot_script with examples
plot_decileplot_script with examples
plot_evalplot_script with examples
Gaussian Mixture Models — AIC, AICc, and BIC Model Selection
visualkeras: Spam Dense example
visualkeras: autoencoder example
visualkeras: EfficientNetV2 example
visualkeras: ResNetV2 example
visualkeras: custom VGG example
visualkeras: Vector Index DB
The three steps of Bayesian data analysis
General Notation for Statistical Inference
Bayesian Inference
Discrete Bayesian Examples – Genetics and Spell Checking (with θ)
Probability as a Measure of Uncertainty
Example — Probabilities from Football Point Spreads
Example — Calibration for Record Linkage
Some Useful Results from Probability Theory
Computation and Software
Bayesian Inference in Applied Statistics
Estimating a Probability from Binomial Data
Posterior as a Compromise Between Data and Prior Information
Summarizing Posterior Inference
Informative Prior Distributions
Normal Distribution with Known Variance
Other Standard Single-Parameter Models
Informative Prior Distribution for Cancer Rates
Noninformative Prior Distributions
Weakly Informative Prior Distributions
Averaging Over Nuisance Parameters
Normal Data with a Noninformative Prior Distribution
Normal Data with a Conjugate Prior Distribution
Multinomial Model for Categorical Data
Multivariate Normal Model with Known Variance
Multivariate Normal with Unknown Mean and Variance
Example: Bayesian analysis of a bioassay experiment (logistic, nonconjugate)
Summary of Elementary Modeling and Computation
Normal Approximations to the Posterior Distribution
Large-Sample Theory
Counterexamples to large-sample (asymptotic) Bayesian theorems
Frequency Evaluations of Bayesian Inferences
Bayesian interpretations of other statistical methods
Constructing a Parameterized Prior Distribution
Exchangeability and hierarchical models
Bayesian analysis of conjugate hierarchical models
Normal model with exchangeable parameters
Example: parallel experiments in eight schools
Hierarchical modeling applied to a meta-analysis
Weakly Informative Priors for Variance Parameters
Why Do We Analyze Data?
The Process of Data Analysis
CRISP-DM for Data Science
Big Data: Definition, Characteristics, Evolution, and Business Impact
The First Step in Knowing Your Data
IEEE 754 Floating-Point Standard
Discovering Associations Through Data: From Everyday Patterns to Chicago Taxi Trips (September 2022)
Taxi Trips – 2022 dataset from the City of Chicago open data portal
Objective Selection of the Bin Width for a Time Histogram
Measuring Associations in Data
Measuring Associations Between Two Continuous Variables
Correlation Coefficients in Python (Pearson, Spearman, Kendall)
Karl Pearson
Harald Cramér
What Are Statistical Tests?
Eta Squared (η²): Effect Size in ANOVA
What is a Neural Network?
Supervised Learning and Neural Networks
Why Deep Learning is Taking Off
Geoffrey Hinton Interview
Binary Classification and Logistic Regression (Neural Network Basics)
Logistic Regression (Binary Classification Model)
Logistic Regression – Loss Function and Cost Function
Probability
Frequentist
Type I Error
Standard Error (SE)
True Mean (Population Mean)
Margin of Error (MoE)
Critical Value
Sample Standard Deviation
Sample Mean
Regression Coefficient
Proportion
True Population Parameter
Compromise Power Analysis
Post Hoc Power Analysis
A Priori Power Analysis
Statistical Significance
Z-Score
Two-Proportion Z-Test
Beta Distribution
Minimum Detectable Lift (MDL)
Trivial Effects
Sample size
Power (1 – β)
Significance Level (α)
Effect Size (δ)
Hypothesis Testing
P-Value (probability value)
Z-Test
T-Test
Statistically Significant
IID (Independent and Identically Distributed)
AI (Artificial Intelligence)
Machine Learning (ML)
Medical AI
LLMs (Large Language Models)
Population Proportion
Target Variable
Probability Density
Normal Distribution
Probability Mass
Probability Distribution
Cumulative Distribution Function (CDF)
Support Vector Machines (SVMs)
Confidence Level
Neural Networks
Logistic Regression
Classification Models
Likelihood
Correlation
Outlier
Regression Models
Median
Mean
Computer Vision (CV)
Natural Language Processing (NLP)
Chi-square (χ²) Test
Kolmogorov–Smirnov (KS) Test
Statistical Tests
Decision Trees
Linear Models
Statistical Power
Clopper–Pearson Interval
Wilson Score Interval
Confidence Intervals (CIs)
Power Analysis
What Are Time Series, and How Are They Used?
Getting Started with R
A Gentle Introduction to Stationarity
Weak and Strong Stationarity
level: intermediate (276)
plot_residuals_distribution with examples
plot_residuals_distribution with examples
Visualkeras: Spam Classification Conv1D Dense Example
visualkeras: custom vgg16 example
visualkeras: custom vgg16 show dimension example
The Place of Model Checking in Applied Bayesian Statistics
Do the Inferences from the Model Make Sense?
Posterior predictive checking
Graphical posterior predictive checks
Model checking for the educational testing example
Measures of predictive accuracy
Model comparison based on predictive performance
Model comparison using Bayes factors
Continuous model expansion
Implicit assumptions and model expansion: an example
Bayesian inference requires a model for data collection
Data-collection models and ignorability
Sample surveys
Designed experiments
Sensitivity and the role of randomization
Observational studies
Censoring and truncation
Bayesian decision theory in different contexts
Using regression predictions: survey incentives
Multistage decision making: medical screening
Hierarchical decision analysis for home radon
Personal vs. institutional decision analysis
Numerical integration
Distributional approximations
Direct simulation and rejection sampling
Importance sampling
How many simulation draws are needed?
Computing environments
Debugging Bayesian computing
Gibbs sampler
Metropolis and Metropolis-Hastings algorithms
Using Gibbs and Metropolis as building blocks
Inference and assessing convergence
Effective number of simulation draws
Example: hierarchical normal model
Efficient Gibbs samplers
Efficient Metropolis jumping rules
Further extensions to Gibbs and Metropolis
Hamiltonian Monte Carlo
Hamiltonian Monte Carlo for a hierarchical model
Stan: developing a computing environment
Finding posterior modes
Boundary-avoiding priors for modal summaries
Normal and related mixture approximations
Finding marginal posterior modes using EM
Conditional and marginal posterior approximations
Example: hierarchical normal model (continued)
Variational inference
Expectation propagation
Other approximations
Unknown normalizing factors
Understanding Market Baskets and Ideal Customers
What Can Association Rules Tell Us?
How Association Rules Are Discovered: Concepts, Scale, Measures, and the Apriori Approach
Apriori: Frequent Itemsets via the Apriori Algorithm
association_rules: Generating Association Rules from Frequent Itemsets (mlxtend)
Cross-Selling
Stratified Random Sampling
Linear Congruential Random Number Generator (LCG)
Partitioning Observations to Train Objective Models
Putting Similar Observations into Clusters
Clustering
Recency, Frequency, and Monetary Value (RFM)
RFM Analysis
Creating Segments of Observations for Business Reasons (RFM)
Least Squares Regression
Multiple Linear Regression
Feature Importance in Linear Regression
Forward Selection: Definition and Core Idea
Forward Selection and Model Interpretation in Linear Regression
Understanding Forward and Backward Stepwise Regression
How Shapley Values Work
Gradient Descent in Logistic Regression
Derivatives
More Derivative Examples
Computation Graph
Derivatives with a Computation Graph
Logistic Regression Gradient Descent
Gradient Descent on m Training Examples
Vectorization in Logistic Regression
More Vectorization Examples
Vectorizing Logistic Regression
Subsampling
Class Weighting
SMOTE (Synthetic Minority Over-sampling Technique)
Oversampling
NearMiss (Distance-based Undersampling)
Cluster-based undersampling
Random Undersampling
Micro AUROC
Multi-label Classification
Micro F1
Single-label Classification
Micro Recall
Micro Precision
One-vs-Rest (OvR) AUROC
Macro AUROC (Macro-Averaged AUROC)
Macro F1
Macro Recall
Macro Precision
Multiclass AUROC
Gini Coefficient
Bootstrap Confidence Intervals (CIs)
Mann–Whitney U Test (also called the Wilcoxon rank-sum test)
Cross-Selling
Upselling
Customer Segmentation
SaaS (Software as a Service)
Valuation Metric
D2C (Direct-to-Consumer)
LTV:CAC Ratio
Net LTV (sometimes called Contribution LTV)
Gross LTV (Customer Lifetime Value)
Predictive LTV (pLTV)
Cohort-Based LTV (Simple Version)
Customer Lifetime
Gross Margin
Fully Loaded CAC (Customer Acquisition Cost)
Organic CAC (Customer Acquisition Cost)
Paid CAC (Customer Acquisition Cost)
Channel-Specific CAC (Customer Acquisition Cost)
Blended CAC (Customer Acquisition Cost)
Lead-Gen Software
Conversion Rate Uplift
Bayesian Stopping Rules
Optimizely
Online Experimentation Platforms
Stopping Rules
Treatment Effect
Bayesian Sequential Testing
Likelihood Ratio (LR)
Group Sequential Testing
Traditional A/B Test (Fixed-Horizon A/B Test)
Fixed-Horizon Testing
True Conversion Rate
Google Experiments
A/B/n Test
Multivariate Test (MVT)
Risk of Peeking
Session Length
Revenue per User (RPU / ARPU)
Churn
Retention
Blocked Splits (Single Holdout)
Sliding Window (Rolling Window) Cross-Validation
Expanding Window Cross-Validation
Data Leakage
Stratified Group K-Fold
Stratified Shuffle Split
Multiclass stratified CV
k-fold cross-validation
Cross-Validation (CV)
Model Distillation (Knowledge Distillation)
Early Stopping
Epochs
Hyperparameter
KYC
FTEs
AWS SageMaker
Vertex AI
OpenAI API (ML API)
Ensemble
Model Weights
FLOPs
OpEx
Active Learning
Lagging Indicators
Leading Indicators
Cramér’s V
Cohort
Discriminatory Power
KS Statistic (Kolmogorov–Smirnov Statistic)
ROI (Return on Investment)
Supplier Constraints
Long Lead Times
Slow-Moving SKUs
SKU
Real-Time Inventory Tracking
Supplier Management
Demand Forecasting
Reorder Point (ROP) Optimization
Safety Stock
Backorder Rate
Lost Sales Value
Fill Rate
Stockout Rate
Classification Probability
Average Absolute Error (AAE)
Relative accuracy
R² (R-squared)
Long-Tail Items
Kaggle
ROC Curve (Receiver Operating Characteristic)
Binary Cross-Entropy (BCE)
Loss Functions
Underflow
Logit Space
Binary Classification
Log-Odds
Softmax Function
Sigmoid Function
Squashing Function
Conversion Rate (CR)
Cost-Per-Click (CPC) Models
Mean Squared Error (MSE)
One-vs-Rest (OvR)
Multiclass Classification
Partial AUC (pAUC)
Micro AUC
Macro AUC
Sensitivity in Feature Engineering
Encode (in Feature Engineering)
Normalize (in Feature Engineering)
Accuracy
Deep Ensembles
Quantization
ONNX (Open Neural Network Exchange)
Full Annotation
Weak Supervision
TPU Clusters
DeLong’s Test
Label Noise
Evaluation Set
Per-class Precision (sometimes called class-wise precision)
Multiclass Precision
Multilabel Precision
Weighted Averaging
Harmonic Mean
F1-score
Model Score
Bootstrap
Average Precision (AP)
Upsampling
Downsampling
Micro Averaging
Macro Averaging
AUC (Area Under the Curve)
LTV (Customer Lifetime Value)
CAC (Customer Acquisition Cost)
Sequential Testing (also called sequential analysis)
A/B Testing
Time-based splits (a.k.a. Temporal Cross-Validation, Rolling Window Validation)
k-fold Stratified Cross-Validation (Stratified CV)
Cannibalization
Crew Overtime
Overstock %
Stockouts
MASE (Mean Absolute Scaled Error)
WMAPE (Weighted Mean Absolute Percentage Error)
sMAPE (Symmetric Mean Absolute Percentage Error)
RMSLE (Root Mean Squared Logarithmic Error)
Mean Absolute Error (MAE)
Coverage
Log Loss (also called Logarithmic Loss or Cross-Entropy Loss)
Logits
CTR (Click-Through Rate)
WAPE (Weighted Absolute Percentage Error)
Recall
Mean Absolute Percentage Error (MAPE)
Root Mean Squared Error (RMSE)
ROC-AUC (Receiver Operating Characteristic – Area Under Curve, = AUROC)
Baseline Heuristics
Precision (a.k.a. Positive Predictive Value, PPV)
Precision–Recall AUC (PR-AUC)
Advanced Sorting in Spreadsheets
Linear Processes
Understanding ARMA Processes
Computing ACFs of Causal AR(2) Processes Using Difference Equations
Understanding ACFs via Difference Equations for AR(p) and ARMA(p, q)
Best Linear Predictor of a Stationary Process
Sample ACF and Sample PACF
model-type: classification (35)
plot_calibration with examples
plot_classifier_eval with examples
plot_confusion_matrix with examples
plot_feature_importances with examples
plot_learning_curve with examples
plot_precision_recall with examples
plot_roc_curve with examples
corpus A Tale of Two Cities .mp3 with examples
corpus Knowledge and Information local .png with examples
corpus WHO European Region local or url per file with examples
corpus WHO European Region YouTube shorts with examples
corpus WHO European Region local .zip with examples
Cython: Realtime compile_and_load (.pyx)
plot_cumulative_gain with examples
plot_ks_statistic with examples
plot_lift with examples
Introduction to modelplotpy (legacy)
Introduction to modelplotpy
plot_report with examples
plot_pca_2d_projection with examples
plot_pca_component_variance with examples
annoy impute with examples
MLflow
plot_aucplot_script with examples
plot_decileplot_script with examples
plot_evalplot_script with examples
Visualkeras: Spam Classification Conv1D Dense Example
visualkeras: Spam Dense example
visualkeras: autoencoder example
visualkeras: custom vgg16 example
visualkeras: custom vgg16 show dimension example
visualkeras: EfficientNetV2 example
visualkeras: ResNetV2 example
visualkeras: custom VGG example
visualkeras: Vector Index DB
model-type: clustering (3)
plot_elbow with examples
plot_silhouette with examples
Gaussian Mixture Models — AIC, AICc, and BIC Model Selection
model-type: regression (3)
plot_pca_component_variance with examples
plot_residuals_distribution with examples
plot_residuals_distribution with examples
model-workflow: corpus (5)
corpus A Tale of Two Cities .mp3 with examples
corpus Knowledge and Information local .png with examples
corpus WHO European Region local or url per file with examples
corpus WHO European Region YouTube shorts with examples
corpus WHO European Region local .zip with examples
model-workflow: feature engineering (2)
plot_pca_2d_projection with examples
plot_pca_component_variance with examples
model-workflow: impute (1)
annoy impute with examples
model-workflow: model building (11)
Cython: Realtime compile_and_load (.pyx)
MLflow
Visualkeras: Spam Classification Conv1D Dense Example
visualkeras: Spam Dense example
visualkeras: autoencoder example
visualkeras: custom vgg16 example
visualkeras: custom vgg16 show dimension example
visualkeras: EfficientNetV2 example
visualkeras: ResNetV2 example
visualkeras: custom VGG example
visualkeras: Vector Index DB
model-workflow: model evaluation (20)
plot_calibration with examples
plot_classifier_eval with examples
plot_confusion_matrix with examples
plot_feature_importances with examples
plot_learning_curve with examples
plot_precision_recall with examples
plot_roc_curve with examples
plot_elbow with examples
plot_silhouette with examples
plot_cumulative_gain with examples
plot_ks_statistic with examples
plot_lift with examples
Introduction to modelplotpy (legacy)
Introduction to modelplotpy
plot_report with examples
plot_residuals_distribution with examples
plot_aucplot_script with examples
plot_decileplot_script with examples
plot_evalplot_script with examples
plot_residuals_distribution with examples
model-workflow: model-selection (1)
Gaussian Mixture Models — AIC, AICc, and BIC Model Selection
model-workflow: vector-db (9)
annoy.Annoy legacy c-api with examples
annoy.Index python-api with examples
Index (cython) python-api benchmark with examples
Index (cython) python-api with examples
Approximate Nearest Neighbors with Annoy — A Hamlet Example
annoy.Index to NPY or CSV with examples
Mmap annoy.AnnoyIndex with examples
Precision annoy.AnnoyIndex with examples
Simple annoy.AnnoyIndex with examples
plot-type: Inertia (sum of squared distances) (1)
plot_elbow with examples
plot-type: PCA (1)
plot_pca_2d_projection with examples
plot-type: WSS (within-cluster sum of squares) (1)
plot_elbow with examples
plot-type: auc (4)
plot_learning_curve with examples
plot_precision_recall with examples
plot_roc_curve with examples
plot_aucplot_script with examples
plot-type: bar (8)
Mmap annoy.AnnoyIndex with examples
Precision annoy.AnnoyIndex with examples
Simple annoy.AnnoyIndex with examples
plot_feature_importances with examples
plot_silhouette with examples
corpus WHO European Region local or url per file with examples
annoy impute with examples
Memory-Mapping Showcase – Basic / Medium / Advanced
plot-type: barh (1)
Misc Showcase
plot-type: cython (11)
Cython quickstart: compile_and_load
Browse and compile templates
Build profiles: fast-debug, release, annotate
Cache and restart reuse
Pin/Alias: stable handles for cached builds
Multi-module package builds (5 package examples)
Multi-file builds: .pxi includes and external headers
C++ mode basics: cppclass and libcpp containers
Vector ops without NumPy: array(‘d’) + memoryviews
Workflow templates (train / hpo / predict) + CLI entry template
Cython: Realtime compile_and_load (.pyx)
plot-type: decile (7)
plot_cumulative_gain with examples
plot_ks_statistic with examples
plot_lift with examples
Introduction to modelplotpy (legacy)
Introduction to modelplotpy
plot_report with examples
plot_decileplot_script with examples
plot-type: density (1)
Gaussian Mixture Models — AIC, AICc, and BIC Model Selection
plot-type: eval (4)
plot_classifier_eval with examples
plot_confusion_matrix with examples
plot_feature_importances with examples
plot_evalplot_script with examples
plot-type: heatmap (2)
plot_classifier_eval with examples
plot_confusion_matrix with examples
plot-type: histogram (2)
plot_residuals_distribution with examples
plot_residuals_distribution with examples
plot-type: line (16)
plot_calibration with examples
plot_learning_curve with examples
plot_precision_recall with examples
plot_roc_curve with examples
plot_elbow with examples
plot_cumulative_gain with examples
plot_ks_statistic with examples
plot_lift with examples
Introduction to modelplotpy (legacy)
Introduction to modelplotpy
plot_report with examples
plot_pca_component_variance with examples
plot_aucplot_script with examples
plot_decileplot_script with examples
plot_evalplot_script with examples
Gaussian Mixture Models — AIC, AICc, and BIC Model Selection
plot-type: model (1)
plot_calibration with examples
plot-type: qqplot (2)
plot_residuals_distribution with examples
plot_residuals_distribution with examples
plot-type: scatter (1)
plot_pca_2d_projection with examples
plot-type: silhouette (1)
plot_silhouette with examples
plot-type: text (6)
corpus A Tale of Two Cities .mp3 with examples
corpus Knowledge and Information local .png with examples
corpus WHO European Region YouTube shorts with examples
corpus WHO European Region local .zip with examples
Misc Showcase
MLflow
plot-type: visualkeras (9)
Visualkeras: Spam Classification Conv1D Dense Example
visualkeras: Spam Dense example
visualkeras: autoencoder example
visualkeras: custom vgg16 example
visualkeras: custom vgg16 show dimension example
visualkeras: EfficientNetV2 example
visualkeras: ResNetV2 example
visualkeras: custom VGG example
visualkeras: Vector Index DB
purpose: reference (897)
Documentation Tagging Guidelines
The three steps of Bayesian data analysis
General Notation for Statistical Inference
Bayesian Inference
Discrete Bayesian Examples – Genetics and Spell Checking (with θ)
Probability as a Measure of Uncertainty
Example — Probabilities from Football Point Spreads
Example — Calibration for Record Linkage
Some Useful Results from Probability Theory
Computation and Software
Bayesian Inference in Applied Statistics
Estimating a Probability from Binomial Data
Posterior as a Compromise Between Data and Prior Information
Summarizing Posterior Inference
Informative Prior Distributions
Normal Distribution with Known Variance
Other Standard Single-Parameter Models
Informative Prior Distribution for Cancer Rates
Noninformative Prior Distributions
Weakly Informative Prior Distributions
Averaging Over Nuisance Parameters
Normal Data with a Noninformative Prior Distribution
Normal Data with a Conjugate Prior Distribution
Multinomial Model for Categorical Data
Multivariate Normal Model with Known Variance
Multivariate Normal with Unknown Mean and Variance
Example: Bayesian analysis of a bioassay experiment (logistic, nonconjugate)
Summary of Elementary Modeling and Computation
Normal Approximations to the Posterior Distribution
Large-Sample Theory
Counterexamples to large-sample (asymptotic) Bayesian theorems
Frequency Evaluations of Bayesian Inferences
Bayesian interpretations of other statistical methods
Constructing a Parameterized Prior Distribution
Exchangeability and hierarchical models
Bayesian analysis of conjugate hierarchical models
Normal model with exchangeable parameters
Example: parallel experiments in eight schools
Hierarchical modeling applied to a meta-analysis
Weakly Informative Priors for Variance Parameters
The Place of Model Checking in Applied Bayesian Statistics
Do the Inferences from the Model Make Sense?
Posterior predictive checking
Graphical posterior predictive checks
Model checking for the educational testing example
Measures of predictive accuracy
Model comparison based on predictive performance
Model comparison using Bayes factors
Continuous model expansion
Implicit assumptions and model expansion: an example
Bayesian inference requires a model for data collection
Data-collection models and ignorability
Sample surveys
Designed experiments
Sensitivity and the role of randomization
Observational studies
Censoring and truncation
Bayesian decision theory in different contexts
Using regression predictions: survey incentives
Multistage decision making: medical screening
Hierarchical decision analysis for home radon
Personal vs. institutional decision analysis
Numerical integration
Distributional approximations
Direct simulation and rejection sampling
Importance sampling
How many simulation draws are needed?
Computing environments
Debugging Bayesian computing
Gibbs sampler
Metropolis and Metropolis-Hastings algorithms
Using Gibbs and Metropolis as building blocks
Inference and assessing convergence
Effective number of simulation draws
Example: hierarchical normal model
Efficient Gibbs samplers
Efficient Metropolis jumping rules
Further extensions to Gibbs and Metropolis
Hamiltonian Monte Carlo
Hamiltonian Monte Carlo for a hierarchical model
Stan: developing a computing environment
Finding posterior modes
Boundary-avoiding priors for modal summaries
Normal and related mixture approximations
Finding marginal posterior modes using EM
Conditional and marginal posterior approximations
Example: hierarchical normal model (continued)
Variational inference
Expectation propagation
Other approximations
Unknown normalizing factors
Conditional modeling
Bayesian analysis of classical regression
Regression for causal inference: incumbency and voting
Goals of regression analysis
Assembling the matrix of explanatory variables
Regularization and dimension reduction
Unequal variances and correlations
Including numerical prior information
Regression coefficients exchangeable in batches
Example: forecasting U.S. presidential elections
Interpreting a normal prior distribution as extra data
Varying intercepts and slopes
Computation: batching and transformation
Analysis of variance and the batching of coefficients
Hierarchical models for batches of variance components
Standard generalized linear model likelihoods
Working with generalized linear models
Weakly informative priors for logistic regression
Overdispersed Poisson regression for police stops
State-level opinons from national polls
Models for multivariate and multinomial responses
Loglinear models for multivariate discrete data
Aspects of robustness
Overdispersed versions of standard models
Posterior inference and computation
Robust inference for the eight schools
Robust regression using t-distributed errors
Notation
Multiple imputation
Missing data in the multivariate normal and t models
Example: multiple imputation for a series of polls
Missing values with counted data
Example: an opinion poll in Slovenia
Example: serial dilution assay
Example: population toxicokinetics
Splines and weighted sums of basis functions
Basis selection and shrinkage of coefficients
Non-normal models and regression surfaces
Gaussian process regression
Example: birthdays and birthdates
Latent Gaussian process models
Functional data analysis
Density estimation and regression
Setting up and interpreting mixture models
Example: reaction times and schizophrenia
Label switching and posterior computation
Unspecified number of mixture components
Mixture models for classification and regression
Bayesian histograms
Dirichlet process prior distributions
Dirichlet process mixtures
Beyond density estimation
Hierarchical dependence
Density regression
Bayesian Data Analysis
Why Data Analytics Matters Today
How Data Analytics Improves the Workplace
Data-Driven Decision-Making
Detectives and Data Analysts
The Six Phases of the Data Analysis Process
The Origins of Data Analysis and the Many Ways to Structure It
Understanding the Data Ecosystem
Understanding the Data Analysis Process and the Data Life Cycle
Understanding the Data Life Cycle
A Review of the Six Stages of the Data Life Cycle
The Stages of the Data Analysis Process and Their Roles
Practical Application of the Data Analysis Process
Analytical Skills and Their Core Components
Applying Analytical Skills in a Business Context
Analytical Thinking and Its Core Components
Analytical Thinking and Questions for Problem Solving
Root Cause Analysis and Business Applications of the Five Whys
Data-Driven Decision-Making and the Role of Analytical Skills
Case Studies in Data Analysis and the Practical Impact of Data-Driven Decision-Making
Overview of Core Tools Used by Data Analysts
The Role of Spreadsheets in Data Analysis and Basic Concepts
The Concept and Basic Use of SQL (Query Language)
The Role and Importance of Data Visualization
Industries Where Data Analysts Work and How Data Is Used
The Role of Business Tasks in Data Analysis
Fairness in Data Analysis
Key Factors to Consider When Choosing a Data Analytics Role
🌱 Foundations
Using Data Analysis to Choose the Right Advertising Strategy
Understanding Common Problem Types in Data Analytics
Applying Data Analytics Problem Types in Real Business Scenarios
Why Asking the Right Questions Matters in Data Analytics
The Relationship Between Data and Decision-Making
Quantitative and Qualitative Data in Decision-Making
Data Creates Value Only When It Is Communicated
The Difference Between Data and Metrics, and the Role of Metrics
Dashboards
Mathematical Thinking
Spreadsheets in Data Analysis
Building and Organizing a Spreadsheet
How Data Analysts Use Spreadsheets
Spreadsheet Calculations with Formulas
Common Spreadsheet Errors and How to Fix Them
Spreadsheet Functions
Defining the Problem Domain
Context and Bias in Data Analysis
Stakeholder Expectations in Data Analysis
Staying Focused on the Project Objective
Clear Communication with Stakeholders and Teams
Adapting to Communication Expectations at Work
Managing Stakeholder Expectations and Project Constraints
Balancing Speed and Accuracy in Data Analysis
Sharing Data to Drive Impact
Effective Meetings
Conflict Resolution in the Workplace
🎯 Data-Driven Decisions
How Data Is Generated and Collected
Choosing the Right Data to Collect
Understanding Data Types and Data Formats
Structured Data and Data Models
Data Types in Spreadsheets
Data Tables (Tabular Data)
Wide Data vs. Long Data
Understanding Bias in Data Analysis
Sampling Bias and Unbiased Data
Common Types of Data Bias
Identifying Good Data Sources (ROCCC Framework)
Identifying Bad Data Sources (When Data Does Not ROCCC)
Data Ethics in Data Analysis
Data Privacy in Data Ethics
Open Data and Openness in Data Ethics
Databases and Relational Database Concepts
Metadata in Databases
Metadata Repositories and Data Governance
Accessing Data: Internal and External Sources
Importing Data into Spreadsheets
Sorting and Filtering Data in Spreadsheets
BigQuery Account Types
Querying Data with SQL
Organizing Data for Personal and Work Projects
Data Security in Spreadsheets
📦 Data Preparation
The Importance of Clean Data
Data Integrity and Its Risks in Data Analysis
Aligning Data with Business Objectives
Handling Insufficient Data in Data Analysis
Population, Sample Size, and Random Sampling
Statistical Power in Data Analysis
Sample Size and Data Integrity
Margin of Error
Dirty Data vs. Clean Data
The Importance of Clean Data (revisited)
Common Issues in Dirty Data
Data Cleaning with Spreadsheets
Cleaning and Merging Multiple Datasets
Spreadsheet Tools for Data Cleaning
Using Spreadsheet Functions for Data Cleaning
Viewing Data Differently for More Effective Data Cleaning
Data Mapping and the Big Picture of Clean Data
Introduction to SQL
Spreadsheets vs. SQL
Core SQL Queries for Data Cleaning and Analysis
Cleaning Data with SQL: Removing Duplicates and Cleaning String Variables
Using CAST to Clean and Format Data in SQL
Advanced SQL Functions for Data Cleaning
COALESCE
Verifying and Reporting Data Integrity
Verifying Data-Cleaning Efforts
Verification Techniques: Using Spreadsheets and SQL to Catch Repeated Errors
Documenting Data-Cleaning Changes
Reporting Data-Cleaning Results
Using Feedback from Data Cleaning to Improve Data Quality
Refining a Resume for Data Analytics Roles
Exploring Data Analyst Job Opportunities
🧽 Data Cleaning & Preparation
Understanding Data Analysis
Data Organization in Analysis
Sorting and Filtering in Data Analysis
Sorting Data in Spreadsheets
Sorting and Filtering Data in SQL Using ORDER BY and WHERE
Data Formatting and Unit Conversion in Spreadsheets
Data Validation in Spreadsheets
Combining Data Validation and Conditional Formatting in Spreadsheets
Using CONCAT in SQL to Combine Text from Multiple Columns
Working with Strings in Spreadsheets (LEN, LEFT, RIGHT, FIND)
Problem-Solving and Seeking Help in Data Analysis
How to Effectively Search for Solutions Online as a Data Analyst
Choosing the Right Tool in Data Analysis
Preparing Data for VLOOKUP in Spreadsheets
Using VLOOKUP to Combine Data Across Spreadsheets
Troubleshooting VLOOKUP and Building a Problem-Solving Framework
Using JOIN in SQL to Combine Tables
Subqueries in SQL
Aggregating Data with Subqueries, HAVING, and CASE in SQL
Using Spreadsheet Formulas for Sales Trend Analysis
Using COUNTIF and SUMIF for Conditional Aggregation in Spreadsheets
Using SUMPRODUCT for Advanced Spreadsheet Calculations
Using Pivot Tables for Calculations and Trend Analysis
Using Pivot Table Filters and Calculated Fields for Deeper Analysis
Comparing Calculations in Spreadsheets and SQL
Embedding Calculations in SQL Queries
Using GROUP BY and ORDER BY for Aggregated Calculations in SQL
Data Validation as an Ongoing Analytical Process
Temporary Tables and the WITH Clause in SQL
Creating Temporary Tables in SQL — Methods, Trade-offs, and Best Practices
📊 Analyze Data
Data Visualization
Connecting Data and Images
Creating Powerful Data Visualizations: Focus, Structure, and Analytical Purpose
Static vs. Dynamic Data Visualizations: Design Tradeoffs, Control, and Interactivity
Elements of Art in Data Visualization: Line, Shape, Color, Space, and Movement
Choosing the Right Visualization: Audience-Centered Design and Chart Selection
Design Thinking in Data Visualization: A User-Centered Framework
Accessibility in Data Visualization: Designing for Everyone
Introduction to Tableau
Getting Started with Tableau Public
Creating a CO₂ Emissions Visualization in Tableau Public
Effective vs. Ineffective Data Visualizations in Tableau
Using Creativity in Tableau
Linking Multiple Datasets in Tableau Public
Data Storytelling: Giving Numbers a Clear and Convincing Voice
Engaging Your Audience in Data Storytelling: Identifying the Key Message
Data Dashboards: Organizing Insight for Real-Time Decision Making
Using Filters to Create Compelling and Focused Visuals
Structuring a Persuasive Data Presentation: Turning Insights into Story
Designing Effective Data Presentation Slides: Structure, Visuals, and Professional Impact
Using a Strategic Framework to Structure Data Presentations
Weaving Data into Presentations: Hypotheses, Context, and the McCandless Method
Presentation Skills for Data Analysts: Delivering Insights with Confidence
Presenting Like a Pro: Best Practices for Data Analysts
Preparing for Q&A: Anticipating and Responding to Stakeholder Questions
Handling Objections in Data Presentations: Responding with Confidence and Clarity
Q&A Best Practices: Answering Questions with Clarity and Confidence
🎨 Data Visualization
Introduction to Python and Programming Fundamentals
Python Fundamentals
Jupyter Notebook and Coding Environments
Object-Oriented Programming (OOP) in Python
Variables in Python
Naming Conventions and Restrictions in Python
Data Types and Type Conversion in Python
Functions in Python
Code Reusability, Modularity, and Clean Code in Python
Comments, Algorithms, and Docstrings in Python
Boolean Data, Comparators, and Logical Operators in Python
Branching and Conditional Statements in Python
While Loops and Iteration in Python
For Loops in Python
range() Function and Loop Control in Python
Strings in Python
String Indexing and Slicing in Python
String Formatting with .format() in Python
Data Types vs Data Structures & Introduction to Lists
Modifying Lists in Python
Tuples in Python
Advanced Use of Loops, Lists, Tuples & List Comprehension
Dictionaries in Python
Advanced Dictionary Usage in Python
Sets in Python
Libraries, Packages, and Modules in Python
Introduction to NumPy and Vectorization
NumPy Arrays (ndarray) and Core Concepts
Introduction to Pandas (Data Analysis Library)
Pandas DataFrame & Series
Boolean Masking in Pandas
Grouping and Aggregation in Pandas (groupby, agg)
Combining Data in Pandas (concat and merge)
🐍 Data Analysis Using Python
Transferable Skills
Career Identity Statement
Career Dreamer (AI Tool for Career Exploration)
Job Search Plan (Using AI Tools)
Tailoring Your Resume
Using AI to Improve and Tailor Your Resume
Building a Professional Online Presence (Personal Brand)
Choosing the Right Job Platforms
Job Application Tracking (Using AI + Spreadsheets)
Networking for Job Search
Interview Preparation
STAR Method (Behavioral Interview)
Using AI (NotebookLM) for Interview Preparation
Practicing Interviews with AI (Gemini Live)
Post-Interview Strategy
💼 Job Search
📈 Data Analytics
Why Do We Analyze Data?
The Process of Data Analysis
CRISP-DM for Data Science
Big Data: Definition, Characteristics, Evolution, and Business Impact
The First Step in Knowing Your Data
IEEE 754 Floating-Point Standard
Discovering Associations Through Data: From Everyday Patterns to Chicago Taxi Trips (September 2022)
Taxi Trips – 2022 dataset from the City of Chicago open data portal
Objective Selection of the Bin Width for a Time Histogram
Measuring Associations in Data
Measuring Associations Between Two Continuous Variables
Correlation Coefficients in Python (Pearson, Spearman, Kendall)
Karl Pearson
Harald Cramér
What Are Statistical Tests?
Eta Squared (η²): Effect Size in ANOVA
Understanding Market Baskets and Ideal Customers
What Can Association Rules Tell Us?
How Association Rules Are Discovered: Concepts, Scale, Measures, and the Apriori Approach
Apriori: Frequent Itemsets via the Apriori Algorithm
association_rules: Generating Association Rules from Frequent Itemsets (mlxtend)
Cross-Selling
Stratified Random Sampling
Linear Congruential Random Number Generator (LCG)
Partitioning Observations to Train Objective Models
Putting Similar Observations into Clusters
Clustering
Recency, Frequency, and Monetary Value (RFM)
RFM Analysis
Creating Segments of Observations for Business Reasons (RFM)
Least Squares Regression
Multiple Linear Regression
Feature Importance in Linear Regression
Forward Selection: Definition and Core Idea
Forward Selection and Model Interpretation in Linear Regression
Understanding Forward and Backward Stepwise Regression
How Shapley Values Work
Logistic Regression: Modeling Binary Outcomes via Odds and Log-Odds
Maximum Likelihood (MLE): Fitting a Distribution to Observed Data
Assessing Model Fit in Logistic Regression
Complete and Quasi-Complete Separation in Logistic Regression
Forward Selection with Nested Models and Deviance Tests
Interpreting and Assessing a Forward-Selection Logistic Regression Model for College Student Retention
Motivation of Decision Trees: An Incremental Model of Decision-Making
The CART Algorithm
Decision Trees as Piecewise Models and Their Predictive Structure
How CART Decision Trees Model Interactions
Cluster Profiling Using Decision Trees
Using Decision Trees to Explain Clustering Results
Assessing the Quality of Prediction Models
Binary Classification Models – Conceptual Framework and Evaluation Metrics
Nominal Classification Models: Model State and Evaluation Metrics
Binary Classification Model Evaluation and Threshold Optimization
Identifying Outliers Using Residuals and Studentized Residuals
AUC–ROC Curve: Evaluating Classification Model Performance
Lift Analysis for Direct Mail Campaigns: Concept, Process, and Business Value
Data Preparation & Analysis
What is a Neural Network?
Supervised Learning and Neural Networks
Why Deep Learning is Taking Off
Geoffrey Hinton Interview
Binary Classification and Logistic Regression (Neural Network Basics)
Logistic Regression (Binary Classification Model)
Logistic Regression – Loss Function and Cost Function
Gradient Descent in Logistic Regression
Derivatives
More Derivative Examples
Computation Graph
Derivatives with a Computation Graph
Logistic Regression Gradient Descent
Gradient Descent on m Training Examples
Vectorization in Logistic Regression
More Vectorization Examples
Vectorizing Logistic Regression
Deep Learning
Subsampling
Class Weighting
SMOTE (Synthetic Minority Over-sampling Technique)
Oversampling
Low-pass Filtering
NearMiss (Distance-based Undersampling)
Cluster-based undersampling
Random Undersampling
Signal Processing
Time Series
Micro AUROC
Multi-label Classification
Micro F1
Single-label Classification
Micro Recall
Micro Precision
One-vs-Rest (OvR) AUROC
Macro AUROC (Macro-Averaged AUROC)
Macro F1
Macro Recall
Macro Precision
Multiclass AUROC
Gini Coefficient
Bootstrap Confidence Intervals (CIs)
Probability
Mann–Whitney U Test (also called the Wilcoxon rank-sum test)
Predictive Parity (Calibration)
Equalized Odds (Fairness)
Equal Opportunity (Fairness)
Demographic Parity (Statistical Parity)
Cross-Selling
Upselling
Customer Segmentation
SaaS (Software as a Service)
Valuation Metric
D2C (Direct-to-Consumer)
LTV:CAC Ratio
Net LTV (sometimes called Contribution LTV)
Gross LTV (Customer Lifetime Value)
Predictive LTV (pLTV)
Cohort-Based LTV (Simple Version)
Customer Lifetime
Gross Margin
Fully Loaded CAC (Customer Acquisition Cost)
Organic CAC (Customer Acquisition Cost)
Paid CAC (Customer Acquisition Cost)
Channel-Specific CAC (Customer Acquisition Cost)
Blended CAC (Customer Acquisition Cost)
Lead-Gen Software
Thompson Sampling (TS) in Bandits (Multi-Armed Bandit Problem (MAB))
Bayesian Decision Theory (BDT)
Bayesian Time Series
Posterior probability of uplift
Gaussian Processes (GPs)
Bayesian Neural Networks (BNNs)
Variational Inference (VI)
MCMC (Markov Chain Monte Carlo)
Sequential Settings
Frequentist
Binomial Likelihood
Posterior belief
Marginal Likelihood (also called The Model Evidence or Integrated Likelihood)
Posterior
Prior Belief (or Prior Probability)
Parameter(s) of Interest
Bayes’ Theorem
Conversion Rate Uplift
Bayesian Stopping Rules
Optimizely
Online Experimentation Platforms
Stopping Rules
Treatment Effect
Posterior Probability
Bayesian Sequential Testing
Likelihood Ratio (LR)
Sequential Probability Ratio Test (SPRT)
Pocock Method
O’Brien–Fleming (OBF) Method
Group Sequential Testing
Type I Error
Traditional A/B Test (Fixed-Horizon A/B Test)
Fixed-Horizon Testing
True Conversion Rate
Standard Error (SE)
True Mean (Population Mean)
Margin of Error (MoE)
Critical Value
Sample Standard Deviation
Sample Mean
Regression Coefficient
Proportion
True Population Parameter
Compromise Power Analysis
Post Hoc Power Analysis
A Priori Power Analysis
Statistical Significance
Z-Score
Two-Proportion Z-Test
Beta Distribution
Google Experiments
Minimum Detectable Lift (MDL)
Trivial Effects
Sample size
Power (1 – β)
Significance Level (α)
Effect Size (δ)
Hypothesis Testing
Ranking Algorithms
Probabilistic Interleaving
Team Draft Interleaving (TDI)
Balanced Interleaving
Causal Impact
Bandit Algorithms
A/B/n Test
Multivariate Test (MVT)
Risk of Peeking
Causal Inference
P-Value (probability value)
Z-Test
T-Test
Session Length
Revenue per User (RPU / ARPU)
Churn
Retention
Statistically Significant
IID (Independent and Identically Distributed)
Temporal autocorrelation (Serial Correlation)
Blocked Splits (Single Holdout)
Sliding Window (Rolling Window) Cross-Validation
Expanding Window Cross-Validation
Data Leakage
Stratified Group K-Fold
Stratified Shuffle Split
Multiclass stratified CV
k-fold cross-validation
Cross-Validation (CV)
Re-scoring
Drift Detection
Model Distillation (Knowledge Distillation)
Early Stopping
Epochs
Hyperparameter
AI (Artificial Intelligence)
Machine Learning (ML)
Medical AI
KYC
FTEs
AWS SageMaker
Vertex AI
OpenAI API (ML API)
AWS SageMaker Endpoints
Cloud Inference with Big Payloads
Cloud Inference
Ensemble
Model Weights
FLOPs
OpEx
LLMs (Large Language Models)
Recalibration
Reweighting
Continuous Retraining
Monitoring Pipelines
Active Learning
Bayesian Correction
Recalibrate Thresholds
Guardrails (in ML & Data Systems)
Model KPIs (Key Performance Indicators)
Lagging Indicators
Leading Indicators
Windows (in Time-Series)
Autoencoder
Frozen Encoder
Embedding
Representation Shift
Classifier Two-Sample Tests (C2STs)
Energy Distance
Maximum Mean Discrepancy (MMD)
Cardinality in Categorical Data
Categorical Drift
Cramér’s V
Macro Shifts
Categorical Explosions
Cohort
Off-Distribution
Discriminatory Power
KS Statistic (Kolmogorov–Smirnov Statistic)
Model Stability
Feature Values
Four-Fifths (80%) Rule
SLI (Service Level Indicator)
ROI (Return on Investment)
Treatment Cost
Incremental Revenue
Incremental Recovery Rate (IRR)
Incremental Sales
Random Targeting Strategy
Causal ML (Causal Machine Learning)
Cumulative Uplift
Population Proportion
Incremental Gain
Total Incremental Benefit (TIB)
Cumulative Incremental Gain (CIG)
Qini Curve
Uplift Score
Uplift Models
Ops Health Dashboard
SLA Breach Rate
SLA (Service Level Agreement)
Supplier Constraints
Long Lead Times
Slow-Moving SKUs
SKU
Real-Time Inventory Tracking
Supplier Management
Demand Forecasting
Reorder Point (ROP) Optimization
Safety Stock
Backorder Rate
Lost Sales Value
Fill Rate
Stockout Rate
Prophet — Time Series Forecasting by Facebook (Meta)
LSTM — Long Short-Term Memory Networks
ARIMA (AutoRegressive Integrated Moving Average)
Return Distribution
Value-at-Risk (VaR)
Risk Forecast
Probabilistic Scoring
Full Distribution
Continuous Probabilistic Forecasts
Classification Probability
Quantile Forecasts
Point Forecasts
Strictly Proper Scoring Rules
Probability Forecasts
Target Variable
Probability Density
Normal Distribution
Probability Mass
Probability Distribution
Probabilistic Forecasts
Deterministic forecasts
Cumulative Distribution Function (CDF)
M-Competitions (Makridakis Competitions)
Forecasting Benchmarks
Average Absolute Error (AAE)
Seasonal Lag
Simple Baseline Methods
Naïve Baseline Forecast
Forecast Error
Forecasting Competitions
Predicting Percentiles
Prediction Intervals (PI)
Quantile Regression
Quantile Level
Time Series Forecasting
Log-Space
Relative accuracy
R² (R-squared)
Long-Tail Items
Self-Information of Popularity
Relevance in Recommender Systems
Genre Overlap
Jaccard index
Cosine Similarity of Item Features
Intra-List Diversity (ILD)
Dominating in Recommender Systems
Catalog Coverage
User Coverage
Item Coverage
Diminishing Utility
DCG (Discounted Cumulative Gain)
Kaggle
TREC (Text REtrieval Conference)
Adaptive ECE (Expected Calibration Error with Adaptive Binning)
Maximum Calibration Error (MCE)
ROC Curve (Receiver Operating Characteristic)
Murphy’s Decomposition
Temperature Scaling
Platt Scaling
Isotonic Regression
Support Vector Machines (SVMs)
Underconfident
Overconfident
Confidence Level
Risk-Based Decisions
Neural Networks
Binary Cross-Entropy (BCE)
Loss Functions
Underflow
Logit Space
Logistic Regression
Binary Classification
Classification Models
Log-Odds
Softmax Function
Sigmoid Function
Squashing Function
Conversion Rate (CR)
Cost-Per-Click (CPC) Models
Causal Trees
Uplift Random Forests
Uplift Curve
Likelihood
Correlation
Causal Effect
Outlier
Mean Squared Error (MSE)
Regression Models
One-vs-Rest (OvR)
Multiclass Classification
Partial AUC (pAUC)
Micro AUC
Macro AUC
Median
Mean
Sensitivity in Feature Engineering
Encode (in Feature Engineering)
Normalize (in Feature Engineering)
Embedding Similarity
Computer Vision (CV)
Natural Language Processing (NLP)
Accuracy
Chi-square (χ²) Test
Kolmogorov–Smirnov (KS) Test
Jensen–Shannon (JS) Divergence
Kullback–Leibler (KL) Divergence
Statistical Tests
Seasonality
Concept Drift
Data Drift
Fair Lending laws
Basel III
High-Stakes Domains
Deep Ensembles
Counterfactual Explanations
LIME (Local Interpretable Model-agnostic Explanations)
SHAP (SHapley Additive exPlanations)
Post-hoc Explainability
Decision Trees
Linear Models
Caching
Quantization
ONNX (Open Neural Network Exchange)
Full Annotation
Weak Supervision
TPU Clusters
Statistical Power
Drift Guardrails
Latency Guardrails
Fairness Guardrails
DeLong’s Test
Dataset Shift
Label Noise
Evaluation Set
Clopper–Pearson Interval
Wilson Score Interval
Per-class Precision (sometimes called class-wise precision)
Multiclass Precision
Multilabel Precision
Weighted Averaging
Harmonic Mean
F1-score
Model Score
Bootstrap
Average Precision (AP)
Upsampling
Downsampling
Micro Averaging
Macro Averaging
AUC (Area Under the Curve)
Fairness parity
LTV (Customer Lifetime Value)
CAC (Customer Acquisition Cost)
Bayesian Inference.
Sequential Testing (also called sequential analysis)
Confidence Intervals (CIs)
Power Analysis
Interleaving Tests
A/B Testing
Time-based splits (a.k.a. Temporal Cross-Validation, Rolling Window Validation)
k-fold Stratified Cross-Validation (Stratified CV)
Compute budgets
Manual review minutes
Inference Cost (Inference $)
Label Drift (a.k.a. Target Drift)
Covariate Drift (a.k.a. Covariate Shift)
KS shift (Kolmogorov–Smirnov shift)
PSI (Population Stability Index)
Selection Rate
SLOs (Service Level Objectives)
Cannibalization
Revenue net of treatment cost
Incremental Conversions
Uplift@k
AUUC (Area Under the Uplift Curve)
Qini Coefficient
Crew Overtime
SLA Breaches
Overstock %
Stockouts
Continuous Ranked Probability Score (CRPS)
MASE (Mean Absolute Scaled Error)
Pinball Loss (a.k.a. Quantile Loss)
WMAPE (Weighted Mean Absolute Percentage Error)
sMAPE (Symmetric Mean Absolute Percentage Error)
RMSLE (Root Mean Squared Logarithmic Error)
Mean Absolute Error (MAE)
Novelty (in Recommender Systems)
Diversity (in Recommender Systems)
Coverage
Hit Rate (HR)
NDCG (Normalized Discounted Cumulative Gain)
Mean Average Precision (MAP)
Expected Calibration Error (ECE)
Reliability Curves (also called Calibration Curves)
Log Loss (also called Logarithmic Loss or Cross-Entropy Loss)
Brier Score
Calibration quality (Model Calibration)
Logits
CTR (Click-Through Rate)
WAPE (Weighted Absolute Percentage Error)
Recall
Uplift
Mean Absolute Percentage Error (MAPE)
Root Mean Squared Error (RMSE)
ROC-AUC (Receiver Operating Characteristic – Area Under Curve, = AUROC)
Baseline Heuristics
Precision (a.k.a. Positive Predictive Value, PPV)
Precision–Recall AUC (PR-AUC)
Advanced Sorting in Spreadsheets
Terminology
What Are Time Series, and How Are They Used?
Getting Started with R
A Gentle Introduction to Stationarity
Weak and Strong Stationarity
Linear Processes
Understanding ARMA Processes
Computing ACFs of Causal AR(2) Processes Using Difference Equations
Understanding ACFs via Difference Equations for AR(p) and ARMA(p, q)
Best Linear Predictor of a Stationary Process
Sample ACF and Sample PACF
Preliminary Estimation for AR Models and the Yule–Walker Equations
Maximum Likelihood Estimation for ARMA Models (Gaussian MLE)
Diagnostics After Fitting a Time Series Model
Order Selection for Time Series Models
ARIMA Models: How Nonstationary Models Are Built from Stationary Ones
SARIMA Models: Seasonal ARIMA
Beyond One-Step Ahead Predictions
Exponential Smoothing Models
Time Series
purpose: showcase (61)
annoy.Annoy legacy c-api with examples
annoy.Index python-api with examples
Index (cython) python-api benchmark with examples
Index (cython) python-api with examples
Approximate Nearest Neighbors with Annoy — A Hamlet Example
annoy.Index to NPY or CSV with examples
Mmap annoy.AnnoyIndex with examples
Precision annoy.AnnoyIndex with examples
Simple annoy.AnnoyIndex with examples
plot_calibration with examples
plot_classifier_eval with examples
plot_confusion_matrix with examples
plot_feature_importances with examples
plot_learning_curve with examples
plot_precision_recall with examples
plot_roc_curve with examples
plot_elbow with examples
plot_silhouette with examples
corpus A Tale of Two Cities .mp3 with examples
corpus Knowledge and Information local .png with examples
corpus WHO European Region local or url per file with examples
corpus WHO European Region YouTube shorts with examples
corpus WHO European Region local .zip with examples
Cython quickstart: compile_and_load
Browse and compile templates
Build profiles: fast-debug, release, annotate
Cache and restart reuse
Pin/Alias: stable handles for cached builds
Multi-module package builds (5 package examples)
Multi-file builds: .pxi includes and external headers
C++ mode basics: cppclass and libcpp containers
Vector ops without NumPy: array(‘d’) + memoryviews
Workflow templates (train / hpo / predict) + CLI entry template
Cython: Realtime compile_and_load (.pyx)
plot_cumulative_gain with examples
plot_ks_statistic with examples
plot_lift with examples
Introduction to modelplotpy (legacy)
Introduction to modelplotpy
plot_report with examples
plot_pca_2d_projection with examples
plot_pca_component_variance with examples
annoy impute with examples
Memory-Mapping Showcase – Basic / Medium / Advanced
Misc Showcase
MLflow
plot_residuals_distribution with examples
plot_aucplot_script with examples
plot_decileplot_script with examples
plot_evalplot_script with examples
Gaussian Mixture Models — AIC, AICc, and BIC Model Selection
plot_residuals_distribution with examples
Visualkeras: Spam Classification Conv1D Dense Example
visualkeras: Spam Dense example
visualkeras: autoencoder example
visualkeras: custom vgg16 example
visualkeras: custom vgg16 show dimension example
visualkeras: EfficientNetV2 example
visualkeras: ResNetV2 example
visualkeras: custom VGG example
visualkeras: Vector Index DB
topic: advanced (3)
Data Validation as an Ongoing Analytical Process
Temporary Tables and the WITH Clause in SQL
Creating Temporary Tables in SQL — Methods, Trade-offs, and Best Practices
topic: analyze (30)
Understanding Data Analysis
Data Organization in Analysis
Sorting and Filtering in Data Analysis
Sorting Data in Spreadsheets
Sorting and Filtering Data in SQL Using ORDER BY and WHERE
Data Formatting and Unit Conversion in Spreadsheets
Data Validation in Spreadsheets
Combining Data Validation and Conditional Formatting in Spreadsheets
Using CONCAT in SQL to Combine Text from Multiple Columns
Working with Strings in Spreadsheets (LEN, LEFT, RIGHT, FIND)
Problem-Solving and Seeking Help in Data Analysis
How to Effectively Search for Solutions Online as a Data Analyst
Choosing the Right Tool in Data Analysis
Preparing Data for VLOOKUP in Spreadsheets
Using VLOOKUP to Combine Data Across Spreadsheets
Troubleshooting VLOOKUP and Building a Problem-Solving Framework
Using JOIN in SQL to Combine Tables
Subqueries in SQL
Aggregating Data with Subqueries, HAVING, and CASE in SQL
Using Spreadsheet Formulas for Sales Trend Analysis
Using COUNTIF and SUMIF for Conditional Aggregation in Spreadsheets
Using SUMPRODUCT for Advanced Spreadsheet Calculations
Using Pivot Tables for Calculations and Trend Analysis
Using Pivot Table Filters and Calculated Fields for Deeper Analysis
Comparing Calculations in Spreadsheets and SQL
Embedding Calculations in SQL Queries
Using GROUP BY and ORDER BY for Aggregated Calculations in SQL
Data Validation as an Ongoing Analytical Process
Temporary Tables and the WITH Clause in SQL
Creating Temporary Tables in SQL — Methods, Trade-offs, and Best Practices
topic: apply (6)
Tailoring Your Resume
Using AI to Improve and Tailor Your Resume
Building a Professional Online Presence (Personal Brand)
Choosing the Right Job Platforms
Job Application Tracking (Using AI + Spreadsheets)
Networking for Job Search
topic: basics (10)
Introduction to Python and Programming Fundamentals
Python Fundamentals
Jupyter Notebook and Coding Environments
Object-Oriented Programming (OOP) in Python
Variables in Python
Naming Conventions and Restrictions in Python
Data Types and Type Conversion in Python
Functions in Python
Code Reusability, Modularity, and Clean Code in Python
Comments, Algorithms, and Docstrings in Python
topic: bias_ethics (8)
Understanding Bias in Data Analysis
Sampling Bias and Unbiased Data
Common Types of Data Bias
Identifying Good Data Sources (ROCCC Framework)
Identifying Bad Data Sources (When Data Does Not ROCCC)
Data Ethics in Data Analysis
Data Privacy in Data Ethics
Open Data and Openness in Data Ethics
topic: calc (8)
Using Spreadsheet Formulas for Sales Trend Analysis
Using COUNTIF and SUMIF for Conditional Aggregation in Spreadsheets
Using SUMPRODUCT for Advanced Spreadsheet Calculations
Using Pivot Tables for Calculations and Trend Analysis
Using Pivot Table Filters and Calculated Fields for Deeper Analysis
Comparing Calculations in Spreadsheets and SQL
Embedding Calculations in SQL Queries
Using GROUP BY and ORDER BY for Aggregated Calculations in SQL
topic: cleaning (32)
The Importance of Clean Data
Data Integrity and Its Risks in Data Analysis
Aligning Data with Business Objectives
Handling Insufficient Data in Data Analysis
Population, Sample Size, and Random Sampling
Statistical Power in Data Analysis
Sample Size and Data Integrity
Margin of Error
Dirty Data vs. Clean Data
The Importance of Clean Data (revisited)
Common Issues in Dirty Data
Data Cleaning with Spreadsheets
Cleaning and Merging Multiple Datasets
Spreadsheet Tools for Data Cleaning
Using Spreadsheet Functions for Data Cleaning
Viewing Data Differently for More Effective Data Cleaning
Data Mapping and the Big Picture of Clean Data
Introduction to SQL
Spreadsheets vs. SQL
Core SQL Queries for Data Cleaning and Analysis
Cleaning Data with SQL: Removing Duplicates and Cleaning String Variables
Using CAST to Clean and Format Data in SQL
Advanced SQL Functions for Data Cleaning
COALESCE
Verifying and Reporting Data Integrity
Verifying Data-Cleaning Efforts
Verification Techniques: Using Spreadsheets and SQL to Catch Repeated Errors
Documenting Data-Cleaning Changes
Reporting Data-Cleaning Results
Using Feedback from Data Cleaning to Improve Data Quality
Refining a Resume for Data Analytics Roles
Exploring Data Analyst Job Opportunities
topic: combine (9)
Problem-Solving and Seeking Help in Data Analysis
How to Effectively Search for Solutions Online as a Data Analyst
Choosing the Right Tool in Data Analysis
Preparing Data for VLOOKUP in Spreadsheets
Using VLOOKUP to Combine Data Across Spreadsheets
Troubleshooting VLOOKUP and Building a Problem-Solving Framework
Using JOIN in SQL to Combine Tables
Subqueries in SQL
Aggregating Data with Subqueries, HAVING, and CASE in SQL
topic: control (5)
Boolean Data, Comparators, and Logical Operators in Python
Branching and Conditional Statements in Python
While Loops and Iteration in Python
For Loops in Python
range() Function and Loop Control in Python
topic: data analysis (202)
The three steps of Bayesian data analysis
General Notation for Statistical Inference
Bayesian Inference
Discrete Bayesian Examples – Genetics and Spell Checking (with θ)
Probability as a Measure of Uncertainty
Example — Probabilities from Football Point Spreads
Example — Calibration for Record Linkage
Some Useful Results from Probability Theory
Computation and Software
Bayesian Inference in Applied Statistics
Estimating a Probability from Binomial Data
Posterior as a Compromise Between Data and Prior Information
Summarizing Posterior Inference
Informative Prior Distributions
Normal Distribution with Known Variance
Other Standard Single-Parameter Models
Informative Prior Distribution for Cancer Rates
Noninformative Prior Distributions
Weakly Informative Prior Distributions
Averaging Over Nuisance Parameters
Normal Data with a Noninformative Prior Distribution
Normal Data with a Conjugate Prior Distribution
Multinomial Model for Categorical Data
Multivariate Normal Model with Known Variance
Multivariate Normal with Unknown Mean and Variance
Example: Bayesian analysis of a bioassay experiment (logistic, nonconjugate)
Summary of Elementary Modeling and Computation
Normal Approximations to the Posterior Distribution
Large-Sample Theory
Counterexamples to large-sample (asymptotic) Bayesian theorems
Frequency Evaluations of Bayesian Inferences
Bayesian interpretations of other statistical methods
Constructing a Parameterized Prior Distribution
Exchangeability and hierarchical models
Bayesian analysis of conjugate hierarchical models
Normal model with exchangeable parameters
Example: parallel experiments in eight schools
Hierarchical modeling applied to a meta-analysis
Weakly Informative Priors for Variance Parameters
The Place of Model Checking in Applied Bayesian Statistics
Do the Inferences from the Model Make Sense?
Posterior predictive checking
Graphical posterior predictive checks
Model checking for the educational testing example
Measures of predictive accuracy
Model comparison based on predictive performance
Model comparison using Bayes factors
Continuous model expansion
Implicit assumptions and model expansion: an example
Bayesian inference requires a model for data collection
Data-collection models and ignorability
Sample surveys
Designed experiments
Sensitivity and the role of randomization
Observational studies
Censoring and truncation
Bayesian decision theory in different contexts
Using regression predictions: survey incentives
Multistage decision making: medical screening
Hierarchical decision analysis for home radon
Personal vs. institutional decision analysis
Numerical integration
Distributional approximations
Direct simulation and rejection sampling
Importance sampling
How many simulation draws are needed?
Computing environments
Debugging Bayesian computing
Gibbs sampler
Metropolis and Metropolis-Hastings algorithms
Using Gibbs and Metropolis as building blocks
Inference and assessing convergence
Effective number of simulation draws
Example: hierarchical normal model
Efficient Gibbs samplers
Efficient Metropolis jumping rules
Further extensions to Gibbs and Metropolis
Hamiltonian Monte Carlo
Hamiltonian Monte Carlo for a hierarchical model
Stan: developing a computing environment
Finding posterior modes
Boundary-avoiding priors for modal summaries
Normal and related mixture approximations
Finding marginal posterior modes using EM
Conditional and marginal posterior approximations
Example: hierarchical normal model (continued)
Variational inference
Expectation propagation
Other approximations
Unknown normalizing factors
Conditional modeling
Bayesian analysis of classical regression
Regression for causal inference: incumbency and voting
Goals of regression analysis
Assembling the matrix of explanatory variables
Regularization and dimension reduction
Unequal variances and correlations
Including numerical prior information
Regression coefficients exchangeable in batches
Example: forecasting U.S. presidential elections
Interpreting a normal prior distribution as extra data
Varying intercepts and slopes
Computation: batching and transformation
Analysis of variance and the batching of coefficients
Hierarchical models for batches of variance components
Standard generalized linear model likelihoods
Working with generalized linear models
Weakly informative priors for logistic regression
Overdispersed Poisson regression for police stops
State-level opinons from national polls
Models for multivariate and multinomial responses
Loglinear models for multivariate discrete data
Aspects of robustness
Overdispersed versions of standard models
Posterior inference and computation
Robust inference for the eight schools
Robust regression using t-distributed errors
Notation
Multiple imputation
Missing data in the multivariate normal and t models
Example: multiple imputation for a series of polls
Missing values with counted data
Example: an opinion poll in Slovenia
Example: serial dilution assay
Example: population toxicokinetics
Splines and weighted sums of basis functions
Basis selection and shrinkage of coefficients
Non-normal models and regression surfaces
Gaussian process regression
Example: birthdays and birthdates
Latent Gaussian process models
Functional data analysis
Density estimation and regression
Setting up and interpreting mixture models
Example: reaction times and schizophrenia
Label switching and posterior computation
Unspecified number of mixture components
Mixture models for classification and regression
Bayesian histograms
Dirichlet process prior distributions
Dirichlet process mixtures
Beyond density estimation
Hierarchical dependence
Density regression
Bayesian Data Analysis
Why Do We Analyze Data?
The Process of Data Analysis
CRISP-DM for Data Science
Big Data: Definition, Characteristics, Evolution, and Business Impact
The First Step in Knowing Your Data
IEEE 754 Floating-Point Standard
Discovering Associations Through Data: From Everyday Patterns to Chicago Taxi Trips (September 2022)
Taxi Trips – 2022 dataset from the City of Chicago open data portal
Objective Selection of the Bin Width for a Time Histogram
Measuring Associations in Data
Measuring Associations Between Two Continuous Variables
Correlation Coefficients in Python (Pearson, Spearman, Kendall)
Karl Pearson
Harald Cramér
What Are Statistical Tests?
Eta Squared (η²): Effect Size in ANOVA
Understanding Market Baskets and Ideal Customers
What Can Association Rules Tell Us?
How Association Rules Are Discovered: Concepts, Scale, Measures, and the Apriori Approach
Apriori: Frequent Itemsets via the Apriori Algorithm
association_rules: Generating Association Rules from Frequent Itemsets (mlxtend)
Cross-Selling
Stratified Random Sampling
Linear Congruential Random Number Generator (LCG)
Partitioning Observations to Train Objective Models
Putting Similar Observations into Clusters
Clustering
Recency, Frequency, and Monetary Value (RFM)
RFM Analysis
Creating Segments of Observations for Business Reasons (RFM)
Least Squares Regression
Multiple Linear Regression
Feature Importance in Linear Regression
Forward Selection: Definition and Core Idea
Forward Selection and Model Interpretation in Linear Regression
Understanding Forward and Backward Stepwise Regression
How Shapley Values Work
Logistic Regression: Modeling Binary Outcomes via Odds and Log-Odds
Maximum Likelihood (MLE): Fitting a Distribution to Observed Data
Assessing Model Fit in Logistic Regression
Complete and Quasi-Complete Separation in Logistic Regression
Forward Selection with Nested Models and Deviance Tests
Interpreting and Assessing a Forward-Selection Logistic Regression Model for College Student Retention
Motivation of Decision Trees: An Incremental Model of Decision-Making
The CART Algorithm
Decision Trees as Piecewise Models and Their Predictive Structure
How CART Decision Trees Model Interactions
Cluster Profiling Using Decision Trees
Using Decision Trees to Explain Clustering Results
Assessing the Quality of Prediction Models
Binary Classification Models – Conceptual Framework and Evaluation Metrics
Nominal Classification Models: Model State and Evaluation Metrics
Binary Classification Model Evaluation and Threshold Optimization
Identifying Outliers Using Residuals and Studentized Residuals
AUC–ROC Curve: Evaluating Classification Model Performance
Lift Analysis for Direct Mail Campaigns: Concept, Process, and Business Value
Data Preparation & Analysis
topic: data analytics (225)
Why Data Analytics Matters Today
How Data Analytics Improves the Workplace
Data-Driven Decision-Making
Detectives and Data Analysts
The Six Phases of the Data Analysis Process
The Origins of Data Analysis and the Many Ways to Structure It
Understanding the Data Ecosystem
Understanding the Data Analysis Process and the Data Life Cycle
Understanding the Data Life Cycle
A Review of the Six Stages of the Data Life Cycle
The Stages of the Data Analysis Process and Their Roles
Practical Application of the Data Analysis Process
Analytical Skills and Their Core Components
Applying Analytical Skills in a Business Context
Analytical Thinking and Its Core Components
Analytical Thinking and Questions for Problem Solving
Root Cause Analysis and Business Applications of the Five Whys
Data-Driven Decision-Making and the Role of Analytical Skills
Case Studies in Data Analysis and the Practical Impact of Data-Driven Decision-Making
Overview of Core Tools Used by Data Analysts
The Role of Spreadsheets in Data Analysis and Basic Concepts
The Concept and Basic Use of SQL (Query Language)
The Role and Importance of Data Visualization
Industries Where Data Analysts Work and How Data Is Used
The Role of Business Tasks in Data Analysis
Fairness in Data Analysis
Key Factors to Consider When Choosing a Data Analytics Role
🌱 Foundations
Using Data Analysis to Choose the Right Advertising Strategy
Understanding Common Problem Types in Data Analytics
Applying Data Analytics Problem Types in Real Business Scenarios
Why Asking the Right Questions Matters in Data Analytics
The Relationship Between Data and Decision-Making
Quantitative and Qualitative Data in Decision-Making
Data Creates Value Only When It Is Communicated
The Difference Between Data and Metrics, and the Role of Metrics
Dashboards
Mathematical Thinking
Spreadsheets in Data Analysis
Building and Organizing a Spreadsheet
How Data Analysts Use Spreadsheets
Spreadsheet Calculations with Formulas
Common Spreadsheet Errors and How to Fix Them
Spreadsheet Functions
Defining the Problem Domain
Context and Bias in Data Analysis
Stakeholder Expectations in Data Analysis
Staying Focused on the Project Objective
Clear Communication with Stakeholders and Teams
Adapting to Communication Expectations at Work
Managing Stakeholder Expectations and Project Constraints
Balancing Speed and Accuracy in Data Analysis
Sharing Data to Drive Impact
Effective Meetings
Conflict Resolution in the Workplace
🎯 Data-Driven Decisions
How Data Is Generated and Collected
Choosing the Right Data to Collect
Understanding Data Types and Data Formats
Structured Data and Data Models
Data Types in Spreadsheets
Data Tables (Tabular Data)
Wide Data vs. Long Data
Understanding Bias in Data Analysis
Sampling Bias and Unbiased Data
Common Types of Data Bias
Identifying Good Data Sources (ROCCC Framework)
Identifying Bad Data Sources (When Data Does Not ROCCC)
Data Ethics in Data Analysis
Data Privacy in Data Ethics
Open Data and Openness in Data Ethics
Databases and Relational Database Concepts
Metadata in Databases
Metadata Repositories and Data Governance
Accessing Data: Internal and External Sources
Importing Data into Spreadsheets
Sorting and Filtering Data in Spreadsheets
BigQuery Account Types
Querying Data with SQL
Organizing Data for Personal and Work Projects
Data Security in Spreadsheets
📦 Data Preparation
The Importance of Clean Data
Data Integrity and Its Risks in Data Analysis
Aligning Data with Business Objectives
Handling Insufficient Data in Data Analysis
Population, Sample Size, and Random Sampling
Statistical Power in Data Analysis
Sample Size and Data Integrity
Margin of Error
Dirty Data vs. Clean Data
The Importance of Clean Data (revisited)
Common Issues in Dirty Data
Data Cleaning with Spreadsheets
Cleaning and Merging Multiple Datasets
Spreadsheet Tools for Data Cleaning
Using Spreadsheet Functions for Data Cleaning
Viewing Data Differently for More Effective Data Cleaning
Data Mapping and the Big Picture of Clean Data
Introduction to SQL
Spreadsheets vs. SQL
Core SQL Queries for Data Cleaning and Analysis
Cleaning Data with SQL: Removing Duplicates and Cleaning String Variables
Using CAST to Clean and Format Data in SQL
Advanced SQL Functions for Data Cleaning
COALESCE
Verifying and Reporting Data Integrity
Verifying Data-Cleaning Efforts
Verification Techniques: Using Spreadsheets and SQL to Catch Repeated Errors
Documenting Data-Cleaning Changes
Reporting Data-Cleaning Results
Using Feedback from Data Cleaning to Improve Data Quality
Refining a Resume for Data Analytics Roles
Exploring Data Analyst Job Opportunities
🧽 Data Cleaning & Preparation
Understanding Data Analysis
Data Organization in Analysis
Sorting and Filtering in Data Analysis
Sorting Data in Spreadsheets
Sorting and Filtering Data in SQL Using ORDER BY and WHERE
Data Formatting and Unit Conversion in Spreadsheets
Data Validation in Spreadsheets
Combining Data Validation and Conditional Formatting in Spreadsheets
Using CONCAT in SQL to Combine Text from Multiple Columns
Working with Strings in Spreadsheets (LEN, LEFT, RIGHT, FIND)
Problem-Solving and Seeking Help in Data Analysis
How to Effectively Search for Solutions Online as a Data Analyst
Choosing the Right Tool in Data Analysis
Preparing Data for VLOOKUP in Spreadsheets
Using VLOOKUP to Combine Data Across Spreadsheets
Troubleshooting VLOOKUP and Building a Problem-Solving Framework
Using JOIN in SQL to Combine Tables
Subqueries in SQL
Aggregating Data with Subqueries, HAVING, and CASE in SQL
Using Spreadsheet Formulas for Sales Trend Analysis
Using COUNTIF and SUMIF for Conditional Aggregation in Spreadsheets
Using SUMPRODUCT for Advanced Spreadsheet Calculations
Using Pivot Tables for Calculations and Trend Analysis
Using Pivot Table Filters and Calculated Fields for Deeper Analysis
Comparing Calculations in Spreadsheets and SQL
Embedding Calculations in SQL Queries
Using GROUP BY and ORDER BY for Aggregated Calculations in SQL
Data Validation as an Ongoing Analytical Process
Temporary Tables and the WITH Clause in SQL
Creating Temporary Tables in SQL — Methods, Trade-offs, and Best Practices
📊 Analyze Data
Data Visualization
Connecting Data and Images
Creating Powerful Data Visualizations: Focus, Structure, and Analytical Purpose
Static vs. Dynamic Data Visualizations: Design Tradeoffs, Control, and Interactivity
Elements of Art in Data Visualization: Line, Shape, Color, Space, and Movement
Choosing the Right Visualization: Audience-Centered Design and Chart Selection
Design Thinking in Data Visualization: A User-Centered Framework
Accessibility in Data Visualization: Designing for Everyone
Introduction to Tableau
Getting Started with Tableau Public
Creating a CO₂ Emissions Visualization in Tableau Public
Effective vs. Ineffective Data Visualizations in Tableau
Using Creativity in Tableau
Linking Multiple Datasets in Tableau Public
Data Storytelling: Giving Numbers a Clear and Convincing Voice
Engaging Your Audience in Data Storytelling: Identifying the Key Message
Data Dashboards: Organizing Insight for Real-Time Decision Making
Using Filters to Create Compelling and Focused Visuals
Structuring a Persuasive Data Presentation: Turning Insights into Story
Designing Effective Data Presentation Slides: Structure, Visuals, and Professional Impact
Using a Strategic Framework to Structure Data Presentations
Weaving Data into Presentations: Hypotheses, Context, and the McCandless Method
Presentation Skills for Data Analysts: Delivering Insights with Confidence
Presenting Like a Pro: Best Practices for Data Analysts
Preparing for Q&A: Anticipating and Responding to Stakeholder Questions
Handling Objections in Data Presentations: Responding with Confidence and Clarity
Q&A Best Practices: Answering Questions with Clarity and Confidence
🎨 Data Visualization
Introduction to Python and Programming Fundamentals
Python Fundamentals
Jupyter Notebook and Coding Environments
Object-Oriented Programming (OOP) in Python
Variables in Python
Naming Conventions and Restrictions in Python
Data Types and Type Conversion in Python
Functions in Python
Code Reusability, Modularity, and Clean Code in Python
Comments, Algorithms, and Docstrings in Python
Boolean Data, Comparators, and Logical Operators in Python
Branching and Conditional Statements in Python
While Loops and Iteration in Python
For Loops in Python
range() Function and Loop Control in Python
Strings in Python
String Indexing and Slicing in Python
String Formatting with .format() in Python
Data Types vs Data Structures & Introduction to Lists
Modifying Lists in Python
Tuples in Python
Advanced Use of Loops, Lists, Tuples & List Comprehension
Dictionaries in Python
Advanced Dictionary Usage in Python
Sets in Python
Libraries, Packages, and Modules in Python
Introduction to NumPy and Vectorization
NumPy Arrays (ndarray) and Core Concepts
Introduction to Pandas (Data Analysis Library)
Pandas DataFrame & Series
Boolean Masking in Pandas
Grouping and Aggregation in Pandas (groupby, agg)
Combining Data in Pandas (concat and merge)
🐍 Data Analysis Using Python
Transferable Skills
Career Identity Statement
Career Dreamer (AI Tool for Career Exploration)
Job Search Plan (Using AI Tools)
Tailoring Your Resume
Using AI to Improve and Tailor Your Resume
Building a Professional Online Presence (Personal Brand)
Choosing the Right Job Platforms
Job Application Tracking (Using AI + Spreadsheets)
Networking for Job Search
Interview Preparation
STAR Method (Behavioral Interview)
Using AI (NotebookLM) for Interview Preparation
Practicing Interviews with AI (Gemini Live)
Post-Interview Strategy
💼 Job Search
📈 Data Analytics
topic: data preparation (57)
Why Do We Analyze Data?
The Process of Data Analysis
CRISP-DM for Data Science
Big Data: Definition, Characteristics, Evolution, and Business Impact
The First Step in Knowing Your Data
IEEE 754 Floating-Point Standard
Discovering Associations Through Data: From Everyday Patterns to Chicago Taxi Trips (September 2022)
Taxi Trips – 2022 dataset from the City of Chicago open data portal
Objective Selection of the Bin Width for a Time Histogram
Measuring Associations in Data
Measuring Associations Between Two Continuous Variables
Correlation Coefficients in Python (Pearson, Spearman, Kendall)
Karl Pearson
Harald Cramér
What Are Statistical Tests?
Eta Squared (η²): Effect Size in ANOVA
Understanding Market Baskets and Ideal Customers
What Can Association Rules Tell Us?
How Association Rules Are Discovered: Concepts, Scale, Measures, and the Apriori Approach
Apriori: Frequent Itemsets via the Apriori Algorithm
association_rules: Generating Association Rules from Frequent Itemsets (mlxtend)
Cross-Selling
Stratified Random Sampling
Linear Congruential Random Number Generator (LCG)
Partitioning Observations to Train Objective Models
Putting Similar Observations into Clusters
Clustering
Recency, Frequency, and Monetary Value (RFM)
RFM Analysis
Creating Segments of Observations for Business Reasons (RFM)
Least Squares Regression
Multiple Linear Regression
Feature Importance in Linear Regression
Forward Selection: Definition and Core Idea
Forward Selection and Model Interpretation in Linear Regression
Understanding Forward and Backward Stepwise Regression
How Shapley Values Work
Logistic Regression: Modeling Binary Outcomes via Odds and Log-Odds
Maximum Likelihood (MLE): Fitting a Distribution to Observed Data
Assessing Model Fit in Logistic Regression
Complete and Quasi-Complete Separation in Logistic Regression
Forward Selection with Nested Models and Deviance Tests
Interpreting and Assessing a Forward-Selection Logistic Regression Model for College Student Retention
Motivation of Decision Trees: An Incremental Model of Decision-Making
The CART Algorithm
Decision Trees as Piecewise Models and Their Predictive Structure
How CART Decision Trees Model Interactions
Cluster Profiling Using Decision Trees
Using Decision Trees to Explain Clustering Results
Assessing the Quality of Prediction Models
Binary Classification Models – Conceptual Framework and Evaluation Metrics
Nominal Classification Models: Model State and Evaluation Metrics
Binary Classification Model Evaluation and Threshold Optimization
Identifying Outliers Using Residuals and Studentized Residuals
AUC–ROC Curve: Evaluating Classification Model Performance
Lift Analysis for Direct Mail Campaigns: Concept, Process, and Business Value
Data Preparation & Analysis
topic: ddd (27)
Using Data Analysis to Choose the Right Advertising Strategy
Understanding Common Problem Types in Data Analytics
Applying Data Analytics Problem Types in Real Business Scenarios
Why Asking the Right Questions Matters in Data Analytics
The Relationship Between Data and Decision-Making
Quantitative and Qualitative Data in Decision-Making
Data Creates Value Only When It Is Communicated
The Difference Between Data and Metrics, and the Role of Metrics
Dashboards
Mathematical Thinking
Spreadsheets in Data Analysis
Building and Organizing a Spreadsheet
How Data Analysts Use Spreadsheets
Spreadsheet Calculations with Formulas
Common Spreadsheet Errors and How to Fix Them
Spreadsheet Functions
Defining the Problem Domain
Context and Bias in Data Analysis
Stakeholder Expectations in Data Analysis
Staying Focused on the Project Objective
Clear Communication with Stakeholders and Teams
Adapting to Communication Expectations at Work
Managing Stakeholder Expectations and Project Constraints
Balancing Speed and Accuracy in Data Analysis
Sharing Data to Drive Impact
Effective Meetings
Conflict Resolution in the Workplace
topic: deep learning (18)
What is a Neural Network?
Supervised Learning and Neural Networks
Why Deep Learning is Taking Off
Geoffrey Hinton Interview
Binary Classification and Logistic Regression (Neural Network Basics)
Logistic Regression (Binary Classification Model)
Logistic Regression – Loss Function and Cost Function
Gradient Descent in Logistic Regression
Derivatives
More Derivative Examples
Computation Graph
Derivatives with a Computation Graph
Logistic Regression Gradient Descent
Gradient Descent on m Training Examples
Vectorization in Logistic Regression
More Vectorization Examples
Vectorizing Logistic Regression
Deep Learning
topic: dirty (9)
Dirty Data vs. Clean Data
The Importance of Clean Data (revisited)
Common Issues in Dirty Data
Data Cleaning with Spreadsheets
Cleaning and Merging Multiple Datasets
Spreadsheet Tools for Data Cleaning
Using Spreadsheet Functions for Data Cleaning
Viewing Data Differently for More Effective Data Cleaning
Data Mapping and the Big Picture of Clean Data
topic: execution (11)
Defining the Problem Domain
Context and Bias in Data Analysis
Stakeholder Expectations in Data Analysis
Staying Focused on the Project Objective
Clear Communication with Stakeholders and Teams
Adapting to Communication Expectations at Work
Managing Stakeholder Expectations and Project Constraints
Balancing Speed and Accuracy in Data Analysis
Sharing Data to Drive Impact
Effective Meetings
Conflict Resolution in the Workplace
topic: foundations (27)
Why Data Analytics Matters Today
How Data Analytics Improves the Workplace
Data-Driven Decision-Making
Detectives and Data Analysts
The Six Phases of the Data Analysis Process
The Origins of Data Analysis and the Many Ways to Structure It
Understanding the Data Ecosystem
Understanding the Data Analysis Process and the Data Life Cycle
Understanding the Data Life Cycle
A Review of the Six Stages of the Data Life Cycle
The Stages of the Data Analysis Process and Their Roles
Practical Application of the Data Analysis Process
Analytical Skills and Their Core Components
Applying Analytical Skills in a Business Context
Analytical Thinking and Its Core Components
Analytical Thinking and Questions for Problem Solving
Root Cause Analysis and Business Applications of the Five Whys
Data-Driven Decision-Making and the Role of Analytical Skills
Case Studies in Data Analysis and the Practical Impact of Data-Driven Decision-Making
Overview of Core Tools Used by Data Analysts
The Role of Spreadsheets in Data Analysis and Basic Concepts
The Concept and Basic Use of SQL (Query Language)
The Role and Importance of Data Visualization
Industries Where Data Analysts Work and How Data Is Used
The Role of Business Tasks in Data Analysis
Fairness in Data Analysis
Key Factors to Consider When Choosing a Data Analytics Role
topic: framing (7)
Using Data Analysis to Choose the Right Advertising Strategy
Understanding Common Problem Types in Data Analytics
Applying Data Analytics Problem Types in Real Business Scenarios
Why Asking the Right Questions Matters in Data Analytics
The Relationship Between Data and Decision-Making
Quantitative and Qualitative Data in Decision-Making
Data Creates Value Only When It Is Communicated
topic: identity (4)
Transferable Skills
Career Identity Statement
Career Dreamer (AI Tool for Career Exploration)
Job Search Plan (Using AI Tools)
topic: integrity (8)
The Importance of Clean Data
Data Integrity and Its Risks in Data Analysis
Aligning Data with Business Objectives
Handling Insufficient Data in Data Analysis
Population, Sample Size, and Random Sampling
Statistical Power in Data Analysis
Sample Size and Data Integrity
Margin of Error
topic: interview (5)
Interview Preparation
STAR Method (Behavioral Interview)
Using AI (NotebookLM) for Interview Preparation
Practicing Interviews with AI (Gemini Live)
Post-Interview Strategy
topic: jobsearch (15)
Transferable Skills
Career Identity Statement
Career Dreamer (AI Tool for Career Exploration)
Job Search Plan (Using AI Tools)
Tailoring Your Resume
Using AI to Improve and Tailor Your Resume
Building a Professional Online Presence (Personal Brand)
Choosing the Right Job Platforms
Job Application Tracking (Using AI + Spreadsheets)
Networking for Job Search
Interview Preparation
STAR Method (Behavioral Interview)
Using AI (NotebookLM) for Interview Preparation
Practicing Interviews with AI (Gemini Live)
Post-Interview Strategy
topic: libraries (8)
Libraries, Packages, and Modules in Python
Introduction to NumPy and Vectorization
NumPy Arrays (ndarray) and Core Concepts
Introduction to Pandas (Data Analysis Library)
Pandas DataFrame & Series
Boolean Masking in Pandas
Grouping and Aggregation in Pandas (groupby, agg)
Combining Data in Pandas (concat and merge)
topic: metrics (3)
The Difference Between Data and Metrics, and the Role of Metrics
Dashboards
Mathematical Thinking
topic: organize (10)
Understanding Data Analysis
Data Organization in Analysis
Sorting and Filtering in Data Analysis
Sorting Data in Spreadsheets
Sorting and Filtering Data in SQL Using ORDER BY and WHERE
Data Formatting and Unit Conversion in Spreadsheets
Data Validation in Spreadsheets
Combining Data Validation and Conditional Formatting in Spreadsheets
Using CONCAT in SQL to Combine Text from Multiple Columns
Working with Strings in Spreadsheets (LEN, LEFT, RIGHT, FIND)
topic: prep (25)
How Data Is Generated and Collected
Choosing the Right Data to Collect
Understanding Data Types and Data Formats
Structured Data and Data Models
Data Types in Spreadsheets
Data Tables (Tabular Data)
Wide Data vs. Long Data
Understanding Bias in Data Analysis
Sampling Bias and Unbiased Data
Common Types of Data Bias
Identifying Good Data Sources (ROCCC Framework)
Identifying Bad Data Sources (When Data Does Not ROCCC)
Data Ethics in Data Analysis
Data Privacy in Data Ethics
Open Data and Openness in Data Ethics
Databases and Relational Database Concepts
Metadata in Databases
Metadata Repositories and Data Governance
Accessing Data: Internal and External Sources
Importing Data into Spreadsheets
Sorting and Filtering Data in Spreadsheets
BigQuery Account Types
Querying Data with SQL
Organizing Data for Personal and Work Projects
Data Security in Spreadsheets
topic: present (9)
Structuring a Persuasive Data Presentation: Turning Insights into Story
Designing Effective Data Presentation Slides: Structure, Visuals, and Professional Impact
Using a Strategic Framework to Structure Data Presentations
Weaving Data into Presentations: Hypotheses, Context, and the McCandless Method
Presentation Skills for Data Analysts: Delivering Insights with Confidence
Presenting Like a Pro: Best Practices for Data Analysts
Preparing for Q&A: Anticipating and Responding to Stakeholder Questions
Handling Objections in Data Presentations: Responding with Confidence and Clarity
Q&A Best Practices: Answering Questions with Clarity and Confidence
topic: principles (8)
Data Visualization
Connecting Data and Images
Creating Powerful Data Visualizations: Focus, Structure, and Analytical Purpose
Static vs. Dynamic Data Visualizations: Design Tradeoffs, Control, and Interactivity
Elements of Art in Data Visualization: Line, Shape, Color, Space, and Movement
Choosing the Right Visualization: Audience-Centered Design and Chart Selection
Design Thinking in Data Visualization: A User-Centered Framework
Accessibility in Data Visualization: Designing for Everyone
topic: process (8)
The Six Phases of the Data Analysis Process
The Origins of Data Analysis and the Many Ways to Structure It
Understanding the Data Ecosystem
Understanding the Data Analysis Process and the Data Life Cycle
Understanding the Data Life Cycle
A Review of the Six Stages of the Data Life Cycle
The Stages of the Data Analysis Process and Their Roles
Practical Application of the Data Analysis Process
topic: python (33)
Introduction to Python and Programming Fundamentals
Python Fundamentals
Jupyter Notebook and Coding Environments
Object-Oriented Programming (OOP) in Python
Variables in Python
Naming Conventions and Restrictions in Python
Data Types and Type Conversion in Python
Functions in Python
Code Reusability, Modularity, and Clean Code in Python
Comments, Algorithms, and Docstrings in Python
Boolean Data, Comparators, and Logical Operators in Python
Branching and Conditional Statements in Python
While Loops and Iteration in Python
For Loops in Python
range() Function and Loop Control in Python
Strings in Python
String Indexing and Slicing in Python
String Formatting with .format() in Python
Data Types vs Data Structures & Introduction to Lists
Modifying Lists in Python
Tuples in Python
Advanced Use of Loops, Lists, Tuples & List Comprehension
Dictionaries in Python
Advanced Dictionary Usage in Python
Sets in Python
Libraries, Packages, and Modules in Python
Introduction to NumPy and Vectorization
NumPy Arrays (ndarray) and Core Concepts
Introduction to Pandas (Data Analysis Library)
Pandas DataFrame & Series
Boolean Masking in Pandas
Grouping and Aggregation in Pandas (groupby, agg)
Combining Data in Pandas (concat and merge)
topic: sources (4)
Databases and Relational Database Concepts
Metadata in Databases
Metadata Repositories and Data Governance
Accessing Data: Internal and External Sources
topic: spreadsheets (6)
Spreadsheets in Data Analysis
Building and Organizing a Spreadsheet
How Data Analysts Use Spreadsheets
Spreadsheet Calculations with Formulas
Common Spreadsheet Errors and How to Fix Them
Spreadsheet Functions
topic: spreadsheets_sql (6)
Importing Data into Spreadsheets
Sorting and Filtering Data in Spreadsheets
BigQuery Account Types
Querying Data with SQL
Organizing Data for Personal and Work Projects
Data Security in Spreadsheets
topic: sql (7)
Introduction to SQL
Spreadsheets vs. SQL
Core SQL Queries for Data Cleaning and Analysis
Cleaning Data with SQL: Removing Duplicates and Cleaning String Variables
Using CAST to Clean and Format Data in SQL
Advanced SQL Functions for Data Cleaning
COALESCE
topic: story (4)
Data Storytelling: Giving Numbers a Clear and Convincing Voice
Engaging Your Audience in Data Storytelling: Identifying the Key Message
Data Dashboards: Organizing Insight for Real-Time Decision Making
Using Filters to Create Compelling and Focused Visuals
topic: structures (10)
Strings in Python
String Indexing and Slicing in Python
String Formatting with .format() in Python
Data Types vs Data Structures & Introduction to Lists
Modifying Lists in Python
Tuples in Python
Advanced Use of Loops, Lists, Tuples & List Comprehension
Dictionaries in Python
Advanced Dictionary Usage in Python
Sets in Python
topic: tableau (6)
Introduction to Tableau
Getting Started with Tableau Public
Creating a CO₂ Emissions Visualization in Tableau Public
Effective vs. Ineffective Data Visualizations in Tableau
Using Creativity in Tableau
Linking Multiple Datasets in Tableau Public
topic: tagging (1)
Documentation Tagging Guidelines
topic: terminology (432)
Subsampling
Class Weighting
SMOTE (Synthetic Minority Over-sampling Technique)
Oversampling
Low-pass Filtering
NearMiss (Distance-based Undersampling)
Cluster-based undersampling
Random Undersampling
Signal Processing
Time Series
Micro AUROC
Multi-label Classification
Micro F1
Single-label Classification
Micro Recall
Micro Precision
One-vs-Rest (OvR) AUROC
Macro AUROC (Macro-Averaged AUROC)
Macro F1
Macro Recall
Macro Precision
Multiclass AUROC
Gini Coefficient
Bootstrap Confidence Intervals (CIs)
Probability
Mann–Whitney U Test (also called the Wilcoxon rank-sum test)
Predictive Parity (Calibration)
Equalized Odds (Fairness)
Equal Opportunity (Fairness)
Demographic Parity (Statistical Parity)
Cross-Selling
Upselling
Customer Segmentation
SaaS (Software as a Service)
Valuation Metric
D2C (Direct-to-Consumer)
LTV:CAC Ratio
Net LTV (sometimes called Contribution LTV)
Gross LTV (Customer Lifetime Value)
Predictive LTV (pLTV)
Cohort-Based LTV (Simple Version)
Customer Lifetime
Gross Margin
Fully Loaded CAC (Customer Acquisition Cost)
Organic CAC (Customer Acquisition Cost)
Paid CAC (Customer Acquisition Cost)
Channel-Specific CAC (Customer Acquisition Cost)
Blended CAC (Customer Acquisition Cost)
Lead-Gen Software
Thompson Sampling (TS) in Bandits (Multi-Armed Bandit Problem (MAB))
Bayesian Decision Theory (BDT)
Bayesian Time Series
Posterior probability of uplift
Gaussian Processes (GPs)
Bayesian Neural Networks (BNNs)
Variational Inference (VI)
MCMC (Markov Chain Monte Carlo)
Sequential Settings
Frequentist
Binomial Likelihood
Posterior belief
Marginal Likelihood (also called The Model Evidence or Integrated Likelihood)
Posterior
Prior Belief (or Prior Probability)
Parameter(s) of Interest
Bayes’ Theorem
Conversion Rate Uplift
Bayesian Stopping Rules
Optimizely
Online Experimentation Platforms
Stopping Rules
Treatment Effect
Posterior Probability
Bayesian Sequential Testing
Likelihood Ratio (LR)
Sequential Probability Ratio Test (SPRT)
Pocock Method
O’Brien–Fleming (OBF) Method
Group Sequential Testing
Type I Error
Traditional A/B Test (Fixed-Horizon A/B Test)
Fixed-Horizon Testing
True Conversion Rate
Standard Error (SE)
True Mean (Population Mean)
Margin of Error (MoE)
Critical Value
Sample Standard Deviation
Sample Mean
Regression Coefficient
Proportion
True Population Parameter
Compromise Power Analysis
Post Hoc Power Analysis
A Priori Power Analysis
Statistical Significance
Z-Score
Two-Proportion Z-Test
Beta Distribution
Google Experiments
Minimum Detectable Lift (MDL)
Trivial Effects
Sample size
Power (1 – β)
Significance Level (α)
Effect Size (δ)
Hypothesis Testing
Ranking Algorithms
Probabilistic Interleaving
Team Draft Interleaving (TDI)
Balanced Interleaving
Causal Impact
Bandit Algorithms
A/B/n Test
Multivariate Test (MVT)
Risk of Peeking
Causal Inference
P-Value (probability value)
Z-Test
T-Test
Session Length
Revenue per User (RPU / ARPU)
Churn
Retention
Statistically Significant
IID (Independent and Identically Distributed)
Temporal autocorrelation (Serial Correlation)
Blocked Splits (Single Holdout)
Sliding Window (Rolling Window) Cross-Validation
Expanding Window Cross-Validation
Data Leakage
Stratified Group K-Fold
Stratified Shuffle Split
Multiclass stratified CV
k-fold cross-validation
Cross-Validation (CV)
Re-scoring
Drift Detection
Model Distillation (Knowledge Distillation)
Early Stopping
Epochs
Hyperparameter
AI (Artificial Intelligence)
Machine Learning (ML)
Medical AI
KYC
FTEs
AWS SageMaker
Vertex AI
OpenAI API (ML API)
AWS SageMaker Endpoints
Cloud Inference with Big Payloads
Cloud Inference
Ensemble
Model Weights
FLOPs
OpEx
LLMs (Large Language Models)
Recalibration
Reweighting
Continuous Retraining
Monitoring Pipelines
Active Learning
Bayesian Correction
Recalibrate Thresholds
Guardrails (in ML & Data Systems)
Model KPIs (Key Performance Indicators)
Lagging Indicators
Leading Indicators
Windows (in Time-Series)
Autoencoder
Frozen Encoder
Embedding
Representation Shift
Classifier Two-Sample Tests (C2STs)
Energy Distance
Maximum Mean Discrepancy (MMD)
Cardinality in Categorical Data
Categorical Drift
Cramér’s V
Macro Shifts
Categorical Explosions
Cohort
Off-Distribution
Discriminatory Power
KS Statistic (Kolmogorov–Smirnov Statistic)
Model Stability
Feature Values
Four-Fifths (80%) Rule
SLI (Service Level Indicator)
ROI (Return on Investment)
Treatment Cost
Incremental Revenue
Incremental Recovery Rate (IRR)
Incremental Sales
Random Targeting Strategy
Causal ML (Causal Machine Learning)
Cumulative Uplift
Population Proportion
Incremental Gain
Total Incremental Benefit (TIB)
Cumulative Incremental Gain (CIG)
Qini Curve
Uplift Score
Uplift Models
Ops Health Dashboard
SLA Breach Rate
SLA (Service Level Agreement)
Supplier Constraints
Long Lead Times
Slow-Moving SKUs
SKU
Real-Time Inventory Tracking
Supplier Management
Demand Forecasting
Reorder Point (ROP) Optimization
Safety Stock
Backorder Rate
Lost Sales Value
Fill Rate
Stockout Rate
Prophet — Time Series Forecasting by Facebook (Meta)
LSTM — Long Short-Term Memory Networks
ARIMA (AutoRegressive Integrated Moving Average)
Return Distribution
Value-at-Risk (VaR)
Risk Forecast
Probabilistic Scoring
Full Distribution
Continuous Probabilistic Forecasts
Classification Probability
Quantile Forecasts
Point Forecasts
Strictly Proper Scoring Rules
Probability Forecasts
Target Variable
Probability Density
Normal Distribution
Probability Mass
Probability Distribution
Probabilistic Forecasts
Deterministic forecasts
Cumulative Distribution Function (CDF)
M-Competitions (Makridakis Competitions)
Forecasting Benchmarks
Average Absolute Error (AAE)
Seasonal Lag
Simple Baseline Methods
Naïve Baseline Forecast
Forecast Error
Forecasting Competitions
Predicting Percentiles
Prediction Intervals (PI)
Quantile Regression
Quantile Level
Time Series Forecasting
Log-Space
Relative accuracy
R² (R-squared)
Long-Tail Items
Self-Information of Popularity
Relevance in Recommender Systems
Genre Overlap
Jaccard index
Cosine Similarity of Item Features
Intra-List Diversity (ILD)
Dominating in Recommender Systems
Catalog Coverage
User Coverage
Item Coverage
Diminishing Utility
DCG (Discounted Cumulative Gain)
Kaggle
TREC (Text REtrieval Conference)
Adaptive ECE (Expected Calibration Error with Adaptive Binning)
Maximum Calibration Error (MCE)
ROC Curve (Receiver Operating Characteristic)
Murphy’s Decomposition
Temperature Scaling
Platt Scaling
Isotonic Regression
Support Vector Machines (SVMs)
Underconfident
Overconfident
Confidence Level
Risk-Based Decisions
Neural Networks
Binary Cross-Entropy (BCE)
Loss Functions
Underflow
Logit Space
Logistic Regression
Binary Classification
Classification Models
Log-Odds
Softmax Function
Sigmoid Function
Squashing Function
Conversion Rate (CR)
Cost-Per-Click (CPC) Models
Causal Trees
Uplift Random Forests
Uplift Curve
Likelihood
Correlation
Causal Effect
Outlier
Mean Squared Error (MSE)
Regression Models
One-vs-Rest (OvR)
Multiclass Classification
Partial AUC (pAUC)
Micro AUC
Macro AUC
Median
Mean
Sensitivity in Feature Engineering
Encode (in Feature Engineering)
Normalize (in Feature Engineering)
Embedding Similarity
Computer Vision (CV)
Natural Language Processing (NLP)
Accuracy
Chi-square (χ²) Test
Kolmogorov–Smirnov (KS) Test
Jensen–Shannon (JS) Divergence
Kullback–Leibler (KL) Divergence
Statistical Tests
Seasonality
Concept Drift
Data Drift
Fair Lending laws
Basel III
High-Stakes Domains
Deep Ensembles
Counterfactual Explanations
LIME (Local Interpretable Model-agnostic Explanations)
SHAP (SHapley Additive exPlanations)
Post-hoc Explainability
Decision Trees
Linear Models
Caching
Quantization
ONNX (Open Neural Network Exchange)
Full Annotation
Weak Supervision
TPU Clusters
Statistical Power
Drift Guardrails
Latency Guardrails
Fairness Guardrails
DeLong’s Test
Dataset Shift
Label Noise
Evaluation Set
Clopper–Pearson Interval
Wilson Score Interval
Per-class Precision (sometimes called class-wise precision)
Multiclass Precision
Multilabel Precision
Weighted Averaging
Harmonic Mean
F1-score
Model Score
Bootstrap
Average Precision (AP)
Upsampling
Downsampling
Micro Averaging
Macro Averaging
AUC (Area Under the Curve)
Fairness parity
LTV (Customer Lifetime Value)
CAC (Customer Acquisition Cost)
Bayesian Inference.
Sequential Testing (also called sequential analysis)
Confidence Intervals (CIs)
Power Analysis
Interleaving Tests
A/B Testing
Time-based splits (a.k.a. Temporal Cross-Validation, Rolling Window Validation)
k-fold Stratified Cross-Validation (Stratified CV)
Compute budgets
Manual review minutes
Inference Cost (Inference $)
Label Drift (a.k.a. Target Drift)
Covariate Drift (a.k.a. Covariate Shift)
KS shift (Kolmogorov–Smirnov shift)
PSI (Population Stability Index)
Selection Rate
SLOs (Service Level Objectives)
Cannibalization
Revenue net of treatment cost
Incremental Conversions
Uplift@k
AUUC (Area Under the Uplift Curve)
Qini Coefficient
Crew Overtime
SLA Breaches
Overstock %
Stockouts
Continuous Ranked Probability Score (CRPS)
MASE (Mean Absolute Scaled Error)
Pinball Loss (a.k.a. Quantile Loss)
WMAPE (Weighted Mean Absolute Percentage Error)
sMAPE (Symmetric Mean Absolute Percentage Error)
RMSLE (Root Mean Squared Logarithmic Error)
Mean Absolute Error (MAE)
Novelty (in Recommender Systems)
Diversity (in Recommender Systems)
Coverage
Hit Rate (HR)
NDCG (Normalized Discounted Cumulative Gain)
Mean Average Precision (MAP)
Expected Calibration Error (ECE)
Reliability Curves (also called Calibration Curves)
Log Loss (also called Logarithmic Loss or Cross-Entropy Loss)
Brier Score
Calibration quality (Model Calibration)
Logits
CTR (Click-Through Rate)
WAPE (Weighted Absolute Percentage Error)
Recall
Uplift
Mean Absolute Percentage Error (MAPE)
Root Mean Squared Error (RMSE)
ROC-AUC (Receiver Operating Characteristic – Area Under Curve, = AUROC)
Baseline Heuristics
Precision (a.k.a. Positive Predictive Value, PPV)
Precision–Recall AUC (PR-AUC)
Advanced Sorting in Spreadsheets
Terminology
topic: thinking (7)
Analytical Skills and Their Core Components
Applying Analytical Skills in a Business Context
Analytical Thinking and Its Core Components
Analytical Thinking and Questions for Problem Solving
Root Cause Analysis and Business Applications of the Five Whys
Data-Driven Decision-Making and the Role of Analytical Skills
Case Studies in Data Analysis and the Practical Impact of Data-Driven Decision-Making
topic: time series (19)
What Are Time Series, and How Are They Used?
Getting Started with R
A Gentle Introduction to Stationarity
Weak and Strong Stationarity
Linear Processes
Understanding ARMA Processes
Computing ACFs of Causal AR(2) Processes Using Difference Equations
Understanding ACFs via Difference Equations for AR(p) and ARMA(p, q)
Best Linear Predictor of a Stationary Process
Sample ACF and Sample PACF
Preliminary Estimation for AR Models and the Yule–Walker Equations
Maximum Likelihood Estimation for ARMA Models (Gaussian MLE)
Diagnostics After Fitting a Time Series Model
Order Selection for Time Series Models
ARIMA Models: How Nonstationary Models Are Built from Stationary Ones
SARIMA Models: Seasonal ARIMA
Beyond One-Step Ahead Predictions
Exponential Smoothing Models
Time Series
topic: tools (8)
Overview of Core Tools Used by Data Analysts
The Role of Spreadsheets in Data Analysis and Basic Concepts
The Concept and Basic Use of SQL (Query Language)
The Role and Importance of Data Visualization
Industries Where Data Analysts Work and How Data Is Used
The Role of Business Tasks in Data Analysis
Fairness in Data Analysis
Key Factors to Consider When Choosing a Data Analytics Role
topic: types (7)
How Data Is Generated and Collected
Choosing the Right Data to Collect
Understanding Data Types and Data Formats
Structured Data and Data Models
Data Types in Spreadsheets
Data Tables (Tabular Data)
Wide Data vs. Long Data
topic: verify (8)
Verifying and Reporting Data Integrity
Verifying Data-Cleaning Efforts
Verification Techniques: Using Spreadsheets and SQL to Catch Repeated Errors
Documenting Data-Cleaning Changes
Reporting Data-Cleaning Results
Using Feedback from Data Cleaning to Improve Data Quality
Refining a Resume for Data Analytics Roles
Exploring Data Analyst Job Opportunities
topic: viz (27)
Data Visualization
Connecting Data and Images
Creating Powerful Data Visualizations: Focus, Structure, and Analytical Purpose
Static vs. Dynamic Data Visualizations: Design Tradeoffs, Control, and Interactivity
Elements of Art in Data Visualization: Line, Shape, Color, Space, and Movement
Choosing the Right Visualization: Audience-Centered Design and Chart Selection
Design Thinking in Data Visualization: A User-Centered Framework
Accessibility in Data Visualization: Designing for Everyone
Introduction to Tableau
Getting Started with Tableau Public
Creating a CO₂ Emissions Visualization in Tableau Public
Effective vs. Ineffective Data Visualizations in Tableau
Using Creativity in Tableau
Linking Multiple Datasets in Tableau Public
Data Storytelling: Giving Numbers a Clear and Convincing Voice
Engaging Your Audience in Data Storytelling: Identifying the Key Message
Data Dashboards: Organizing Insight for Real-Time Decision Making
Using Filters to Create Compelling and Focused Visuals
Structuring a Persuasive Data Presentation: Turning Insights into Story
Designing Effective Data Presentation Slides: Structure, Visuals, and Professional Impact
Using a Strategic Framework to Structure Data Presentations
Weaving Data into Presentations: Hypotheses, Context, and the McCandless Method
Presentation Skills for Data Analysts: Delivering Insights with Confidence
Presenting Like a Pro: Best Practices for Data Analysts
Preparing for Q&A: Anticipating and Responding to Stakeholder Questions
Handling Objections in Data Presentations: Responding with Confidence and Clarity
Q&A Best Practices: Answering Questions with Clarity and Confidence
topic: why (4)
Why Data Analytics Matters Today
How Data Analytics Improves the Workplace
Data-Driven Decision-Making
Detectives and Data Analysts
Tags overview
My tags: component: model
My tags: component: model
#
With this tag
plot_calibration with examples
Edit on GitHub
Show Source