My tags: purpose: reference#
With this tag
- Documentation Tagging Guidelines
- The three steps of Bayesian data analysis
- General Notation for Statistical Inference
- Bayesian Inference
- Discrete Bayesian Examples – Genetics and Spell Checking (with θ)
- Probability as a Measure of Uncertainty
- Example — Probabilities from Football Point Spreads
- Example — Calibration for Record Linkage
- Some Useful Results from Probability Theory
- Computation and Software
- Bayesian Inference in Applied Statistics
- Estimating a Probability from Binomial Data
- Posterior as a Compromise Between Data and Prior Information
- Summarizing Posterior Inference
- Informative Prior Distributions
- Normal Distribution with Known Variance
- Other Standard Single-Parameter Models
- Informative Prior Distribution for Cancer Rates
- Noninformative Prior Distributions
- Weakly Informative Prior Distributions
- Averaging Over Nuisance Parameters
- Normal Data with a Noninformative Prior Distribution
- Normal Data with a Conjugate Prior Distribution
- Multinomial Model for Categorical Data
- Multivariate Normal Model with Known Variance
- Multivariate Normal with Unknown Mean and Variance
- Example: Bayesian analysis of a bioassay experiment (logistic, nonconjugate)
- Summary of Elementary Modeling and Computation
- Normal Approximations to the Posterior Distribution
- Large-Sample Theory
- Counterexamples to large-sample (asymptotic) Bayesian theorems
- Frequency Evaluations of Bayesian Inferences
- Bayesian interpretations of other statistical methods
- Constructing a Parameterized Prior Distribution
- Exchangeability and hierarchical models
- Bayesian analysis of conjugate hierarchical models
- Normal model with exchangeable parameters
- Example: parallel experiments in eight schools
- Hierarchical modeling applied to a meta-analysis
- Weakly Informative Priors for Variance Parameters
- The Place of Model Checking in Applied Bayesian Statistics
- Do the Inferences from the Model Make Sense?
- Posterior predictive checking
- Graphical posterior predictive checks
- Model checking for the educational testing example
- Measures of predictive accuracy
- Model comparison based on predictive performance
- Model comparison using Bayes factors
- Continuous model expansion
- Implicit assumptions and model expansion: an example
- Bayesian inference requires a model for data collection
- Data-collection models and ignorability
- Sample surveys
- Designed experiments
- Sensitivity and the role of randomization
- Observational studies
- Censoring and truncation
- Bayesian decision theory in different contexts
- Using regression predictions: survey incentives
- Multistage decision making: medical screening
- Hierarchical decision analysis for home radon
- Personal vs. institutional decision analysis
- Numerical integration
- Distributional approximations
- Direct simulation and rejection sampling
- Importance sampling
- How many simulation draws are needed?
- Computing environments
- Debugging Bayesian computing
- Gibbs sampler
- Metropolis and Metropolis-Hastings algorithms
- Using Gibbs and Metropolis as building blocks
- Inference and assessing convergence
- Effective number of simulation draws
- Example: hierarchical normal model
- Efficient Gibbs samplers
- Efficient Metropolis jumping rules
- Further extensions to Gibbs and Metropolis
- Hamiltonian Monte Carlo
- Hamiltonian Monte Carlo for a hierarchical model
- Stan: developing a computing environment
- Finding posterior modes
- Boundary-avoiding priors for modal summaries
- Normal and related mixture approximations
- Finding marginal posterior modes using EM
- Conditional and marginal posterior approximations
- Example: hierarchical normal model (continued)
- Variational inference
- Expectation propagation
- Other approximations
- Unknown normalizing factors
- Conditional modeling
- Bayesian analysis of classical regression
- Regression for causal inference: incumbency and voting
- Goals of regression analysis
- Assembling the matrix of explanatory variables
- Regularization and dimension reduction
- Unequal variances and correlations
- Including numerical prior information
- Regression coefficients exchangeable in batches
- Example: forecasting U.S. presidential elections
- Interpreting a normal prior distribution as extra data
- Varying intercepts and slopes
- Computation: batching and transformation
- Analysis of variance and the batching of coefficients
- Hierarchical models for batches of variance components
- Standard generalized linear model likelihoods
- Working with generalized linear models
- Weakly informative priors for logistic regression
- Overdispersed Poisson regression for police stops
- State-level opinons from national polls
- Models for multivariate and multinomial responses
- Loglinear models for multivariate discrete data
- Aspects of robustness
- Overdispersed versions of standard models
- Posterior inference and computation
- Robust inference for the eight schools
- Robust regression using t-distributed errors
- Notation
- Multiple imputation
- Missing data in the multivariate normal and t models
- Example: multiple imputation for a series of polls
- Missing values with counted data
- Example: an opinion poll in Slovenia
- Example: serial dilution assay
- Example: population toxicokinetics
- Splines and weighted sums of basis functions
- Basis selection and shrinkage of coefficients
- Non-normal models and regression surfaces
- Gaussian process regression
- Example: birthdays and birthdates
- Latent Gaussian process models
- Functional data analysis
- Density estimation and regression
- Setting up and interpreting mixture models
- Example: reaction times and schizophrenia
- Label switching and posterior computation
- Unspecified number of mixture components
- Mixture models for classification and regression
- Bayesian histograms
- Dirichlet process prior distributions
- Dirichlet process mixtures
- Beyond density estimation
- Hierarchical dependence
- Density regression
- Bayesian Data Analysis
- Why Data Analytics Matters Today
- How Data Analytics Improves the Workplace
- Data-Driven Decision-Making
- Detectives and Data Analysts
- The Six Phases of the Data Analysis Process
- The Origins of Data Analysis and the Many Ways to Structure It
- Understanding the Data Ecosystem
- Understanding the Data Analysis Process and the Data Life Cycle
- Understanding the Data Life Cycle
- A Review of the Six Stages of the Data Life Cycle
- The Stages of the Data Analysis Process and Their Roles
- Practical Application of the Data Analysis Process
- Analytical Skills and Their Core Components
- Applying Analytical Skills in a Business Context
- Analytical Thinking and Its Core Components
- Analytical Thinking and Questions for Problem Solving
- Root Cause Analysis and Business Applications of the Five Whys
- Data-Driven Decision-Making and the Role of Analytical Skills
- Case Studies in Data Analysis and the Practical Impact of Data-Driven Decision-Making
- Overview of Core Tools Used by Data Analysts
- The Role of Spreadsheets in Data Analysis and Basic Concepts
- The Concept and Basic Use of SQL (Query Language)
- The Role and Importance of Data Visualization
- Industries Where Data Analysts Work and How Data Is Used
- The Role of Business Tasks in Data Analysis
- Fairness in Data Analysis
- Key Factors to Consider When Choosing a Data Analytics Role
- 🌱 Foundations
- Using Data Analysis to Choose the Right Advertising Strategy
- Understanding Common Problem Types in Data Analytics
- Applying Data Analytics Problem Types in Real Business Scenarios
- Why Asking the Right Questions Matters in Data Analytics
- The Relationship Between Data and Decision-Making
- Quantitative and Qualitative Data in Decision-Making
- Data Creates Value Only When It Is Communicated
- The Difference Between Data and Metrics, and the Role of Metrics
- Dashboards
- Mathematical Thinking
- Spreadsheets in Data Analysis
- Building and Organizing a Spreadsheet
- How Data Analysts Use Spreadsheets
- Spreadsheet Calculations with Formulas
- Common Spreadsheet Errors and How to Fix Them
- Spreadsheet Functions
- Defining the Problem Domain
- Context and Bias in Data Analysis
- Stakeholder Expectations in Data Analysis
- Staying Focused on the Project Objective
- Clear Communication with Stakeholders and Teams
- Adapting to Communication Expectations at Work
- Managing Stakeholder Expectations and Project Constraints
- Balancing Speed and Accuracy in Data Analysis
- Sharing Data to Drive Impact
- Effective Meetings
- Conflict Resolution in the Workplace
- 🎯 Data-Driven Decisions
- How Data Is Generated and Collected
- Choosing the Right Data to Collect
- Understanding Data Types and Data Formats
- Structured Data and Data Models
- Data Types in Spreadsheets
- Data Tables (Tabular Data)
- Wide Data vs. Long Data
- Understanding Bias in Data Analysis
- Sampling Bias and Unbiased Data
- Common Types of Data Bias
- Identifying Good Data Sources (ROCCC Framework)
- Identifying Bad Data Sources (When Data Does Not ROCCC)
- Data Ethics in Data Analysis
- Data Privacy in Data Ethics
- Open Data and Openness in Data Ethics
- Databases and Relational Database Concepts
- Metadata in Databases
- Metadata Repositories and Data Governance
- Accessing Data: Internal and External Sources
- Importing Data into Spreadsheets
- Sorting and Filtering Data in Spreadsheets
- BigQuery Account Types
- Querying Data with SQL
- Organizing Data for Personal and Work Projects
- Data Security in Spreadsheets
- 📦 Data Preparation
- The Importance of Clean Data
- Data Integrity and Its Risks in Data Analysis
- Aligning Data with Business Objectives
- Handling Insufficient Data in Data Analysis
- Population, Sample Size, and Random Sampling
- Statistical Power in Data Analysis
- Sample Size and Data Integrity
- Margin of Error
- Dirty Data vs. Clean Data
- The Importance of Clean Data (revisited)
- Common Issues in Dirty Data
- Data Cleaning with Spreadsheets
- Cleaning and Merging Multiple Datasets
- Spreadsheet Tools for Data Cleaning
- Using Spreadsheet Functions for Data Cleaning
- Viewing Data Differently for More Effective Data Cleaning
- Data Mapping and the Big Picture of Clean Data
- Introduction to SQL
- Spreadsheets vs. SQL
- Core SQL Queries for Data Cleaning and Analysis
- Cleaning Data with SQL: Removing Duplicates and Cleaning String Variables
- Using CAST to Clean and Format Data in SQL
- Advanced SQL Functions for Data Cleaning
- COALESCE
- Verifying and Reporting Data Integrity
- Verifying Data-Cleaning Efforts
- Verification Techniques: Using Spreadsheets and SQL to Catch Repeated Errors
- Documenting Data-Cleaning Changes
- Reporting Data-Cleaning Results
- Using Feedback from Data Cleaning to Improve Data Quality
- Refining a Resume for Data Analytics Roles
- Exploring Data Analyst Job Opportunities
- 🧽 Data Cleaning & Preparation
- Understanding Data Analysis
- Data Organization in Analysis
- Sorting and Filtering in Data Analysis
- Sorting Data in Spreadsheets
- Sorting and Filtering Data in SQL Using ORDER BY and WHERE
- Data Formatting and Unit Conversion in Spreadsheets
- Data Validation in Spreadsheets
- Combining Data Validation and Conditional Formatting in Spreadsheets
- Using CONCAT in SQL to Combine Text from Multiple Columns
- Working with Strings in Spreadsheets (LEN, LEFT, RIGHT, FIND)
- Problem-Solving and Seeking Help in Data Analysis
- How to Effectively Search for Solutions Online as a Data Analyst
- Choosing the Right Tool in Data Analysis
- Preparing Data for VLOOKUP in Spreadsheets
- Using VLOOKUP to Combine Data Across Spreadsheets
- Troubleshooting VLOOKUP and Building a Problem-Solving Framework
- Using JOIN in SQL to Combine Tables
- Subqueries in SQL
- Aggregating Data with Subqueries, HAVING, and CASE in SQL
- Using Spreadsheet Formulas for Sales Trend Analysis
- Using COUNTIF and SUMIF for Conditional Aggregation in Spreadsheets
- Using SUMPRODUCT for Advanced Spreadsheet Calculations
- Using Pivot Tables for Calculations and Trend Analysis
- Using Pivot Table Filters and Calculated Fields for Deeper Analysis
- Comparing Calculations in Spreadsheets and SQL
- Embedding Calculations in SQL Queries
- Using GROUP BY and ORDER BY for Aggregated Calculations in SQL
- Data Validation as an Ongoing Analytical Process
- Temporary Tables and the WITH Clause in SQL
- Creating Temporary Tables in SQL — Methods, Trade-offs, and Best Practices
- 📊 Analyze Data
- Data Visualization
- Connecting Data and Images
- Creating Powerful Data Visualizations: Focus, Structure, and Analytical Purpose
- Static vs. Dynamic Data Visualizations: Design Tradeoffs, Control, and Interactivity
- Elements of Art in Data Visualization: Line, Shape, Color, Space, and Movement
- Choosing the Right Visualization: Audience-Centered Design and Chart Selection
- Design Thinking in Data Visualization: A User-Centered Framework
- Accessibility in Data Visualization: Designing for Everyone
- Introduction to Tableau
- Getting Started with Tableau Public
- Creating a CO₂ Emissions Visualization in Tableau Public
- Effective vs. Ineffective Data Visualizations in Tableau
- Using Creativity in Tableau
- Linking Multiple Datasets in Tableau Public
- Data Storytelling: Giving Numbers a Clear and Convincing Voice
- Engaging Your Audience in Data Storytelling: Identifying the Key Message
- Data Dashboards: Organizing Insight for Real-Time Decision Making
- Using Filters to Create Compelling and Focused Visuals
- Structuring a Persuasive Data Presentation: Turning Insights into Story
- Designing Effective Data Presentation Slides: Structure, Visuals, and Professional Impact
- Using a Strategic Framework to Structure Data Presentations
- Weaving Data into Presentations: Hypotheses, Context, and the McCandless Method
- Presentation Skills for Data Analysts: Delivering Insights with Confidence
- Presenting Like a Pro: Best Practices for Data Analysts
- Preparing for Q&A: Anticipating and Responding to Stakeholder Questions
- Handling Objections in Data Presentations: Responding with Confidence and Clarity
- Q&A Best Practices: Answering Questions with Clarity and Confidence
- 🎨 Data Visualization
- Introduction to Python and Programming Fundamentals
- Python Fundamentals
- Jupyter Notebook and Coding Environments
- Object-Oriented Programming (OOP) in Python
- Variables in Python
- Naming Conventions and Restrictions in Python
- Data Types and Type Conversion in Python
- Functions in Python
- Code Reusability, Modularity, and Clean Code in Python
- Comments, Algorithms, and Docstrings in Python
- Boolean Data, Comparators, and Logical Operators in Python
- Branching and Conditional Statements in Python
- While Loops and Iteration in Python
- For Loops in Python
- range() Function and Loop Control in Python
- Strings in Python
- String Indexing and Slicing in Python
- String Formatting with .format() in Python
- Data Types vs Data Structures & Introduction to Lists
- Modifying Lists in Python
- Tuples in Python
- Advanced Use of Loops, Lists, Tuples & List Comprehension
- Dictionaries in Python
- Advanced Dictionary Usage in Python
- Sets in Python
- Libraries, Packages, and Modules in Python
- Introduction to NumPy and Vectorization
- NumPy Arrays (ndarray) and Core Concepts
- Introduction to Pandas (Data Analysis Library)
- Pandas DataFrame & Series
- Boolean Masking in Pandas
- Grouping and Aggregation in Pandas (groupby, agg)
- Combining Data in Pandas (concat and merge)
- 🐍 Data Analysis Using Python
- Transferable Skills
- Career Identity Statement
- Career Dreamer (AI Tool for Career Exploration)
- Job Search Plan (Using AI Tools)
- Tailoring Your Resume
- Using AI to Improve and Tailor Your Resume
- Building a Professional Online Presence (Personal Brand)
- Choosing the Right Job Platforms
- Job Application Tracking (Using AI + Spreadsheets)
- Networking for Job Search
- Interview Preparation
- STAR Method (Behavioral Interview)
- Using AI (NotebookLM) for Interview Preparation
- Practicing Interviews with AI (Gemini Live)
- Post-Interview Strategy
- 💼 Job Search
- 📈 Data Analytics
- Why Do We Analyze Data?
- The Process of Data Analysis
- CRISP-DM for Data Science
- Big Data: Definition, Characteristics, Evolution, and Business Impact
- The First Step in Knowing Your Data
- IEEE 754 Floating-Point Standard
- Discovering Associations Through Data: From Everyday Patterns to Chicago Taxi Trips (September 2022)
- Taxi Trips – 2022 dataset from the City of Chicago open data portal
- Objective Selection of the Bin Width for a Time Histogram
- Measuring Associations in Data
- Measuring Associations Between Two Continuous Variables
- Correlation Coefficients in Python (Pearson, Spearman, Kendall)
- Karl Pearson
- Harald Cramér
- What Are Statistical Tests?
- Eta Squared (η²): Effect Size in ANOVA
- Understanding Market Baskets and Ideal Customers
- What Can Association Rules Tell Us?
- How Association Rules Are Discovered: Concepts, Scale, Measures, and the Apriori Approach
- Apriori: Frequent Itemsets via the Apriori Algorithm
- association_rules: Generating Association Rules from Frequent Itemsets (mlxtend)
- Cross-Selling
- Stratified Random Sampling
- Linear Congruential Random Number Generator (LCG)
- Partitioning Observations to Train Objective Models
- Putting Similar Observations into Clusters
- Clustering
- Recency, Frequency, and Monetary Value (RFM)
- RFM Analysis
- Creating Segments of Observations for Business Reasons (RFM)
- Least Squares Regression
- Multiple Linear Regression
- Feature Importance in Linear Regression
- Forward Selection: Definition and Core Idea
- Forward Selection and Model Interpretation in Linear Regression
- Understanding Forward and Backward Stepwise Regression
- How Shapley Values Work
- Logistic Regression: Modeling Binary Outcomes via Odds and Log-Odds
- Maximum Likelihood (MLE): Fitting a Distribution to Observed Data
- Assessing Model Fit in Logistic Regression
- Complete and Quasi-Complete Separation in Logistic Regression
- Forward Selection with Nested Models and Deviance Tests
- Interpreting and Assessing a Forward-Selection Logistic Regression Model for College Student Retention
- Motivation of Decision Trees: An Incremental Model of Decision-Making
- The CART Algorithm
- Decision Trees as Piecewise Models and Their Predictive Structure
- How CART Decision Trees Model Interactions
- Cluster Profiling Using Decision Trees
- Using Decision Trees to Explain Clustering Results
- Assessing the Quality of Prediction Models
- Binary Classification Models – Conceptual Framework and Evaluation Metrics
- Nominal Classification Models: Model State and Evaluation Metrics
- Binary Classification Model Evaluation and Threshold Optimization
- Identifying Outliers Using Residuals and Studentized Residuals
- AUC–ROC Curve: Evaluating Classification Model Performance
- Lift Analysis for Direct Mail Campaigns: Concept, Process, and Business Value
- Data Preparation & Analysis
- What is a Neural Network?
- Supervised Learning and Neural Networks
- Why Deep Learning is Taking Off
- Geoffrey Hinton Interview
- Binary Classification and Logistic Regression (Neural Network Basics)
- Logistic Regression (Binary Classification Model)
- Logistic Regression – Loss Function and Cost Function
- Gradient Descent in Logistic Regression
- Derivatives
- More Derivative Examples
- Computation Graph
- Derivatives with a Computation Graph
- Logistic Regression Gradient Descent
- Gradient Descent on m Training Examples
- Vectorization in Logistic Regression
- More Vectorization Examples
- Vectorizing Logistic Regression
- Deep Learning
- Subsampling
- Class Weighting
- SMOTE (Synthetic Minority Over-sampling Technique)
- Oversampling
- Low-pass Filtering
- NearMiss (Distance-based Undersampling)
- Cluster-based undersampling
- Random Undersampling
- Signal Processing
- Time Series
- Micro AUROC
- Multi-label Classification
- Micro F1
- Single-label Classification
- Micro Recall
- Micro Precision
- One-vs-Rest (OvR) AUROC
- Macro AUROC (Macro-Averaged AUROC)
- Macro F1
- Macro Recall
- Macro Precision
- Multiclass AUROC
- Gini Coefficient
- Bootstrap Confidence Intervals (CIs)
- Probability
- Mann–Whitney U Test (also called the Wilcoxon rank-sum test)
- Predictive Parity (Calibration)
- Equalized Odds (Fairness)
- Equal Opportunity (Fairness)
- Demographic Parity (Statistical Parity)
- Cross-Selling
- Upselling
- Customer Segmentation
- SaaS (Software as a Service)
- Valuation Metric
- D2C (Direct-to-Consumer)
- LTV:CAC Ratio
- Net LTV (sometimes called Contribution LTV)
- Gross LTV (Customer Lifetime Value)
- Predictive LTV (pLTV)
- Cohort-Based LTV (Simple Version)
- Customer Lifetime
- Gross Margin
- Fully Loaded CAC (Customer Acquisition Cost)
- Organic CAC (Customer Acquisition Cost)
- Paid CAC (Customer Acquisition Cost)
- Channel-Specific CAC (Customer Acquisition Cost)
- Blended CAC (Customer Acquisition Cost)
- Lead-Gen Software
- Thompson Sampling (TS) in Bandits (Multi-Armed Bandit Problem (MAB))
- Bayesian Decision Theory (BDT)
- Bayesian Time Series
- Posterior probability of uplift
- Gaussian Processes (GPs)
- Bayesian Neural Networks (BNNs)
- Variational Inference (VI)
- MCMC (Markov Chain Monte Carlo)
- Sequential Settings
- Frequentist
- Binomial Likelihood
- Posterior belief
- Marginal Likelihood (also called The Model Evidence or Integrated Likelihood)
- Posterior
- Prior Belief (or Prior Probability)
- Parameter(s) of Interest
- Bayes’ Theorem
- Conversion Rate Uplift
- Bayesian Stopping Rules
- Optimizely
- Online Experimentation Platforms
- Stopping Rules
- Treatment Effect
- Posterior Probability
- Bayesian Sequential Testing
- Likelihood Ratio (LR)
- Sequential Probability Ratio Test (SPRT)
- Pocock Method
- O’Brien–Fleming (OBF) Method
- Group Sequential Testing
- Type I Error
- Traditional A/B Test (Fixed-Horizon A/B Test)
- Fixed-Horizon Testing
- True Conversion Rate
- Standard Error (SE)
- True Mean (Population Mean)
- Margin of Error (MoE)
- Critical Value
- Sample Standard Deviation
- Sample Mean
- Regression Coefficient
- Proportion
- True Population Parameter
- Compromise Power Analysis
- Post Hoc Power Analysis
- A Priori Power Analysis
- Statistical Significance
- Z-Score
- Two-Proportion Z-Test
- Beta Distribution
- Google Experiments
- Minimum Detectable Lift (MDL)
- Trivial Effects
- Sample size
- Power (1 – β)
- Significance Level (α)
- Effect Size (δ)
- Hypothesis Testing
- Ranking Algorithms
- Probabilistic Interleaving
- Team Draft Interleaving (TDI)
- Balanced Interleaving
- Causal Impact
- Bandit Algorithms
- A/B/n Test
- Multivariate Test (MVT)
- Risk of Peeking
- Causal Inference
- P-Value (probability value)
- Z-Test
- T-Test
- Session Length
- Revenue per User (RPU / ARPU)
- Churn
- Retention
- Statistically Significant
- IID (Independent and Identically Distributed)
- Temporal autocorrelation (Serial Correlation)
- Blocked Splits (Single Holdout)
- Sliding Window (Rolling Window) Cross-Validation
- Expanding Window Cross-Validation
- Data Leakage
- Stratified Group K-Fold
- Stratified Shuffle Split
- Multiclass stratified CV
- k-fold cross-validation
- Cross-Validation (CV)
- Re-scoring
- Drift Detection
- Model Distillation (Knowledge Distillation)
- Early Stopping
- Epochs
- Hyperparameter
- AI (Artificial Intelligence)
- Machine Learning (ML)
- Medical AI
- KYC
- FTEs
- AWS SageMaker
- Vertex AI
- OpenAI API (ML API)
- AWS SageMaker Endpoints
- Cloud Inference with Big Payloads
- Cloud Inference
- Ensemble
- Model Weights
- FLOPs
- OpEx
- LLMs (Large Language Models)
- Recalibration
- Reweighting
- Continuous Retraining
- Monitoring Pipelines
- Active Learning
- Bayesian Correction
- Recalibrate Thresholds
- Guardrails (in ML & Data Systems)
- Model KPIs (Key Performance Indicators)
- Lagging Indicators
- Leading Indicators
- Windows (in Time-Series)
- Autoencoder
- Frozen Encoder
- Embedding
- Representation Shift
- Classifier Two-Sample Tests (C2STs)
- Energy Distance
- Maximum Mean Discrepancy (MMD)
- Cardinality in Categorical Data
- Categorical Drift
- Cramér’s V
- Macro Shifts
- Categorical Explosions
- Cohort
- Off-Distribution
- Discriminatory Power
- KS Statistic (Kolmogorov–Smirnov Statistic)
- Model Stability
- Feature Values
- Four-Fifths (80%) Rule
- SLI (Service Level Indicator)
- ROI (Return on Investment)
- Treatment Cost
- Incremental Revenue
- Incremental Recovery Rate (IRR)
- Incremental Sales
- Random Targeting Strategy
- Causal ML (Causal Machine Learning)
- Cumulative Uplift
- Population Proportion
- Incremental Gain
- Total Incremental Benefit (TIB)
- Cumulative Incremental Gain (CIG)
- Qini Curve
- Uplift Score
- Uplift Models
- Ops Health Dashboard
- SLA Breach Rate
- SLA (Service Level Agreement)
- Supplier Constraints
- Long Lead Times
- Slow-Moving SKUs
- SKU
- Real-Time Inventory Tracking
- Supplier Management
- Demand Forecasting
- Reorder Point (ROP) Optimization
- Safety Stock
- Backorder Rate
- Lost Sales Value
- Fill Rate
- Stockout Rate
- Prophet — Time Series Forecasting by Facebook (Meta)
- LSTM — Long Short-Term Memory Networks
- ARIMA (AutoRegressive Integrated Moving Average)
- Return Distribution
- Value-at-Risk (VaR)
- Risk Forecast
- Probabilistic Scoring
- Full Distribution
- Continuous Probabilistic Forecasts
- Classification Probability
- Quantile Forecasts
- Point Forecasts
- Strictly Proper Scoring Rules
- Probability Forecasts
- Target Variable
- Probability Density
- Normal Distribution
- Probability Mass
- Probability Distribution
- Probabilistic Forecasts
- Deterministic forecasts
- Cumulative Distribution Function (CDF)
- M-Competitions (Makridakis Competitions)
- Forecasting Benchmarks
- Average Absolute Error (AAE)
- Seasonal Lag
- Simple Baseline Methods
- Naïve Baseline Forecast
- Forecast Error
- Forecasting Competitions
- Predicting Percentiles
- Prediction Intervals (PI)
- Quantile Regression
- Quantile Level
- Time Series Forecasting
- Log-Space
- Relative accuracy
- R² (R-squared)
- Long-Tail Items
- Self-Information of Popularity
- Relevance in Recommender Systems
- Genre Overlap
- Jaccard index
- Cosine Similarity of Item Features
- Intra-List Diversity (ILD)
- Dominating in Recommender Systems
- Catalog Coverage
- User Coverage
- Item Coverage
- Diminishing Utility
- DCG (Discounted Cumulative Gain)
- Kaggle
- TREC (Text REtrieval Conference)
- Adaptive ECE (Expected Calibration Error with Adaptive Binning)
- Maximum Calibration Error (MCE)
- ROC Curve (Receiver Operating Characteristic)
- Murphy’s Decomposition
- Temperature Scaling
- Platt Scaling
- Isotonic Regression
- Support Vector Machines (SVMs)
- Underconfident
- Overconfident
- Confidence Level
- Risk-Based Decisions
- Neural Networks
- Binary Cross-Entropy (BCE)
- Loss Functions
- Underflow
- Logit Space
- Logistic Regression
- Binary Classification
- Classification Models
- Log-Odds
- Softmax Function
- Sigmoid Function
- Squashing Function
- Conversion Rate (CR)
- Cost-Per-Click (CPC) Models
- Causal Trees
- Uplift Random Forests
- Uplift Curve
- Likelihood
- Correlation
- Causal Effect
- Outlier
- Mean Squared Error (MSE)
- Regression Models
- One-vs-Rest (OvR)
- Multiclass Classification
- Partial AUC (pAUC)
- Micro AUC
- Macro AUC
- Median
- Mean
- Sensitivity in Feature Engineering
- Encode (in Feature Engineering)
- Normalize (in Feature Engineering)
- Embedding Similarity
- Computer Vision (CV)
- Natural Language Processing (NLP)
- Accuracy
- Chi-square (χ²) Test
- Kolmogorov–Smirnov (KS) Test
- Jensen–Shannon (JS) Divergence
- Kullback–Leibler (KL) Divergence
- Statistical Tests
- Seasonality
- Concept Drift
- Data Drift
- Fair Lending laws
- Basel III
- High-Stakes Domains
- Deep Ensembles
- Counterfactual Explanations
- LIME (Local Interpretable Model-agnostic Explanations)
- SHAP (SHapley Additive exPlanations)
- Post-hoc Explainability
- Decision Trees
- Linear Models
- Caching
- Quantization
- ONNX (Open Neural Network Exchange)
- Full Annotation
- Weak Supervision
- TPU Clusters
- Statistical Power
- Drift Guardrails
- Latency Guardrails
- Fairness Guardrails
- DeLong’s Test
- Dataset Shift
- Label Noise
- Evaluation Set
- Clopper–Pearson Interval
- Wilson Score Interval
- Per-class Precision (sometimes called class-wise precision)
- Multiclass Precision
- Multilabel Precision
- Weighted Averaging
- Harmonic Mean
- F1-score
- Model Score
- Bootstrap
- Average Precision (AP)
- Upsampling
- Downsampling
- Micro Averaging
- Macro Averaging
- AUC (Area Under the Curve)
- Fairness parity
- LTV (Customer Lifetime Value)
- CAC (Customer Acquisition Cost)
- Bayesian Inference.
- Sequential Testing (also called sequential analysis)
- Confidence Intervals (CIs)
- Power Analysis
- Interleaving Tests
- A/B Testing
- Time-based splits (a.k.a. Temporal Cross-Validation, Rolling Window Validation)
- k-fold Stratified Cross-Validation (Stratified CV)
- Compute budgets
- Manual review minutes
- Inference Cost (Inference $)
- Label Drift (a.k.a. Target Drift)
- Covariate Drift (a.k.a. Covariate Shift)
- KS shift (Kolmogorov–Smirnov shift)
- PSI (Population Stability Index)
- Selection Rate
- SLOs (Service Level Objectives)
- Cannibalization
- Revenue net of treatment cost
- Incremental Conversions
- Uplift@k
- AUUC (Area Under the Uplift Curve)
- Qini Coefficient
- Crew Overtime
- SLA Breaches
- Overstock %
- Stockouts
- Continuous Ranked Probability Score (CRPS)
- MASE (Mean Absolute Scaled Error)
- Pinball Loss (a.k.a. Quantile Loss)
- WMAPE (Weighted Mean Absolute Percentage Error)
- sMAPE (Symmetric Mean Absolute Percentage Error)
- RMSLE (Root Mean Squared Logarithmic Error)
- Mean Absolute Error (MAE)
- Novelty (in Recommender Systems)
- Diversity (in Recommender Systems)
- Coverage
- Hit Rate (HR)
- NDCG (Normalized Discounted Cumulative Gain)
- Mean Average Precision (MAP)
- Expected Calibration Error (ECE)
- Reliability Curves (also called Calibration Curves)
- Log Loss (also called Logarithmic Loss or Cross-Entropy Loss)
- Brier Score
- Calibration quality (Model Calibration)
- Logits
- CTR (Click-Through Rate)
- WAPE (Weighted Absolute Percentage Error)
- Recall
- Uplift
- Mean Absolute Percentage Error (MAPE)
- Root Mean Squared Error (RMSE)
- ROC-AUC (Receiver Operating Characteristic – Area Under Curve, = AUROC)
- Baseline Heuristics
- Precision (a.k.a. Positive Predictive Value, PPV)
- Precision–Recall AUC (PR-AUC)
- Advanced Sorting in Spreadsheets
- Terminology
- What Are Time Series, and How Are They Used?
- Getting Started with R
- A Gentle Introduction to Stationarity
- Weak and Strong Stationarity
- Linear Processes
- Understanding ARMA Processes
- Computing ACFs of Causal AR(2) Processes Using Difference Equations
- Understanding ACFs via Difference Equations for AR(p) and ARMA(p, q)
- Best Linear Predictor of a Stationary Process
- Sample ACF and Sample PACF
- Preliminary Estimation for AR Models and the Yule–Walker Equations
- Maximum Likelihood Estimation for ARMA Models (Gaussian MLE)
- Diagnostics After Fitting a Time Series Model
- Order Selection for Time Series Models
- ARIMA Models: How Nonstationary Models Are Built from Stationary Ones
- SARIMA Models: Seasonal ARIMA
- Beyond One-Step Ahead Predictions
- Exponential Smoothing Models
- Time Series