Logistic Regression – Loss Function and Cost Function#
Stage 2 · 🔵 Logistic Regression as a Neuron · Lesson 07 of 17 · beginner
◀ Previous · Logistic Regression (Binary Classification Model) · Next · Gradient Descent in Logistic Regression ▶ · ↑ Section
Important
✨ AI-generated content. This page was written with the assistance of an AI language model and is provided as a learning aid. Despite careful review, it may still contain mistakes, omissions, or out-of-date information. Whether you are new to the topic, a team lead, or a senior practitioner, treat it as a starting point rather than an authoritative reference: read it critically and independently verify anything you act on (code, commands, figures, and factual claims) against official documentation and primary sources before relying on it.
Measuring one prediction#
To learn, the neuron needs a score of how wrong it is. A loss function \(\mathcal{L}(\hat{y}, y)\) measures the error on a single example — comparing the predicted probability \(\hat{y}\) against the true label \(y\). Training then adjusts \(\mathbf{w}\) and \(b\) to make that loss small.
Why not squared error#
The obvious choice, squared error \(\tfrac{1}{2}(\hat{y} - y)^2\), works for linear regression but is a poor fit here. Because \(\hat{y}\) passes through the sigmoid, squared error makes the cost non-convex — a bumpy surface with many local minima where gradient descent can get stuck. We want a loss whose surface is a single smooth bowl.
The cross-entropy loss#
The right choice is the cross-entropy (log) loss:
Read it by cases. If \(y = 1\) it reduces to \(-\log \hat{y}\), which is small only when \(\hat{y}\) is close to 1; if \(y = 0\) it reduces to \(-\log(1 - \hat{y})\), small only when \(\hat{y}\) is close to 0. Confident wrong answers are punished hard — and for logistic regression this loss is convex. (It is exactly the negative log-likelihood of a Bernoulli label.)
From loss to cost#
Loss scores one example; the cost function averages it over all \(m\):
Loss is per-example; cost is the whole-set average. Training means finding the \(\mathbf{w}, b\) that minimise \(J\) — the job of the next lesson.
Hint
Related lessons: Logistic Regression (Binary Classification Model) · Gradient Descent in Logistic Regression · Logistic Regression Gradient Descent · Binary Classification and Logistic Regression (Neural Network Basics)
See also
Source article Adapted (context, re-expressed) in our own words from: https://insightful-data-lab.com/2025/04/07/logistic-regression-loss-function-and-cost-function/ (insightful-data-lab.com).