🚀 Supercharge your YouTube channel's growth with AI.
Try YTGrowAI FreeA Quick Guide to Pytorch Loss Functions

I keep running into the same problem when I train neural networks: my model learns something, but how do I actually measure whether it is learning the right thing? The loss function is the answer to that question. It is the single number that tells a model how far off its predictions are from the truth, and every weight update in the network tries to make that number smaller.
This article covers PyTorch’s built-in loss functions, from basic ones like MSELoss and CrossEntropyLoss to specialized losses like HuberLoss and TripletMarginLoss. By the end, you will know which loss to reach for in different training scenarios and how to wire them up in your training loop.
TLDR
- Loss functions measure how wrong a model’s predictions are
- Use MSELoss for regression, CrossEntropyLoss for multi-class classification
- BCEWithLogitsLoss combines sigmoid and binary cross-entropy in one numerically stable call
- Custom losses are just nn.Module subclasses with a forward method
- Pair NLLLoss with LogSoftmax, and use KL Divergence for distribution matching tasks
What are Loss Functions in Deep Learning?
A loss function takes the model’s predictions and the true labels, then boils them down to a single scalar value. During training, PyTorch computes this loss after each forward pass, then runs backpropagation to calculate gradients for every weight in the network. Those gradients tell each weight how much it should increase or decrease to reduce the loss on the next batch.
Smaller loss means better predictions. If the loss drops over time, the model is learning. If it plateaus or rises, something is wrong with the data, the learning rate, or the loss choice itself. The loss is the training signal, so picking the right one matters more than almost any other architectural decision.
PyTorch ships every common loss function under the torch.nn namespace. All of them inherit from nn.Module, which means they plug straight into the training loop just like any other layer.
Regression Losses: MSELoss and L1Loss
Regression problems predict continuous values. If your model outputs a raw number (house prices, temperatures, stock returns), you typically reach for one of two main losses.
L1Loss, also called Mean Absolute Error, computes the average absolute difference between predicted and true values. It is robust to outliers because it does not square the errors, so a single wildly wrong prediction does not dominate the loss.
import torch
import torch.nn as nn
criterion = nn.L1Loss()
predictions = torch.tensor([1.2, 3.4, 2.1])
targets = torch.tensor([1.0, 3.0, 2.0])
loss = criterion(predictions, targets)
print(loss)
tensor(0.2333)
MSELoss, or Mean Squared Error, squares each error before averaging. This punishes larger errors much more than smaller ones, which can help when you want the model to be confident in its predictions. The trade-off is that outliers in your training data can inflate the loss dramatically. Both L1Loss and SmoothL1Loss (which behaves like L1 for large errors and L2 for small ones) are also available in torch.nn.
import torch
import torch.nn as nn
criterion = nn.MSELoss()
predictions = torch.tensor([1.2, 3.4, 2.1])
targets = torch.tensor([1.0, 3.0, 2.0])
loss = criterion(predictions, targets)
print(loss)
tensor(0.0700)
Classification Losses: CrossEntropyLoss, BCEWithLogitsLoss, NLLLoss
Classification tasks group predictions into discrete categories. The right loss depends on how many classes you have and how the model outputs its predictions.
CrossEntropyLoss is the workhorse for any classification task with more than two classes. It combines LogSoftmax and NLLLoss into a single call, which is both more numerically stable and more convenient than chaining them manually. It expects raw logits as input, not probabilities. Under the hood, cross-entropy loss measures how close the model’s predicted probability distribution is to the one-hot ground truth distribution.
import torch
import torch.nn as nn
criterion = nn.CrossEntropyLoss()
logits = torch.tensor([[2.0, 0.5, -1.0], [-0.5, 1.5, 0.0]])
targets = torch.tensor([0, 2])
loss = criterion(logits, targets)
print(loss)
tensor(1.0238)
For binary classification, BCEWithLogitsLoss takes the raw sigmoid output and computes binary cross-entropy in one step. Passing logits directly to this loss is more numerically stable than passing pre-sigmoid probabilities because the internal sigmoid computation is fused with the loss calculation.
import torch
import torch.nn as nn
criterion = nn.BCEWithLogitsLoss()
logits = torch.tensor([1.5, -0.8, 3.2])
targets = torch.tensor([1.0, 0.0, 1.0])
loss = criterion(logits, targets)
print(loss)
tensor(0.2042)
NLLLoss expects log-probabilities (the output of LogSoftmax) as input. It is usually paired with a LogSoftmax layer in the model’s forward pass rather than being used standalone. The loss decreases when the model assigns a higher probability to the correct class, making it a natural fit for multi-class problems when you want to apply a specific temperature or regularization to the softmax.
import torch
import torch.nn as nn
criterion = nn.NLLLoss()
log_probs = torch.tensor([[-0.2, -1.5, -0.5], [-0.8, -0.1, -2.0]])
targets = torch.tensor([0, 1])
loss = criterion(log_probs, targets)
print(loss)
tensor(0.1500)
KL Divergence and Advanced Losses
Beyond basic regression and classification, PyTorch provides losses for more specialized training scenarios.
KL Divergence measures how one probability distribution diverges from a reference distribution. KLDivLoss expects the model output in log-probability form. It shows up in