Building Neural Networks with torch.nn in PyTorch

The first `nn.Linear` layer maps two input features to eight hidden values. I find the two feature counts a clear way to see what this layer changes.

You’ll place that layer in a model structure that can hold its parameters and connect it to another layer. Next, see how `torch.nn` contributes as you define the classifier.

What torch.nn contributes to a model

torch.nn is PyTorch’s namespace for neural-network modules, layers, containers, activation functions, and loss functions. nn.Module is the base class for a custom model, while torch.nn itself is not a model class.

Assigning a layer to a module attribute registers its parameters with the model, so model.parameters() can pass them to an optimizer. The optimizer lives in torch.optim, a separate namespace covered in this torch.optim guide.

Registration also lets a parent module move its registered parameters and buffers together and include them in its state dictionary. A plain tensor attribute is not a learnable parameter unless you register it as an nn.Parameter.

ComponentWhat it doesHas learnable state by default
nn.ModuleBase class for a model or reusable model componentOnly when you add parameters or child modules
nn.LinearMaps the final input dimension to a new feature dimensionYes, weight and bias by default
torch.nn.functionalProvides function calls such as activations and tensor operationsNo module-owned parameters
torch.optimUpdates registered parameters using their gradientsOptimizer state is separate from the model

I chose a small Linear model rather than a CNN because its feature-to-logit shape is visible without introducing image dimensions. Use a module when a layer needs registered parameters, while a torch.nn.functional operation suits a step with no module-owned state, as the PyTorch forum discussion about the two APIs illustrates.

nn.Sequential is itself a module container that calls child modules in order, which is why the example can store both Linear layers under one attribute.

Check the inputs before defining layers

Each row is one example, and the last dimension holds its input features. A Linear layer replaces that feature dimension with out_features, which means the batch dimension stays unchanged.

Linear applies an affine transformation to the final dimension with a weight matrix and optional bias, as its reference defines.

For a two-class classifier, the target stores one integer class ID per example. CrossEntropyLoss compares those IDs with two raw scores per row, so do not apply Softmax before the loss, as its reference specifies.

Here Linear carries six feature rows from (6, 2) to (6, 8) and back to (6, 2), so the batch stays at six while the feature count changes.

import torch
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"device: {device}")
TensorShape and typeMeaning
features(batch, 2), float32Two input values for each example
targets(batch,), longOne class ID, 0 or 1, for each example
logits(batch, 2), float32Two unnormalized class scores per example

The device check selects CUDA when available or keeps the example on CPU, and the model, features, and labels must share that device before the forward pass.

Build and train a classifier in four steps

The example uses eight two-feature rows, with two reserved for validation. The same shapes work with larger data as long as each row still has two features.

Step 1: Match the input and class dimensions

The first Linear layer maps two input features to eight hidden values, then the second maps eight values to two class scores that match the class IDs.

Step 2: Register layers inside nn.Module

Classifier inherits from nn.Module and stores an nn.Sequential container as self.layers. The container registers both Linear layers, so model.parameters() includes their weights and biases without a separate parameter list.

Step 3: Connect logits, loss, and optimizer

CrossEntropyLoss takes the model’s raw logits and integer class labels. Each loop clears old gradients, calculates a loss, backpropagates it, and lets stochastic gradient descent adjust the registered parameters.

import torch
from torch import nn

torch.manual_seed(7)

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
features = torch.tensor(
    [
        [-2.0, -1.0],
        [-1.0, -2.0],
        [-1.5, -0.5],
        [-2.0, -2.0],
        [1.0, 1.0],
        [2.0, 1.0],
        [1.0, 2.0],
        [1.5, 0.5],
    ],
    dtype=torch.float32,
    device=device,
)
labels = torch.tensor([0, 0, 0, 0, 1, 1, 1, 1], dtype=torch.long, device=device)
train_indices = torch.tensor([0, 1, 2, 4, 5, 6], device=device)
validation_indices = torch.tensor([3, 7], device=device)
train_features = features[train_indices]
train_labels = labels[train_indices]
validation_features = features[validation_indices]
validation_labels = labels[validation_indices]


class Classifier(nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = nn.Sequential(
            nn.Linear(2, 8),
            nn.ReLU(),
            nn.Linear(8, 2),
        )

    def forward(self, x):
        return self.layers(x)


model = Classifier().to(device)
loss_fn = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.1)

with torch.no_grad():
    initial_loss = loss_fn(model(train_features), train_labels).item()

model.train()
for _ in range(120):
    optimizer.zero_grad()
    logits = model(train_features)
    loss = loss_fn(logits, train_labels)
    loss.backward()
    optimizer.step()

model.eval()
with torch.inference_mode():
    train_logits = model(train_features)
    validation_logits = model(validation_features)
    final_loss = loss_fn(train_logits, train_labels).item()
    train_correct = (train_logits.argmax(dim=1) == train_labels).sum().item()
    validation_correct = (
        validation_logits.argmax(dim=1) == validation_labels
    ).sum().item()

parameter_count = sum(parameter.numel() for parameter in model.parameters())
print(f"PyTorch: {torch.__version__}")
print(f"device: {device}")
print(f"input shape: {tuple(train_features.shape)}")
print(f"logits shape: {tuple(train_logits.shape)}")
print(f"validation logits shape: {tuple(validation_logits.shape)}")
print(f"trainable parameters: {parameter_count}")
print(f"cross-entropy: {initial_loss:.4f} -> {final_loss:.4f}")
print(f"training accuracy: {train_correct}/{len(train_labels)}")
print(f"validation accuracy: {validation_correct}/{len(validation_labels)}")

The two Linear layers register 42 trainable values through their weights and biases. Their matrix shapes are (8, 2) and (2, 8).

PyTorch accumulates gradients by default, so optimizer.zero_grad() clears the previous values before the next loss. loss.backward() computes parameter gradients and optimizer.step() applies them, as the PyTorch optimization tutorial explains.

Step 4: Evaluate without updating the weights

I kept the validation indices out of the optimizer loop. Those rows do not contribute gradients or weight updates, and model.eval() changes layers such as Dropout and BatchNorm while torch.inference_mode() skips gradient tracking.

Terminal output showing PyTorch classifier input and logits shapes, loss, and train and validation counts
The six training examples produce six two-class logit rows. Validation examples stay outside the optimizer updates.

Cross-entropy falls from 0.6426 to 0.0092, and predictions match all six training labels and both validation labels in this toy example.

I used matching validation indices for features and labels, so the 2/2 count checks target alignment as well as output shape. That two-row holdout still cannot estimate performance on a representative dataset.

Read shape, target, and device errors

A PyTorch forum question asks, “How could i change my output size as i expected?” Linear changes the final feature dimension, while the batch dimension passes through unchanged.

SymptomWhat to inspectCorrection
Linear reports incompatible matrix shapesThe last dimension of the input tensorSet in_features to that dimension or reshape the feature data before the layer
CrossEntropyLoss rejects target values or shapeLogits should be (batch, classes). Class-index targets should be (batch,) with integer IDs from 0 through classes minus 1Use torch.long labels and keep one label for each input row
The model and input are on different devicesCheck the device for the model, features, and labelsMove each to the same device before calling the model
Validation changes the weightsCheck that validation batches are outside the optimizer loopUse model.eval() and torch.inference_mode() for evaluation

A new tensor or module is normally created on the CPU, not on a GPU. Selecting a device does not move objects that already exist, so move the model and all tensors before training when you choose an accelerator.

For a deeper loss breakdown, see this PyTorch loss-function reference. The complete PyTorch overview covers tensor operations that sit before the model.

A training score does not measure unseen data

Replace the hand-picked tensors with examples from your task and preserve the feature order. Keep representative validation rows out of optimizer updates, then run the same file.

python torch_nn_walkthrough.py

torch.nn questions that change the next step

torch.nn names a namespace, nn.Module is a base class, and nn.Linear creates a layer instance. Calling the model runs its forward pass, while calling the loss compares its scores with targets.

What is torch.nn in PyTorch?

torch.nn is PyTorch’s namespace for neural-network modules, layers, containers, activation functions, and losses. nn.Module is the base class for custom models.

Does torch.nn use the GPU by default?

No. New tensors and modules are normally created on the CPU. Move the model and its inputs to the same selected device to use an accelerator.

What does nn.Linear do?

nn.Linear applies an affine transformation to the last input dimension. Its in_features value must match that dimension, and its out_features value sets the output dimension.

Should I use nn.Module or torch.nn.functional?

Use an nn.Module for a parameterized layer that should register with the model. Use torch.nn.functional for an operation that does not need module-owned parameters.

Should I apply Softmax before CrossEntropyLoss?

No. CrossEntropyLoss expects unnormalized logits and class-index targets for this classification task, so pass the model scores directly.

Rishabh Das
Rishabh Das
Articles: 47