PyTorch Autograd Automatic Differentiation

The training of deep learning is essentially a process of repeatedly computing gradients and updating parameters.

Manually deriving the gradients for each layer is tedious and error-prone. PyTorch'sAutograd(automatic differentiation) engine was created to solve this problem—it canautomatically compute the gradients of any computation graph, allowing you to focus on model design rather than calculus derivation.


Core Concepts

1. What is Automatic Differentiation

Automatic Differentiation is not numerical differentiation (finite difference method), nor symbolic differentiation (algebraic derivation), but ratherrecording the computation process and applying the chain rule in reverse step by stepto compute derivatives precisely.

PyTorch's Autograd uses adynamic computation graph(Define-by-Run) approach: each forward pass constructs a directed acyclic graph (DAG) in real time, recording every operation and its inputs and outputs; during backpropagation, it traverses the graph in reverse and computes the gradient of each node in order.

2. requires_grad Attribute

The Tensor'srequires_gradattribute controls whether gradients need to be tracked for that tensor:

Example

import torch

# Create a tensor that requires gradient tracking (default requires_grad=False)
x = torch.tensor(3.0, requires_grad=True)
print(x)              # tensor(3., requires_grad=True)
print(x.requires_grad)  # True

# You can also modify it after creation
y = torch.tensor(2.0)
print(y.requires_grad)   # False
y.requires_grad_(True)   # In-place modification (note the trailing underscore)
print(y.requires_grad)   # True

# Results of operations involving tensors with requires_grad=True automatically inherit requires_grad=True
z = x * y
print(z.requires_grad)   # True

3. grad_fn and Computation Graph

Every tensor produced by an operation records agrad_fn, which points to the operation node that created it. This is the "skeleton" of the computation graph:

Example

import torch

x = torch.tensor(2.0, requires_grad=True)
y = torch.tensor(3.0, requires_grad=True)

z = x ** 2 + y * 3    # z = x² + 3y

print(z)           # tensor(13., grad_fn=<AddBackward0>)
print(z.grad_fn)   # <AddBackward0 object>

# Trace the chain of operations that created z
print(z.grad_fn.next_functions)
# ((<PowBackward0 object>, 0), (<MulBackward0 object>, 0))
# You can see that z is composed of a power operation and a multiplication operation

backward() Backpropagation

1. Calling backward() for Scalar Output

Call.backward()on the final scalar (loss value), and Autograd will automatically compute the gradients of all leaf nodes in reverse along the computation graph, storing the results in each tensor's.gradattribute:

Example

import torch

x = torch.tensor(2.0, requires_grad=True)
y = torch.tensor(3.0, requires_grad=True)

# Forward pass: z = x² + 3y
z = x ** 2 + y * 3

# Backward pass: automatically compute dz/dx and dz/dy
z.backward()

# View gradients
print(x.grad)   # tensor(4.)  ← dz/dx = 2x = 2×2 = 4
print(y.grad)   # tensor(3.)  ← dz/dy = 3

# Mathematical verification:
# z = x² + 3y
# dz/dx = 2x = 2×2 = 4  ✓
# dz/dy = 3             ✓

2. Accumulation Issue with Multiple backward() Calls

Autograd gradients areaccumulated, not overwritten. Each timebackward()is called, the gradients will be added to.gradthe existing values. This is the most common pitfall in training loops:

Example

import torch

x = torch.tensor(2.0, requires_grad=True)

# First backpropagation
loss = x ** 2
loss.backward()
print(x.grad)    # tensor(4.)  ← dL/dx = 2x = 4

# Second backpropagation (without zeroing!)
loss = x ** 2
loss.backward()
print(x.grad)    # tensor(8.) ← Accumulated! Not 4, but 4+4=8

# ✅ Correct approach: zero the gradients before each backpropagation
x.grad.zero_()   # Zero in-place (note the trailing underscore)
loss = x ** 2
loss.backward()
print(x.grad)    # tensor(4.) ← Correct

When training a neural network, before eachbackward()calloptimizer.zero_grad()to zero the gradients, otherwise gradients will keep accumulating and cause incorrect parameter updates.

3. Calling backward(gradient) for Non-Scalar Output

If the output is a vector or matrix rather than a scalar,backward()you need to pass agradientparameter with the same shape as the output (i.e., the "upstream gradient"), which is essentially computing the vector-Jacobian product (VJP):

Example

import torch

x = torch.tensor([1.0, 2.0, 3.0], requires_grad=True)

# Forward pass: y is a vector
y = x ** 2   # y = [1, 4, 9]

# For non-scalar output, you must pass the gradient parameter (shape same as y)
# The gradient can be understood as "the gradient of the loss with respect to y"
y.backward(gradient=torch.ones_like(y))   # Assume the upstream gradient is all 1s

print(x.grad)
# tensor([2., 4., 6.]) ← dy/dx = 2x, computed element-wise

# If the upstream gradient is not all 1s (e.g., weighted)
x.grad.zero_()
y.backward(gradient=torch.tensor([1.0, 0.5, 2.0]))  # Different weights
# Actual computation: x.grad = 2x * gradient = [2×1, 4×0.5, 6×2]
print(x.grad)
# tensor([2., 2., 12.])

# More common approach: first use sum/mean to convert to a scalar, then call backward()
x.grad.zero_()
loss = (x ** 2).sum()   # Aggregate the vector into a scalar
loss.backward()
print(x.grad)
# tensor([2., 4., 6.]) ← Equivalent to the first approach

torch.no_grad() Stopping Gradient Tracking

During model inference (prediction), there is no need to compute gradients. Usingtorch.no_grad()can skip the construction of the computation graph, significantly saving memory and computation:

Example

import torch

x = torch.tensor(3.0, requires_grad=True)

# In a no_grad context, no operations will track gradients
with torch.no_grad():
    y = x ** 2
    print(y.requires_grad)  # False ← No longer tracking gradients
    print(y.grad_fn)        # None ← No computation graph node

# After exiting the no_grad context, normal tracking resumes
z = x ** 2
print(z.requires_grad)  # True

# Common use: wrap the entire inference process during model evaluation
model = torch.nn.Linear(10, 1)
inputs = torch.randn(32, 10)

with torch.no_grad():
    outputs = model(inputs)   # No computation graph is built; faster speed and lower memory usage

@torch.no_grad() Decorator Syntax

You can also use the decorator form, which is suitable for marking an entire inference function as no-gradient:

Example

import torch
import torch.nn as nn

model = nn.Linear(10, 1)

@torch.no_grad()
def predict(model, x):
    """Inference function, no gradient computation needed"""
    return model(x)

x = torch.randn(5, 10)
output = predict(model, x)
print(output.requires_grad)   # False

detach() Separating from the Computation Graph

.detach()Returns a new tensor that shares data with the original tensor but does not track gradients. Commonly used in the following scenarios:

Scenario Description
Converting intermediate results to numpy arrays numpy does not support tensors with gradients; you must detach() first
Recording training loss (logs) Avoid keeping the entire computation graph to prevent memory leaks
Freezing gradient propagation for part of the network Scenarios such as GAN training and transfer learning

Example

import torch

x = torch.tensor([1.0, 2.0, 3.0], requires_grad=True)
y = x ** 2 + x * 3   # y has grad_fn

# The tensor after detach shares data with y but is detached from the computation graph
y_detached = y.detach()
print(y_detached.requires_grad)  # False
print(y_detached.grad_fn)        # None

# ✅ Convert to numpy (tensors with gradients cannot be directly converted)
# y.numpy() # ❌ Error: RuntimeError
y_detached.numpy()    # ✅ Normal

# ✅ When recording loss values, detach (to avoid retaining the computation graph and consuming memory)
losses = []
for i in range(3):
    loss = (x ** 2).sum()
    losses.append(loss.detach().item())  # .item() converts a scalar tensor to a Python float
    loss.backward()
    x.grad.zero_()

print(losses)   # [14.0, 14.0, 14.0]

retain_graph Retaining the Computation Graph

By default,backward()after execution, the computation graph will beautomatically released(to save memory). If you need to backpropagate multiple times on the same computation graph (e.g., in some GAN training), you need to passretain_graph=True:

Example

import torch

x = torch.tensor(2.0, requires_grad=True)
y = x ** 3   # y = x³

# First backward (retain computation graph)
y.backward(retain_graph=True)
print(x.grad)    # tensor(12.)  ← dy/dx = 3x² = 3×4 = 12

# Second backward (computation graph still exists)
x.grad.zero_()
y.backward(retain_graph=True)
print(x.grad)    # tensor(12.) ← Same result

# No need to retain last time
x.grad.zero_()
y.backward()     # After this, the computation graph is released
print(x.grad)    # tensor(12.)

# Attempting backward again will raise an error (computation graph already released)
# y.backward()   # &#x274c; RuntimeError: Trying to backward through the graph a second time

Unnecessary useretain_graph=Truewill cause memory to keep growing, because the computation graph cannot be released. Only use it when multiple backward passes are actually needed.


Application of Gradients in Neural Network Training

The following is a complete example of manually implementing gradient descent using Autograd, demonstrating the full workflow of Autograd in actual training:

Example

import torch

# Construct training data: y = 2x + 1 plus noise
torch.manual_seed(42)
X = torch.randn(100, 1)
y_true = 2 * X + 1 + 0.1 * torch.randn(100, 1)

# Initialize model parameters (requires gradient tracking)
w = torch.zeros(1, requires_grad=True)   # Weight
b = torch.zeros(1, requires_grad=True)   # Bias

lr = 0.1    # Learning rate
epochs = 50  # Number of training epochs

for epoch in range(epochs):
    # 1. Forward propagation: compute predictions
    y_pred = X * w + b

    # 2. Compute loss (mean squared error)
    loss = ((y_pred - y_true) ** 2).mean()

    # 3. Backward propagation: automatically compute d(loss)/dw and d(loss)/db
    loss.backward()

    # 4. Manually update parameters (wrap with no_grad to avoid update operations being tracked in the computation graph)
    with torch.no_grad():
        w -= lr * w.grad
        b -= lr * b.grad

    # 5. Clear gradients (must be cleared before the next backward)
    w.grad.zero_()
    b.grad.zero_()

    if (epoch + 1) % 10 == 0:
        print(f"Epoch {epoch+1:3d} | Loss: {loss.item():.4f} | w={w.item():.3f}, b={b.item():.3f}")

print(f"\nTraining complete: w ≈ {w.item():.3f} (true value 2.0), b ≈ {b.item():.3f} (true value 1.0))

The execution result of the above code is similar to the following:

Epoch  10 | Loss: 0.1064 | w=1.587, b=0.805
Epoch  20 | Loss: 0.0281 | w=1.876, b=0.949
Epoch  30 | Loss: 0.0152 | w=1.954, b=0.983
Epoch  40 | Loss: 0.0128 | w=1.978, b=0.993
Epoch  50 | Loss: 0.0122 | w=1.987, b=0.997

训练完成:w ≈ 1.987(真实值 2.0),b ≈ 0.997(真实值 1.0)

Common API Quick Reference Table

API Function Common scenarios
tensor.requires_grad_(True) Enable gradient tracking in-place Enable Autograd for an existing tensor
loss.backward() Backward propagation, compute gradients of all leaf nodes Each iteration in the training loop
tensor.grad Access the gradient value of a tensor View or manually update parameters
tensor.grad.zero_() Clear gradients in-place Must clear before each backward()
torch.no_grad() Context manager, disable gradient tracking Inference phase, manual parameter updates
tensor.detach() Return a new tensor detached from the computation graph Convert to numpy, log records, freeze gradients
tensor.item() Convert a scalar tensor to a Python number Print loss values, record metrics
loss.backward(retain_graph=True) Retain the computation graph, allow multiple backward passes Scenarios such as GAN training that require multiple backward passes

Common Questions and Precautions

1. Difference between leaf nodes and non-leaf nodes

Onlyleaf nodes(i.e., tensors directly created by the user, not the result of operations) have their gradients saved in.grad. Non-leaf nodes produced by intermediate operations do not retain gradients by default (to save memory). If you need to view the gradient of an intermediate node, call.retain_grad():

Example

import torch

x = torch.tensor(2.0, requires_grad=True)   # Leaf node

# Intermediate node (non-leaf node)
y = x ** 2      # y is an intermediate node
y.retain_grad() # Explicitly declare to retain y's gradient

z = y * 3       # z is the final output
z.backward()

print(x.grad)   # tensor(12.) ← Leaf node, saved normally
print(y.grad)   # tensor(3.) ← Saved because of retain_grad()
# Intermediate nodes without retain_grad(), .grad is None

2. In-place operations may break the computation graph

Performing in-place operations on tensors that require gradient tracking (e.g.+=、.add_()) may prevent Autograd from correctly propagating backwards, so they should be avoided as much as possible:

Example

import torch

x = torch.tensor([1.0, 2.0], requires_grad=True)

# Dangerous: in-place operation on a leaf node
# x += 1 # May raise an error: a leaf Variable that requires grad has been used in an in-place operation

# Safe: use a non-in-place operation
y = x + 1   # Create a new tensor, do not modify x

# It is safe to perform in-place operations in a no_grad context (e.g., parameter updates)
with torch.no_grad():
    x += 0.01   # Used for manual parameter updates, this is safe

3. Only floating-point tensors support gradients

Integer types (e.g.torch.int64) tensors do not supportrequires_grad=True, only floating-point types (float32、float64、float16) can participate in automatic differentiation.

Other extensions