PyTorch GPU / CUDA Acceleration
The core operations of deep learning are large-scale matrix multiplication and element-wise operations. CPUs are designed to handle complex serial logic, typically with 8 to 64 cores; GPUs, on the other hand, have thousands of simple parallel cores, making them naturally suited for this kind of highly parallel numerical computation. PyTorch, through NVIDIA'sCUDA(Compute Unified Device Architecture) framework, invokes the GPU, which can speed up training by dozens or even a hundred times.
1. Differences between CPU and GPU
The improvement in GPU training speed mainly comes from two aspects: first, the parallel execution of a large number of identical computations; second, high-bandwidth GPU memory enables data transfer much faster than CPU memory.
For compute-intensive operations such as matrix multiplication, the acceleration effect is particularly significant.
| Comparison item | CPU | GPU(NVIDIA) |
|---|---|---|
| Number of cores | 8-64 large cores | Thousands of small cores (CUDA Cores) |
| Design goal | Low-latency serial processing | High-throughput parallel computing |
| Memory bandwidth | ~50~100 GB/s | ~500~3000 GB/s |
| Matrix multiplication speed | Baseline | 10x-100x faster |
| PyTorch interface | "cpu" |
"cuda" |
If using Apple Silicon (M1/M2/M3), PyTorch, through thempsbackend, supports Metal GPU acceleration. The usage is almost identical to CUDA; just change the device to"mps"。
2. Detecting the CUDA Environment
Before using the GPU, you need to confirm whether the current environment has a CUDA-enabled PyTorch installed.
Also, check whether there are any available GPU devices on the system.
Example
# Whether CUDA is supported
print("CUDA available:", torch.cuda.is_available())
# Number of GPU devices
print("GPU count:", torch.cuda.device_count())
# Index of the current default GPU
print("Current GPU:", torch.cuda.current_device())
# GPU model name
print("GPU model:", torch.cuda.get_device_name(0))
# PyTorch version and the CUDA version it was compiled with
print("PyTorch version:", torch.__version__)
print("CUDA version:", torch.version.cuda)
Example output:
CUDA 可用: True GPU 数量: 1 当前 GPU: 0 GPU 型号: NVIDIA GeForce RTX 4090 PyTorch 版本: 2.3.0+cu121 CUDA 版本: 12.1
2.1 Dynamically Selecting a Device (Recommended Approach)
Hardcoding in code"cuda"This will cause machines without a GPU to directly throw an error.
Example
# Method 1: Classic approach, best compatibility
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# Method 2: Recommended for PyTorch 2.0+, automatically supports CUDA / MPS / CPU
device = (
"cuda" if torch.cuda.is_available()
else "mps" if torch.backends.mps.is_available()
else "cpu"
)
print(f"Using device: {device}")
3. Moving Tensors Between Devices
Tensors in PyTorch are created on the CPU by default.
To use GPU computation, you need to explicitly move tensors to the GPU, or create them directly on the GPU.
3.1 Basic Moving Methods
Example
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# Create a CPU tensor
cpu_tensor = torch.tensor([1.0, 2.0, 3.0])
print(cpu_tensor.device) # cpu
# Method 1: .to(device) — recommended, most versatile
gpu_tensor = cpu_tensor.to(device)
# Method 2: .cuda() — CUDA environment only
gpu_tensor = cpu_tensor.cuda()
# Method 3: Specify the device directly when creating
gpu_tensor = torch.tensor([1.0, 2.0, 3.0], device=device)
gpu_tensor = torch.randn(3, 4, device=device)
# Move back to CPU (for printing, numpy conversion, saving, etc.)
back_to_cpu = gpu_tensor.cpu()
print(gpu_tensor.device) # cuda:0
print(back_to_cpu.device) # cpu
When converting a GPU tensor to numpy, you must first move it back to the CPU, and if the tensor has gradients, you also need to detach it first:
Example
arr = gpu_tensor.cpu().numpy()
# Convert a GPU tensor with gradients to numpy
arr = gpu_tensor.detach().cpu().numpy()
3.2 Device Consistency Constraint
Tensors on different devices cannot directly participate in the same operation.
Otherwise, it will throwRuntimeError:
Example
b = torch.randn(3) # On CPU
# c = a + b # RuntimeError: Expected all tensors to be on the same device
# Correct approach: unify the device first
c = a + b.to("cuda")
Check the device where the tensor is located:
Example
print(x.device) # cuda:0
print(x.is_cuda) # True
print(x.get_device()) # 0 (GPU index)
3.3 Speed Comparison Verification
Example
import time
device = torch.device("cuda")
n = 5000
# CPU matrix multiplication
a_cpu = torch.randn(n, n)
b_cpu = torch.randn(n, n)
start = time.time()
c_cpu = torch.matmul(a_cpu, b_cpu)
print(f"CPU time: {time.time() - start:.3f}s")
# GPU matrix multiplication
a_gpu = a_cpu.to(device)
b_gpu = b_cpu.to(device)
torch.cuda.synchronize() # Make sure data transfer is complete before starting the timer
start = time.time()
c_gpu = torch.matmul(a_gpu, b_gpu)
torch.cuda.synchronize() # Wait for the GPU to finish execution before stopping the timer
print(f"GPU time: {time.time() - start:.3f}s")
Example output:
CPU 耗时: 1.847s GPU 耗时: 0.021s
GPU computation isasynchronously executed—after the Python call returns, the GPU operation may not be complete. When timing, you must call
torch.cuda.synchronize()to wait for the GPU to actually finish; otherwise, the measurement results are inaccurate.
4. Moving the Model to GPU
All parameters of the model (weight、bias) are essentially tensors.
These parameters also need to be moved to the GPU in order to perform forward and backward propagation on the GPU.
Call.to(device)on the entire model, and PyTorch will automatically traverse and move all internal parameters.
Example
import torch.nn as nn
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
class SimpleNet(nn.Module):
def __init__(self):
super().__init__()
self.net = nn.Sequential(
nn.Linear(784, 256),
nn.ReLU(),
nn.Linear(256, 128),
nn.ReLU(),
nn.Linear(128, 10),
)
def forward(self, x):
return self.net(x)
# Move the model to GPU, only need to call once
model = SimpleNet().to(device)
# Verify that all parameters are on the GPU
for name, param in model.named_parameters():
print(f"{name}: {param.device}")
# net.0.weight: cuda:0
# net.0.bias: cuda:0
# ...
# Input data must also be on the same device
x = torch.randn(32, 784).to(device)
output = model(x)
print(output.shape) # torch.Size([32, 10])
When the model is on the GPU but the input data is still on the CPU, the forward pass will raise an error. Be sure to, after DataLoader reads each batch,
inputsandlabelscall .to(device) on all of them..to(device)。
5. Complete Training Workflow
The following is a standard GPU training template that includes data loading, model training, and validation evaluation.
Example
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
# 1. Device configuration
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Training device: {device}")
# 2. Data loading
# pin_memory=True: Locks data in memory to speed up CPU -> GPU transfer
# num_workers: Multiprocess preloading to reduce data waiting time
transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.5,), (0.5,)),
])
train_dataset = datasets.MNIST(root="./data", train=True,
download=True, transform=transform)
test_dataset = datasets.MNIST(root="./data", train=False,
download=True, transform=transform)
train_loader = DataLoader(train_dataset, batch_size=256, shuffle=True,
num_workers=4, pin_memory=True)
test_loader = DataLoader(test_dataset, batch_size=256, shuffle=False,
num_workers=4, pin_memory=True)
# 3. Define the model
class CNN(nn.Module):
def __init__(self):
super().__init__()
self.features = nn.Sequential(
nn.Conv2d(1, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
)
self.classifier = nn.Sequential(
nn.Flatten(),
nn.Linear(64 * 7 * 7, 256), nn.ReLU(), nn.Dropout(0.5),
nn.Linear(256, 10),
)
def forward(self, x):
return self.classifier(self.features(x))
model = CNN().to(device) # Move the model to GPU
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=1e-3)
# 4. Training function
def train_epoch(model, loader, optimizer, criterion):
model.train()
total_loss, correct = 0.0, 0
for inputs, labels in loader:
# non_blocking=True: Asynchronous transfer, CPU can continue preparing the next batch of data
inputs = inputs.to(device, non_blocking=True)
labels = labels.to(device, non_blocking=True)
optimizer.zero_grad()
outputs = model(inputs)
loss = criterion(outputs, labels)
loss.backward()
optimizer.step()
total_loss += loss.item() * inputs.size(0)
correct += (outputs.argmax(1) == labels).sum().item()
n = len(loader.dataset)
return total_loss / n, correct / n
# 5. Validation function
def eval_epoch(model, loader, criterion):
model.eval()
total_loss, correct = 0.0, 0
with torch.no_grad():
for inputs, labels in loader:
inputs = inputs.to(device, non_blocking=True)
labels = labels.to(device, non_blocking=True)
outputs = model(inputs)
loss = criterion(outputs, labels)
total_loss += loss.item() * inputs.size(0)
correct += (outputs.argmax(1) == labels).sum().item()
n = len(loader.dataset)
return total_loss / n, correct / n
# 6. Main training loop
for epoch in range(1, 11):
train_loss, train_acc = train_epoch(model, train_loader, optimizer, criterion)
val_loss, val_acc = eval_epoch(model, test_loader, criterion)
print(f"Epoch {epoch:02d} | "
f"Train Loss: {train_loss:.4f}, Acc: {train_acc:.4f} | "
f"Val Loss: {val_loss:.4f}, Acc: {val_acc:.4f}")
Key points:
deviceVariables are declared at the very top of the code, referenced globally, never hardcoded"cuda"- Model
.to(device)only needs to be called once - For each batch,
inputsandlabelsmust also be.to(device) - DataLoader settings
pin_memory=Trueandnon_blocking=Trueaccelerate data transfer - Use in validation phase
torch.no_grad()disable gradients to save memory
6. Multi-GPU Training
When a single GPU's memory is insufficient, consider using multi-GPU parallel training.
In addition, multi-GPU can further improve training speed. PyTorch provides two main approaches.
6.1 DataParallel
This is the simplest multi-GPU approach.
It uses a single process, splits each batch evenly across GPUs, runs forward propagation in parallel, and aggregates gradients on the main GPU for updates. It is easy to use, but the main GPU is heavily loaded, utilization across GPUs is uneven, and it is suitable for quick starts.
Example
import torch.nn as nn
model = CNN()
if torch.cuda.device_count() > 1:
print(f"Using {torch.cuda.device_count()} GPUs")
model = nn.DataParallel(model)
# You can also specify which GPUs to use
# model = nn.DataParallel(model, device_ids=[0, 1])
model = model.to("cuda")
# The subsequent training code is exactly the same as single-GPU
# DataParallel automatically splits the batch evenly across GPUs and aggregates the results
If you need to access the original model's attributes (such as custom methods), you need to.moduleaccess them via:
Example
print(model.module.classifier)
# When saving the model, it is recommended to save model.module for easy single-GPU loading
torch.save(model.module.state_dict(), "model.pth")
6.2 DistributedDataParallel
This is the multi-process approach.
Each process is bound to one GPU, each card holds a full model copy, and gradients are synchronized via All-Reduce. It is the recommended solution for production environments and large-scale training, with far higher efficiency than DataParallel.
Example
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.utils.data.distributed import DistributedSampler
def main(rank, world_size):
# Initialize process group; the NCCL backend is optimized for NVIDIA GPUs
dist.init_process_group(backend="nccl", rank=rank, world_size=world_size)
# Bind each process to its corresponding GPU
torch.cuda.set_device(rank)
device = torch.device(f"cuda:{rank}")
# Wrap with DDP after moving the model to the corresponding GPU
model = CNN().to(device)
model = DDP(model, device_ids=[rank])
# Use DistributedSampler in DataLoader to ensure no data overlap across GPUs
sampler = DistributedSampler(train_dataset,
num_replicas=world_size, rank=rank)
loader = DataLoader(train_dataset, batch_size=64,
sampler=sampler, pin_memory=True)
# Training logic is identical to single-GPU
# ...
dist.destroy_process_group()
# How to launch (recommended: torchrun):
# torchrun --nproc_per_node=4 train_ddp.py
| Comparison item | DataParallel | DistributedDataParallel |
|---|---|---|
| Number of processes | Single process | Multi-process (one per GPU) |
| Communication backend | Python GIL limitation | NCCL (efficient) |
| Main GPU load | Heavy (gradient aggregation) | Balanced (All-Reduce) |
| Code changes | Minimal | Moderate |
| Applicable scenarios | Rapid experimentation | Production training |
7. Mixed Precision Training AMP
Training uses float32 (FP32) precision by default.
Automatic Mixed Precision (AMP)Uses float16 (FP16) or bfloat16 (BF16) for some computations, yielding significant benefits with almost no loss of precision:
- Memory usage reduced by about 50%
- Training speed improved 2x–3x (Tensor Core hardware acceleration)
- Very few code changes required; only three modifications
Example
model = CNN().to(device)
optimizer = optim.Adam(model.parameters(), lr=1e-3)
scaler = GradScaler() # FP16 gradient scaler to prevent gradients from underflowing to zero
for epoch in range(num_epochs):
model.train()
for inputs, labels in train_loader:
inputs = inputs.to(device, non_blocking=True)
labels = labels.to(device, non_blocking=True)
optimizer.zero_grad()
# Modification 1: automatically select FP16/FP32 within the autocast region
with autocast(device_type="cuda"):
outputs = model(inputs)
loss = criterion(outputs, labels)
# Modification 2: scale the loss with scaler before backpropagation
scaler.scale(loss).backward()
# Modification 3: scaler updates parameters and automatically handles gradient scaling internally
scaler.step(optimizer)
scaler.update()
7.1 Choosing Between FP16 and BF16
| Format | Precision bits | Exponent bits | Suitable hardware | Characteristics |
|---|---|---|---|---|
| float16 | 10 bits | 5 bits | RTX 20/30/40 series | Requires GradScaler to prevent overflow |
| bfloat16 | 7 bits | 8 bits | A100 / H100 / RTX 4090 | Same range as FP32, more stable training |
If your GPU supports BF16, use it preferentially; no GradScaler needed:
Example
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
outputs = model(inputs)
loss = criterion(outputs, labels)
loss.backward()
optimizer.step()
8. Performance Optimization Tips
This section introduces several techniques for optimizing memory usage and training speed.
8.1 GPU Memory Management
Example
print(f"Allocated: {torch.cuda.memory_allocated() / 1024**2:.1f} MB")
print(f"Cached: {torch.cuda.memory_reserved() / 1024**2:.1f} MB")
# Print detailed memory report
print(torch.cuda.memory_summary())
# Empty the cache pool (does not release allocated memory)
torch.cuda.empty_cache()
# Use inference_mode during inference (faster than no_grad, completely disables the gradient engine)
with torch.inference_mode():
output = model(x)
# Gradient checkpointing: trade recomputation for memory (commonly used in large model training)
from torch.utils.checkpoint import checkpoint
output = checkpoint(model_block, x)
8.2 DataLoader Optimization
Example
dataset,
batch_size=256,
num_workers=4, # Multi-process prefetching; recommended to set to 50% of CPU cores
pin_memory=True, # Pinned memory to speed up CPU -> GPU transfer
persistent_workers=True, # Keep worker processes alive to avoid recreating processes every epoch
prefetch_factor=2, # Number of batches prefetched per worker
)
8.3 torch.compile Model Compilation (PyTorch 2.0+)
One line of code compiles and optimizes the computation graph, improving speed by 30%–200% (depending on model architecture):
Example
model = torch.compile(model)
# Trade-offs between different compilation modes
model = torch.compile(model, mode="default") # Balanced, suitable for most scenarios
model = torch.compile(model, mode="reduce-overhead") # Reduces scheduling overhead
model = torch.compile(model, mode="max-autotune") # Maximum optimization, longer compilation time
8.4 Gradient Accumulation (Simulating Large Batches)
When memory is insufficient, gradient accumulation can simulate a larger batch size without actually increasing memory usage:
Example
for i, (inputs, labels) in enumerate(train_loader):
inputs = inputs.to(device)
labels = labels.to(device)
outputs = model(inputs)
loss = criterion(outputs, labels) / accumulation_steps # Divide loss equally
loss.backward() # Accumulate gradients without zeroing
if (i + 1) % accumulation_steps == 0:
optimizer.step()
optimizer.zero_grad() # Only zero gradients after parameter update
8.5 Strategies for Insufficient GPU Memory
| Strategy | Method | Effect |
|---|---|---|
| Reduce batch size | 256 -> 64 | Linearly reduces memory |
| Mixed precision | AMP + FP16/BF16 | Memory reduced by about 50% |
| Gradient accumulation | Update every N steps | Simulate large batch without increasing memory |
| Gradient checkpointing | torch.utils.checkpoint |
Greatly reduces memory, but reduces training speed |
| Freeze some parameters | Freeze backbone in transfer learning | Reduces memory usage of backpropagation |
9. Common Errors and Troubleshooting
This section summarizes the most common errors in GPU training and their solutions.
9.1 RuntimeError: Expected all tensors to be on the same device
Cause: The tensors involved in the operation are on different devices (one on CPU, one on GPU).
Example
print(inputs.device, labels.device, next(model.parameters()).device)
# Solution: ensure .to(device) is executed for every batch
inputs = inputs.to(device)
labels = labels.to(device)
9.2 RuntimeError: CUDA out of memory
Cause: Out of memory, commonly due to batch size being too large, model being too large, or accidental accumulation of computation graphs.
Example
# Incorrect code
total_loss += loss # loss is a tensor that holds the entire computation graph
# Correct code
total_loss += loss.item() # .item() extracts a Python scalar and releases the computation graph
# Other troubleshooting steps:
# 1. Reduce batch_size
# 2. Ensure the validation loop uses torch.no_grad()
# 3. Call torch.cuda.empty_cache() to clear cache
# 4. Use torch.cuda.memory_summary() to locate the source of GPU memory usage
9.3 Can't call numpy() on Tensor that requires grad
Reason: Tensors with gradients cannot be directly converted to numpy arrays.
Example
arr = gpu_tensor.numpy()
# Correct code
arr = gpu_tensor.detach().cpu().numpy()
9.4 CUDA error: device-side assert triggered
Reason: Usually it is label values out of bounds (e.g., 10 classes but the label value equals 10), or array index out of bounds. The error message is generated on the GPU, and the default display location is inaccurate.
Example
import os
os.environ["CUDA_LAUNCH_BLOCKING"] = "1"
9.5 Training Speed Not Improving (Low GPU Utilization)
Reason: Data loading becomes the bottleneck; the GPU spends most of its time waiting for the CPU to prepare data.
Troubleshooting steps:
- Run
nvidia-smiorwatch -n 1 nvidia-smiObserve GPU utilization- Utilization consistently below 80% indicates a data bottleneck
- Increase DataLoader's num_workers
- Enable pin_memory=True
- Move data preprocessing to the GPU (torchvision.transforms supports GPU operations)
- Preload small-file datasets into memory
9.6 Using PyTorch Profiler to Pinpoint Bottlenecks
Example
with profile(
activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
record_shapes=True,
) as prof:
for i, (inputs, labels) in enumerate(train_loader):
if i >= 10:
break
inputs = inputs.to(device)
labels = labels.to(device)
loss = criterion(model(inputs), labels)
loss.backward()
optimizer.step()
optimizer.zero_grad()
# Sort by CUDA time, print top 15 items
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=15))
10. API Quick Reference
The following is a quick reference table for common PyTorch GPU-related APIs.
10.1 Device Management
| Operation | Code |
|---|---|
| Detect CUDA availability | torch.cuda.is_available() |
| Select device | torch.device("cuda" if ... else "cpu") |
| Number of GPUs | torch.cuda.device_count() |
| GPU name | torch.cuda.get_device_name(0) |
| Set current GPU | torch.cuda.set_device(0) |
| Wait for GPU to complete | torch.cuda.synchronize() |
10.2 Tensor Operations
| Operation | Code |
|---|---|
| Move tensor to GPU | tensor.to(device)ortensor.cuda() |
| Move tensor to CPU | tensor.cpu() |
| View current device | tensor.device |
| Whether it is on GPU | tensor.is_cuda |
| Convert GPU tensor to numpy | tensor.detach().cpu().numpy() |
| Asynchronous transfer | tensor.to(device, non_blocking=True) |
10.3 Model Operations
| Operation | Code |
|---|---|
| Move model to GPU | model.to(device) |
| Enable training mode | model.train() |
| Enable inference mode | model.eval() |
| Compile acceleration | torch.compile(model) |
| Disable gradients | with torch.no_grad(): |
| Inference mode | with torch.inference_mode(): |
10.4 GPU Memory Management
| Operation | Code |
|---|---|
| Allocated GPU memory | torch.cuda.memory_allocated() |
| Cached GPU memory | torch.cuda.memory_reserved() |
| GPU memory report | torch.cuda.memory_summary() |
| Clear cache | torch.cuda.empty_cache() |
10.5 Core Principles Quick Reference
1. 用 device 变量统一管理,不要硬编码 "cuda" 2. 模型和数据必须在同一设备上,每个 batch 都要 .to(device) 3. 验证和推理时一定使用 torch.no_grad() 或 torch.inference_mode() 4. 生产训练推荐开启 AMP 混合精度,几乎免费获得 2x 加速 5. DataLoader 设置 pin_memory=True 和 num_workers >= 4 减少数据瓶颈 6. PyTorch 2.0+ 可以用 torch.compile(model) 一行提速 30%+ 7. GPU 计算是异步的,精确计时需要调用 torch.cuda.synchronize()Other extensions