PyTorch torch.nn.Dropout Function

PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual


torch.nn.DropoutIt is a module in PyTorch used for regularization.

It prevents overfitting by randomly zeroing input elements to reduce co-adaptation between neurons.

Function Definition

torch.nn.Dropout(p=0.5, inplace=False)

Parameter Description:

  • p(float): The probability of each element being zeroed. Defaults to 0.5.
  • inplace(bool): Whether to perform the operation in-place. Defaults to False.

Usage Examples

Example 1: Basic Usage

Create and use a Dropout layer:

Example

import torch
import torch.nn as nn

# Create Dropout layer with drop probability 0.5
dropout = nn.Dropout(p=0.5)

# Training mode (Dropout active)
dropout.train()

# Create input
input_tensor = torch.ones(1, 10)
print("Input:", input_tensor.squeeze().tolist())

# Forward pass multiple times, observe randomness
for i in range(3):
    output = dropout(input_tensor)
    print(f"Output {i+1}:", output.squeeze().tolist())

It can be seen that approximately half of the elements are randomly zeroed on each call.

Example 2: Training vs Evaluation Mode

Dropout behaves differently during training and evaluation:

Example

import torch
import torch.nn as nn

dropout = nn.Dropout(p=0.5)

# Training mode
dropout.train()
train_output = dropout(torch.ones(4, 10))
print("Training mode - activation ratio:", (train_output != 0).float().mean().item())

# Evaluation mode
dropout.eval()
eval_output = dropout(torch.ones(4, 10))
print("Evaluation mode - activation ratio:", (eval_output != 0).float().mean().item())
print("Evaluation mode output:", eval_output[0].tolist())

During evaluation, Dropout has no effect, and the output remains unchanged.

Example 3: Use in Neural Networks

A typical fully connected network with Dropout:

Example

import torch
import torch.nn as nn

class DropoutNet(nn.Module):
    def __init__(self, input_dim=784, hidden_dim=256, output_dim=10, dropout_rate=0.5):
        super(DropoutNet, self).__init__()
        self.fc1 = nn.Linear(input_dim, hidden_dim)
        self.dropout1 = nn.Dropout(p=dropout_rate)
        self.fc2 = nn.Linear(hidden_dim, hidden_dim)
        self.dropout2 = nn.Dropout(p=dropout_rate)
        self.fc3 = nn.Linear(hidden_dim, output_dim)
        self.relu = nn.ReLU()

    def forward(self, x):
        x = self.relu(self.fc1(x))
        x = self.dropout1(x)  # First Dropout
        x = self.relu(self.fc2(x))
        x = self.dropout2(x)  # Second Dropout
        x = self.fc3(x)
        return x

model = DropoutNet()

# Training mode
model.train()
input_data = torch.randn(32, 784)
output = model(input_data)
print("Training mode output shape:", output.shape)

# Evaluation mode
model.eval()
output = model(input_data)
print("Evaluation mode output shape:", output.shape)

Example 4: Using Dropout2d in CNN

nn.Dropout2dDrop entire feature maps by channel:

Example

import torch
import torch.nn as nn

# Dropout2d drops by channel
dropout2d = nn.Dropout2d(p=0.5)

# Input: batch=1, channels=4, height=4, width=4
input_tensor = torch.ones(1, 4, 4, 4)
dropout2d.train()

output = dropout2d(input_tensor)
print("Dropout2d output shape:", output.shape)
print("Number of non-zero channels:", (output.sum(dim=(2, 3)) != 0).sum().item())

Example 5: Effects of Different Dropout Rates

The effect of the dropout rate on the network:

Example

import torch
import torch.nn as nn

for p in [0.1, 0.3, 0.5, 0.7]:
    dropout = nn.Dropout(p=p)
    dropout.train()

    # Average over multiple runs
    total_active = 0
    for _ in range(100):
        output = dropout(torch.ones(1000))
        total_active += (output != 0).float().sum().item()

    avg_active = total_active / 100 / 1000
    print(f"p={p} - average activation ratio: {avg_active:.2%} (expected: {1-p:.2%})")

Dropout Type Comparison

Type Dropout Method Applicable Scenarios
nn.Dropout Randomly zero individual elements Fully connected layers, feature vectors
nn.Dropout2d Randomly zero entire channels Convolutional layer feature maps
nn.Dropout3d Randomly zero entire 3D channels 3D convolutional features

FAQ

Q1: How to choose the Dropout rate?

  • 0.1-0.3: Lighter regularization, suitable for large datasets
  • 0.4-0.5: Common default values
  • 0.5+: Stronger regularization, suitable for small datasets

Q2: Where should Dropout be placed?

Usually placed after the fully connected layer and after the activation function. It can also be placed before the activation function.

Q3: Do I need to disable Dropout during evaluation?

Yes, when evaluating, use `model.eval()`model.eval()to automatically disable Dropout.


Use Cases

nn.DropoutMain use cases include:

  • Prevent overfitting: Reduce dependency between neurons
  • Model ensembling: Approximate the effect of multiple networks
  • Fully connected layers: Most commonly used in FC layers
  • Feature dropout: Improve model robustness

Note: Dropout is enabled during training; during evaluation, be sure to switch to eval mode, otherwise the output will be unstable.


PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual

Other Extensions