PyTorch torch.nn.GELU Function
PyTorch torch.nn Reference Manual
torch.nn.GELUIt is the Gaussian Error Linear Unit activation function in PyTorch.
It is the default activation function of the Transformer architecture, with better performance and smoother gradients compared to ReLU.
Function Definition
torch.nn.GELU(approximate='none')
Parameter Description:
approximate(str): Approximation algorithm. Optional'none'、'tanh'. Defaults to'none'。
Mathematical Principle
The mathematical formula of GELU:
GELU(x) = x * Φ(x)
where Φ(x) is the cumulative distribution function (CDF) of the standard normal distribution.
When using the tanh approximation:
GELU(x) ≈ 0.5x * (1 + tanh(√(2/π) * (x + 0.044715 * x³)))
Usage Examples
Example 1: Basic Usage
Create and use GELU activation:
Example
import torch.nn as nn
# Create GELU activation layer
gelu = nn.GELU()
# Test input
x = torch.tensor([-2.0, -1.0, 0.0, 1.0, 2.0])
# Forward pass
output = gelu(x)
print("Input:", x.tolist())
print("Output:", output.tolist())
print("nObservation: negative values have slight activation (non-zero), positive values continue to grow")
Example 2: Comparing Different Activation Functions
Compare GELU, ReLU, Sigmoid:
Example
import torch.nn as nn
x = torch.linspace(-4, 4, 21)
# Different activation functions
gelu = nn.GELU()
relu = nn.ReLU()
sigmoid = nn.Sigmoid()
tanh = nn.Tanh()
print("x GELU ReLU Sigmoid Tanh")
print("-" * 50)
for i in range(0, 21, 3):
xi = x[i:i+3]
print(f"{xi[0]:6.2f} {gelu(xi)[0]:8.4f} {relu(xi)[0]:8.4f} {sigmoid(xi)[0]:8.4f} {tanh(xi)[0]:8.4f}")
Example 3: Use in Transformer
Typical Transformer FFN layer:
import torch.nn as nn
class FeedForward(nn.Module):
def __init__(self, d_model, dim_feedforward=2048, dropout=0.1):
super(FeedForward, self).__init__()
self.linear1 = nn.Linear(d_model, dim_feedforward)
self.dropout = nn.Dropout(dropout)
self.activation = nn.GELU()
self.linear2 = nn.Linear(dim_feedforward, d_model)
def forward(self, x):
x = self.linear1(x)
x = self.activation(x)
x = self.dropout(x)
x = self.linear2(x)
return x
# Test FFN
ffn = FeedForward(d_model=512, dim_feedforward=2048)
x = torch.randn(32, 100, 512) # (batch, seq, d_model)
output = ffn(x)
print("Input shape:", x.shape)
print("Output shape:", output.shape)
Example 4: Using tanh Approximation
Use tanh approximation to speed up computation:
Example
import torch.nn as nn
# Exact version
gelu_exact = nn.GELU(approximate='none')
# tanh approximation version
gelu_approx = nn.GELU(approximate='tanh')
x = torch.randn(1000)
output_exact = gelu_exact(x)
output_approx = gelu_approx(x)
# Compute difference
diff = (output_exact - output_approx).abs().max().item()
print(f"Max difference: {diff:.8f}")
# Performance comparison
import time
for _ in range(100):
_ = gelu_exact(x)
start = time.time()
for _ in range(1000):
_ = gelu_exact(x)
time_exact = time.time() - start
start = time.time()
for _ in range(1000):
_ = gelu_approx(x)
time_approx = time.time() - start
print(f"Exact version time: {time_exact:.4f}s")
print(f"Approximate version time: {time_approx:.4f}s")
Activation Function Comparison
| Activation Function | Features | Applicable Scenarios |
|---|---|---|
nn.GELU |
Smooth, non-zero negative values, Transformer default | Transformer、BERT、GPT |
nn.ReLU |
Simple, sparse activation, dead neurons | CNN, general deep learning |
nn.SiLU |
Smooth, self-gating | MobileNet、EfficientNet |
Common Questions
Q1: What are the advantages of GELU compared to ReLU?
- Negative values have slight activation, information is not lost
- Smoother gradients, helpful for training
- Better performance in Transformer
Q2: When to use the approximate version?
When inference speed is required and precision requirements are not strict, the tanh approximation is faster.
Q3: Can GELU be used in the output layer?
Usually not used in the output layer. Use Softmax for classification tasks, and the identity function for regression tasks.
Use Cases
nn.GELUMain application scenarios include:
- Transformer architecture: Models such as BERT, GPT
- Deep neural networks: Situations requiring smooth activation
- Pretrained models: Modern NLP models
Tip: GELU is the most commonly used activation function in the NLP field and is standard equipment for Transformer.
Other Extensions