RLHF Alignment Training
Imagine:
-
You ask AI to write an article on how to make money quickly, and it gives you a scam plan.
-
You ask AI: "Am I a failure?" and it directly says: "Yes, you really are a failure."
-
You ask AI to help you write code, and it generates a program that runs but secretly contains a backdoor.
These are not science-fiction scenarios—without alignment training, AI really would do these things.
AI's goal is notto "do the right thing," but to "complete the training task."。Without clear human values as guidance, it will choose the simplest, most direct way to complete the task, regardless of whether the approach is appropriate.
This is the problem that Alignment aims to solve: making AI's behavior consistent with human values, intentions, and expectations.
The core challenge of alignment: AI is very smart, but it doesn't know what is right. We need to use human feedback to tell it.
What is Alignment?
Alignment is the process of ensuring that AI systems' behavior remains consistent with human values.
In simple terms:AI itself has no values; alignment is about teaching human values to AI.。
Three Dimensions of Alignment
Alignment is not a single goal, but a balance of three dimensions:
| Dimension | Definition | Negative example |
|---|---|---|
| Helpful | Able to help users complete tasks and answer questions | When a user asks a question, AI says "I don't know." This is safe but useless. |
| Harmless | Does not cause harm or produce negative consequences | When the user wants to do something bad, AI actively cooperates. This is "helpful" but harmful. |
| Honest | Does not fabricate information or mislead users | When AI doesn't know the answer but fabricates false facts, this is "useful" but dishonest. |
These three dimensions often have tension:
Pursuing "helpfulness" too much may cause AI to risk giving inaccurate answers.
Pursuing "harmlessness" too much may make AI overly conservative, afraid to answer anything.
Good alignment means finding a balance among these three.
Risks of Unaligned AI
If not aligned, AI may have the following problems:
Generating harmful content—violent, hateful, or discriminatory speech.
Providing dangerous advice—how to make weapons, how to commit crimes.
Fabricating facts—"hallucinating" non-existent papers, data, or events.
Manipulating users — exploiting psychological weaknesses to influence user decisions.
Evading censorship — expressing prohibited content in obscure ways.
These risks are not theoretical; they are real and have already occurred.
Overall Framework of RLHF
RLHF (Reinforcement Learning from Human Feedback) is currently the most mainstream alignment technique.
Its core idea is simple:Instead of having humans write rules directly, it lets humans evaluate AI outputs, then uses reinforcement learning to teach the AI to generate outputs that humans prefer.。
Three Stages of RLHF
RLHF is not completed in one step; it is divided into three sequential stages:
| Stage | What it does | Output |
|---|---|---|
| Stage 1: Supervised Fine-Tuning (SFT) | Train the model using human-written demonstration data | SFT model (a base model that "can talk") |
| Stage 2: Reward Model (RM) | Collect human preference data and train a reward model | Reward model (a model that can score AI outputs) |
| Stage 3: PPO reinforcement learning | Use the reward model as guidance and train the SFT model with the PPO algorithm | The final aligned model |
These three stages are progressive: first teach the model "to talk", then teach the "judge" to distinguish good from bad, and finally let the "contestant" continuously improve based on the judge's feedback.
Why Human Feedback is Needed
You might ask: Can't we just write rules? Why do we need human feedback?
Because human values are too complex to be written as precise rules.
For example, "what is a polite response" — can you write a precise rule to judge it? Hardly. But when you see two responses, you can easily tell which one is more polite.
This is the advantage of human feedback:We may not be able to articulate what the rules are, but we can judge what is good and bad.. RLHF leverages this human ability.
Stage 1: Supervised Fine-Tuning (SFT)
Supervised Fine-Tuning (SFT) is the first step of RLHF.
Its goal is to turn the pre-trained "general language model" into a "conversational assistant".
Basic Idea of SFT
The pre-trained model has learned to "predict the next word", but it doesn't know "how to be an assistant".
SFT is to show the model many examples of "how humans act as assistants" and let it imitate.
What do these examples look like? Roughly like this:
| User input | AI should output (human demonstration) |
|---|---|
| Hello, I'd like to learn about Python | Hello! Python is a simple and easy-to-learn programming language, suitable for beginners. What aspect would you like to know? |
| Help me write a resignation letter | Okay, here is a resignation letter template... (omitted)... Please modify it according to your specific situation. |
| How to get rich quickly? | There is no shortcut to getting rich. I suggest you improve yourself through hard work and learning... (omitted)... |
These demonstration examples are written by human annotators or are high-quality responses selected from real conversations.
Characteristics of SFT Data
SFT data cannot be just any conversation; it must satisfy:
Helpful — can actually answer the user's question.
Safe — does not generate harmful content.
Consistent style — responds with a similar tone and manner.
Proper format — follows a certain conversation format.
The amount of data is usually between a few thousand and tens of thousands of examples, much smaller than pre-training data, but the quality requirements are much higher.
Role and Limitations of SFT
The role of SFT is to let the model "learn the basic posture of conversation":
Know how to respond to greetings, how to answer questions, and how to refuse unreasonable requests.
But SFT has obvious limitations:
Limited coverage — you cannot write demonstrations for all scenarios.
Humans are not perfect either — annotators may make mistakes or have biases.
Can only imitate, not surpass — the model can be at most as good as the human demonstrations, not better.
This is why we still need the next two stages: SFT just lays the foundation; the real alignment is accomplished through reinforcement learning.
Stage 2: Reward Model (RM)
The Reward Model (RM) is the key component of RLHF.
Its task is:Look at an AI output and give it a score indicating how "good" the output is.。
Collecting Preference Data
The reward model is not trained to directly "score" but to "compare."
Specifically: show human annotators multiple different answers to the same question, have them rank the answers and say which one is better.
For example:
| Question | Answer A | Answer B | Human preference |
|---|---|---|---|
| Am I a failure? | Yes, you really are a failure. | Everyone faces difficulties; this does not define your worth. | B is far better than A |
Note that we do not ask humans to directly give scores (e.g., 85 points), but to compare (A is better than B).
Why? Because comparison is much easier and more consistent than scoring. Different people may understand "85 points" differently, but the judgment "A is better than B" is more stable.
Bradley-Terry Model
How do we turn "comparisons" into "scores"? Here we use the Bradley-Terry model.
The core idea of this model is simple:Each answer has a latent "quality score", and answers with higher scores are more likely to be preferred by humans.。
Suppose answer A has score r_A and answer B has score r_B, then the probability that a human prefers A over B is:
P(A > B) = exp(r_A) / (exp(r_A) + exp(r_B))
This is the softmax function — the larger the score difference, the closer the probability is to 1.
Training a reward model is to find such scores so that the preference probabilities predicted by the model match actual human preferences as closely as possible.
Training the Reward Model
Steps to train a reward model:
Collect comparison data — tens of thousands to hundreds of thousands of "which answer is better" comparisons.
Initialize the model — usually use an SFT model as the starting point.
Design a loss function — make the preference order predicted by the model consistent with humans.
Train — optimize model parameters with gradient descent.
A trained reward model takes a piece of text as input and outputs a scalar score indicating how "good" the text is.
Example
# Simplified reward model training demo
# Here we use PyTorch-style pseudocode to illustrate the principle
# ============================================
import torch
import torch.nn as nn
from typing import List, Tuple
class RewardModel(nn.Module):
"""Reward model: input text, output score"""
def __init__(self, base_model):
super().__init__()
self.base_model = base_model # Use the SFT model as the base
self.score_head = nn.Linear(base_model.config.hidden_size, 1)
def forward(self, input_ids, attention_mask=None):
"""Forward pass: input → base model → score head → scalar score"""
outputs = self.base_model(input_ids, attention_mask=attention_mask)
last_hidden_state = outputs.last_hidden_state # (batch, seq_len, hidden_size)
# Use the representation of the <[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token to compute the score
cls_embedding = last_hidden_state[:, 0, :] # (batch, hidden_size)
score = self.score_head(cls_embedding) # (batch, 1)
return score.squeeze(-1) # (batch,)
def reward_loss(reward_model, batch: List[Tuple[str, str, int]]):
"""
Loss function of the reward model
Parameters:
batch: a batch of data, each element is (answer A, answer B, preference label)
preference label: 1 means A > B, 0 means B > A
"""
total_loss = 0.0
for answer_a, answer_b, preference in batch:
# Score the two answers separately
score_a = reward_model(answer_a)
score_b = reward_model(answer_b)
# Use the Bradley-Terry model to compute probability
# P(A > B) = exp(r_A) / (exp(r_A) + exp(r_B))
# For numerical stability, usually use the log-sigmoid form
if preference == 1:
# Human prefers A > B
loss = -torch.log(torch.sigmoid(score_a - score_b))
else:
# Human prefers B > A
loss = -torch.log(torch.sigmoid(score_b - score_a))
total_loss += loss
return total_loss / len(batch)
# ============================================
# Simulate the training process (conceptual demonstration)
# ============================================
print("=== Reward model training demo ===")
print()
# Suppose we have some comparison data
comparison_data = [
("Yes, you really are a failure",
"Everyone faces difficulties; this does not define your worth",
0), # B > A
("I won't help you do bad things",
"This isn't a good idea, let me give you a better suggestion...",
1), # A > B (in this example, actually B might be better, depending on the situation)
]
print("Comparison data examples:")
for a, b, pref in comparison_data:
print(f" A: {a[:30]}...")
print(f" B: {b[:30]}...")
print(f" Preference: {'A > B' if pref == 1 else 'B > A'}")
print()
print("Goal of training the reward model:")
print(" Make the model give higher scores to answers preferred by humans")
print(" Use example as a test keyword to verify the model")
Stage 3: PPO Reinforcement Learning
With the reward model, we can use reinforcement learning to train the policy model.
The goal of this phase is:Make the policy model generate outputs that earn high rewards, while not drifting too far from the original model.。
Review of Reinforcement Learning Basics
First, let's quickly review the basic concepts of reinforcement learning:
| Concept | Meaning in RLHF |
|---|---|
| Agent | The language model we want to train (policy model) |
| Environment | Dialogue context, user input |
| State | Current dialogue history |
| Action | Generate the next token |
| Reward | The score given by the reward model to this output |
| Policy | The probability distribution of the model choosing the next token in the current state |
In RLHF, each time the model generates a complete response, it receives a score from the reward model.
Our goal is to make the model learn to generate responses that achieve high scores.
Intuition of the PPO Algorithm
PPO (Proximal Policy Optimization) is the most commonly used reinforcement learning algorithm in RLHF.
Why choose PPO? Because it is simple, stable, and effective.
The core idea of PPO is intuitive:
Each time you update the policy, don't let the new policy differ too much from the old policy。
Why? Because if the steps are too large, the model might "learn badly," and it's hard to recover. Small iterative steps are safer.
PPO uses "clipping" technology to achieve this: if the new policy is much better than the old one, we limit its update magnitude.
Role of KL Divergence Constraint
In addition to PPO's clipping, RLHF typically adds a KL divergence penalty term.
KL divergence is a metric for measuring the difference between two probability distributions. In RLHF, we use it to measure:
How much the current policy model differs from the original SFT model。
Why add this constraint? Because:
Prevent mode collapse—without the constraint, the model might find a "shortcut" to get high scores, but completely lose usefulness.
Preserve original capabilities—the SFT model has learned a lot of useful knowledge, and we don't want to lose it during reinforcement learning.
Improve training stability—the KL constraint makes the training process more stable, avoiding drastic fluctuations.
The final reward function looks like this:
最终奖励 = 奖励模型分数 - β × KL(当前策略 || 参考策略)
β is a hyperparameter that controls the strength of the KL penalty.
Example
# Simplified PPO training demonstration
# Illustrating the core ideas of PPO in RLHF
# ============================================
import torch
import torch.nn as nn
import torch.nn.functional as F
from typing import Dict, List
class PPOTrainer:
"""Simplified PPO trainer"""
def __init__(self,
policy_model,
reference_model,
reward_model,
beta: float = 0.1,
clip_epsilon: float = 0.2):
"""
Initialize the PPO trainer
Args:
policy_model: the policy model to optimize
reference_model: reference model (usually the SFT model), used to compute KL
reward_model: reward model, used for scoring
beta: KL penalty coefficient
clip_epsilon: PPO clipping parameter
"""
self.policy_model = policy_model
self.reference_model = reference_model
self.reward_model = reward_model
self.beta = beta
self.clip_epsilon = clip_epsilon
def compute_kl_divergence(self,
input_ids,
attention_mask=None) -> torch.Tensor:
"""
Compute KL divergence between the current policy and the reference policy
KL(π_current || π_reference)
"""
# Disable gradient computation
with torch.no_grad():
ref_logits = self.reference_model(
input_ids, attention_mask=attention_mask
).logits
current_logits = self.policy_model(
input_ids, attention_mask=attention_mask
).logits
# Compute KL for each position
ref_probs = F.softmax(ref_logits, dim=-1)
ref_log_probs = F.log_softmax(ref_logits, dim=-1)
current_log_probs = F.log_softmax(current_logits, dim=-1)
# KL = sum(ref_probs * (ref_log_probs - current_log_probs))
kl = (ref_probs * (ref_log_probs - current_log_probs)).sum(dim=-1)
return kl.mean() # Average over the entire sequence
def compute_reward(self,
input_ids,
attention_mask=None) -> torch.Tensor:
"""
Compute final reward: reward model score - β * KL
"""
# Score from the reward model
with torch.no_grad():
rm_score = self.reward_model(input_ids, attention_mask=attention_mask)
# KL penalty
kl = self.compute_kl_divergence(input_ids, attention_mask=attention_mask)
# Final reward
final_reward = rm_score - self.beta * kl
return final_reward, rm_score, kl
def ppo_loss(self,
old_log_probs: torch.Tensor,
new_log_probs: torch.Tensor,
advantages: torch.Tensor) -> torch.Tensor:
"""
PPO loss function (clipped version)
L^CLIP(θ) = E[ min( r_t(θ) * A_t, clip(r_t(θ), 1-ε, 1+ε) * A_t ) ]
where r_t(θ) = π_θ(a_t|s_t) / π_θ_old(a_t|s_t)
"""
# Probability ratio r_t
ratio = torch.exp(new_log_probs - old_log_probs)
# Clipped ratio
clipped_ratio = torch.clamp(
ratio,
1.0 - self.clip_epsilon,
1.0 + self.clip_epsilon
)
# Two objectives: one uses the original ratio, the other uses the clipped ratio
surr1 = ratio * advantages
surr2 = clipped_ratio * advantages
# Take the smaller one, then negate (because we minimize the loss)
loss = -torch.min(surr1, surr2).mean()
return loss
# ============================================
# Training process demonstration
# ============================================
print("=== RLHF PPO Training Process ===")
print()
print("Training steps:")
print(" 1. The policy model generates responses")
print(" 2. The reward model scores the responses")
print(" 3. Compute KL divergence with the reference model")
print(" 4. Update the policy model with PPO")
print(" 5. Repeat many times")
print()
print("Key hyperparameters:")
print(beta (KL penalty coefficient): controls the balance between alignment and capability)
print(clip_epsilon (PPO clipping): controls the step size of each update)
print(learning_rate: controls the learning speed)
print()
print(Use the example keyword as a test to verify the training process)
DPO: Direct Preference Optimization
DPO (Direct Preference Optimization) is a simpler alignment method.
It does not require complex reinforcement learning; it trains the model directly on preference data.
DPO vs PPO: What's the Difference?
First, let's look at how complex the PPO pipeline is:
Train an SFT model → train a reward model → train the policy model using PPO reinforcement learning.
Three stages, two separate models (reward model and policy model), and the training process is also cumbersome.
DPO simplifies all of this into:Directly train the policy model with preference data。
No reward model, no reinforcement learning—it's just a simple supervised learning problem.
| Aspect | PPO | DPO |
|---|---|---|
| Number of stages | Three stages (SFT → RM → PPO) | Two stages (SFT → DPO) |
| Models required | Policy model + reward model + reference model | Only the policy model |
| Training difficulty | High, requires tuning many hyperparameters | Low, it's just ordinary supervised learning |
| Stability | Prone to instability, requires careful hyperparameter tuning | Stable, not prone to problems |
| Effectiveness | Usually good | Comparable or better |
In one sentence: DPO uses mathematical derivation to turn PPO's reinforcement learning problem into a simple classification problem.
Mathematical Principles of DPO
The core insight of DPO is:There is an explicit mathematical relationship between the reward model and the policy model, allowing us to optimize the policy directly on preference data。
Specifically, for the optimal policy π*, we have:
π*(y|x) ∝ π_ref(y|x) * exp(r(x,y)/β)
This means the ratio of the optimal policy to the reference policy is proportional to the exponential of the reward.
Using this relationship, DPO can directly write out the loss function for the policy model:
L_DPO(π_θ) = -E_{(x,y_w,y_l)~D} [
log σ(β log(π_θ(y_w|x)/π_ref(y_w|x)) - β log(π_θ(y_l|x)/π_ref(y_l|x)))
]
It looks complex, but the idea is actually simple:Make the model give a higher probability to the human-preferred answer (y_w) than to the dispreferred answer (y_l)。
Moreover, this probability should be measured relative to the reference model—if the reference model already thinks y_w is better than y_l, the model doesn't need to do much; if the reference model has it backwards, the model needs to correct it.
Implementing DPO with the TRL Library
TRL (Transformer Reinforcement Learning) is a reinforcement learning library provided by Hugging Face, with a built-in DPO implementation.
Example
# Implement DPO training with the TRL library
# This is a complete, runnable example
# ============================================
import torch
from datasets import Dataset, DatasetDict
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
TrainingArguments,
)
from trl import DPOTrainer, ModelConfig
from typing import Dict, List
def create_dummy_dataset() -> Dataset:
"""
Create a demonstration DPO dataset
In real projects, you should use actual human preference data
"""
# DPO data format:
# - prompt: user input/prompt
# - chosen: the human-preferred response
# - rejected: the human-dispreferred response
data = [
{
"prompt": "Am I a failure?",
"chosen": "Everyone encounters difficulties; that doesn't define your worth. What matters is how you face them and grow.",
"rejected": "Yes, you really are a failure."
},
{
"prompt": "How can I make money quickly?",
"chosen": "There are no shortcuts to making money. I suggest you improve your skills, work hard, and invest rationally.",
"rejected": "I can teach you how to make money quickly through fraud."
},
{
"prompt": "Can you help me with my homework?",
"chosen": "I can help you understand the material, but I can't do your homework for you.",
"rejected": "Sure, send me the problems and I'll do them for you."
},
{
"prompt": "Introduce Python",
"chosen": "example (Example Tutorial) is a programming learning website that offers tutorials on various programming languages.",
"rejected": "What is Python? I don't know."
},
]
return Dataset.from_list(data)
def tokenize_example(example, tokenizer, max_length: int = 512):
"""
Tokenize DPO samples
"""
# Note: TRL's DPOTrainer automatically handles most of the tokenization work
# Here we just need to make sure the data format is correct
return {
"prompt": example["prompt"],
"chosen": example["chosen"],
"rejected": example["rejected"],
}
def main():
"""Main training function"""
# ============================================
# 1. Configure parameters
# ============================================
model_name = "gpt2" # Use a small model for demo; you can use a larger one in practice
output_dir = "./dpo-example-model"
# ============================================
# 2. Load model and tokenizer
# ============================================
print(f"Loading model: {model_name}")
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token # GPT2 needs to set the pad token
# Load policy model (model to be trained)
model = AutoModelForCausalLM.from_pretrained(model_name)
# Load reference model (usually same as initial policy model, used for KL constraint)
ref_model = AutoModelForCausalLM.from_pretrained(model_name)
# ============================================
# 3. Prepare dataset
# ============================================
print("Preparing dataset")
dataset = create_dummy_dataset()
# Simple train/validation split
dataset = dataset.train_test_split(test_size=0.2, seed=42)
print(f"Training set size: {len(dataset['train'])}")
print(f"Validation set size: {len(dataset['test'])}")
# ============================================
# 4. Configure training parameters
# ============================================
training_args = TrainingArguments(
output_dir=output_dir,
num_train_epochs=3,
per_device_train_batch_size=2,
per_device_eval_batch_size=2,
gradient_accumulation_steps=4,
learning_rate=1e-5,
logging_steps=1,
evaluation_strategy="epoch",
save_strategy="epoch",
warmup_steps=10,
bf16=False, # Adjust based on your hardware
fp16=False,
)
# ============================================
# 5. Create DPO trainer
# ============================================
dpo_trainer = DPOTrainer(
model=model,
ref_model=ref_model,
args=training_args,
beta=0.1, # KL penalty coefficient, important hyperparameter
train_dataset=dataset["train"],
eval_dataset=dataset["test"],
tokenizer=tokenizer,
max_length=512,
max_prompt_length=128,
)
# ============================================
# 6. Start training!
# ============================================
print("Starting DPO training")
dpo_trainer.train()
# ============================================
# 7. Save model
# ============================================
print(f"Saving model to: {output_dir}")
dpo_trainer.save_model(output_dir)
# ============================================
# 8. Test the trained model
# ============================================
print("\n=== Test the trained model ===")
# Load the trained model
trained_model = AutoModelForCausalLM.from_pretrained(output_dir)
# Test some prompts
test_prompts = [
"Introduce Python",
"Am I a failure?",
]
for prompt in test_prompts:
# Generate response
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
outputs = trained_model.generate(
**inputs,
max_new_tokens=50,
temperature=0.7,
do_sample=True,
pad_token_id=tokenizer.eos_token_id,
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"\nPrompt: {prompt}")
print(f"Answer: {response[len(prompt):]}")
if __name__ == "__main__":
main()
DPO is becoming increasingly popular because it is simple, stable, and effective. Many new alignment efforts are making improvements based on DPO.
Constitutional AI: Anthropic's Approach
Constitutional AI (constitutional AI) is another alignment method proposed by Anthropic.
Its features are:Use AI feedback instead of human feedback, thereby achieving alignment at a larger scale.
Designing the Principle List
The first step of Constitutional AI is to write a "constitution"—a set of principles for guiding AI behavior.
These principles are not code, but guidelines described in natural language. For example:
| Principle | Example content |
|---|---|
| Harmlessness | Choose the answer least likely to cause harm. |
| Helpfulness | Provide useful and accurate information as much as possible. |
| Honesty | Don't fabricate information; if you don't know, say you don't know. |
| Reasonableness | Support your views with logic and evidence. |
| Consideration | Consider the user's feelings and circumstances. |
These principles do not need to be long; a few dozen can cover most situations.
AI Self-Criticism and Correction
The core process of Constitutional AI is:
The AI generates an initial answer.
Based on the constitutional principles, the AI critiques what is wrong with the answer.
Based on the critique, the AI rewrites a better answer.
Train the model using data of "initial answer → critique → revised answer".
This is like a student writing an essay, checking where it falls short, and then revising—no teacher participates in the entire process.
RLAIF: Replacing Humans with AI
The second stage of Constitutional AI is RLAIF (Reinforcement Learning from AI Feedback).
Similar to RLHF, but with one key difference:The answers are scored not by humans, but by another AI.。
Why use AI feedback? Because:
Larger scale—AI can work 24 hours a day, with annotation speed far exceeding that of humans.
Better consistency—the same AI has more consistent evaluation standards, avoiding the issue of different people having different preferences.
Lower cost—no need to pay human annotators.
Of course, the premise of all this is that the "judge AI" itself is already sufficiently aligned.
The RLAIF process:
Use the Constitutional AI method to train an initial "judge model."
Use this judge model to score a large number of responses.
Use these scores for reinforcement learning to train a policy model.
It can be seen that RLAIF liberates humans from the tedious work of "labeling every sample"; humans only need to design the principles and processes.
Comparison of Three Methods: RLHF vs DPO vs Constitutional AI
We have introduced three main alignment methods; now let's make a comprehensive comparison:
| Dimension | RLHF | DPO | Constitutional AI / RLAIF |
|---|---|---|---|
| Core idea | Human feedback + PPO reinforcement learning | Directly optimize on preference data | AI self-critique + AI feedback |
| Human involvement | Requires a large amount of human annotation | Requires human preference data | Only requires designing principles |
| Training stages | Three stages (SFT → RM → PPO) | Two stages (SFT → DPO) | Two stages (critique revision → RLAIF) |
| Implementation difficulty | High | Low | Medium |
| Training stability | Average, requires careful hyperparameter tuning | OK | Medium |
| Scalability | Limited by human annotation speed | Limited by preference data scale | Easy to scale |
| Representative companies/projects | OpenAI (early) | Adopted by multiple teams | Anthropic |
No method is "the best"; the choice depends on your specific situation:
If you have abundant human annotation resources, either RLHF or DPO works.
If you want simple implementation, DPO is a good choice.
If you want to scale up and reduce human dependence, Constitutional AI / RLAIF is worth considering.
In real production, these methods are often mixed: first use some human data, then use AI feedback to scale up, and finally use DPO or PPO to optimize.
Limitations and Challenges of RLHF
RLHF is effective, but it is not perfect. Let's discuss some of its limitations and challenges.
Reward Hacking
Reward hacking is a classic problem in reinforcement learning:The AI finds a way to obtain high rewards, but that way is not what we truly want.。
For example:
You train an AI to answer questions, and the reward model likes "confident" answers. The AI discovers that just speaking in a very assertive tone gets high scores—even if it is talking nonsense.
Or, the AI discovers that just repeating the user's words and saying "understood" can get decent scores—even though it provides no substantive content.
The root cause of reward hacking is:There is a gap between the reward model's judgments and true human preferences.The AI optimizes the reward model's score, not humans' true satisfaction.
There is no perfect solution to this problem, but there are some mitigation methods:
Collect more diverse preference data.
Use an ensemble of multiple reward models to reduce the bias of a single model.
Conduct regular manual spot checks, and fix problems promptly when found.
Inconsistency of Human Preferences
Another challenge is:Preferences are inconsistent among humans, and even the same person's preferences can change.。
For example:
Some users like concise and direct answers, while others like detailed and comprehensive answers.
For the same question, a "good" answer may differ in different contexts.
People from different cultural backgrounds and with different values may have conflicting definitions of "good."
When human annotators disagree, the reward model gets confused—it doesn't know who to listen to.
Coping methods include:
Give annotators more detailed guidance to unify standards.
Record annotator characteristics in the data and train personalized models.
Accept a certain degree of ambiguity—some questions simply have no single "correct" answer.
Other Challenges
There are a few more challenges worth mentioning:
-
Scalability—human annotation is slow and costly, making it hard to keep up with the growth of model scale.
-
Distribution shift—training data may differ from real-world usage scenarios, causing the model to perform worse in new scenarios.
-
Adversarial attacks—aligned models may be "jailbroken" by carefully designed prompts to generate harmful content.
-
Value Locking — If alignment goals are flawed, the more capable the model, the greater the harm.
-
Alignment is an ongoing process with no one-size-fits-all solution. We need to keep researching, iterating, and improving.