Explore indexBack to Terms

Reinforcement Learning from Human Feedback

Reinforcement Learning from Human Feedback (RLHF) trains a reward model from pairwise human preference annotations and optimizes the policy model using reinforcement learning (such as PPO) with KL-divergence penalties. It is widely used to align frontier models for reasoning, helpfulness, and safety.

No public content is connected to this entity yet.