Skip to content
Learn
News
Tools
Jobs
Mission
Search
⌘K
English
Donate
Menu
×
Start here
Start learning
Learn
Go to Learn
Courses
Guides
AI Tutor
Prompt library
Quizzes
Certification
Glossary
News
Go to News
Latest AI news
Topic trackers
Research & datasets
Blog
Editorial standards
Corrections
Tools
Go to Tools
Tool directory
Best-of lists
Compare tools
Cost calculator
Prompt Refiner
Submit AI Tool
Jobs
Go to Jobs
Browse AI jobs
Post a job
Sponsor AIU
Mission
Go to Mission
Mission
Impact & stats
Authors
Help centre
What's new
Support
Donate
Home
/
Glossary
/
Reinforcement Learning from Human Feedback (RLHF)
AI Glossary Term
What does “
Reinforcement Learning from Human Feedback (RLHF)
” mean?
Definition
A training method that uses human preference signals to shape model behavior.
Related terms
Reinforcement Learning
Training by reward signals where an agent learns actions that maximize long-term return.
Reward Model
A model that scores outputs based on preference signals, often used in RLHF pipelines.
DPO (Direct Preference Optimization)
A training method that fine-tunes models directly on preference pairs without needing a separate reward model.
Continual Learning
Training approaches that let a model keep learning from new data without forgetting prior knowledge.
Few-Shot Learning
Learning or adapting behavior from only a small number of examples.
Model Drift
Performance degradation over time as real-world conditions diverge from training assumptions.
Learn more in our free guides
Deep Learning
Reinforcement Learning
Machine Learning Basics
Supervised Learning
See also
Retrieval
Red Teaming
Recommendation System
Robustness
Recall
Safety Filter
RAG (Retrieval-Augmented Generation)
Scaling Law
Previous
Reinforcement Learning
Next
Retrieval
Browse the full AI Glossary
Explore all free AI guides